Flame graph
Also called: CPU profile, profiler, flame chart.
A picture of a CPU profile: a profiler samples the call stack many times a second, and each function becomes a bar as wide as its share of the samples. Callers sit below the functions they call. The widest bars at the top are where the CPU time really goes.
Take a profile: the wider a bar, the more CPU time that function used. Then cache the fonts and profile again.
- Profile
- 0 s of 10
- Samples
- 0
- CPU per request
- –
No profile yet. The PDF service is busy and nobody knows why.
Say it in a prompt
Profile the PDF service under load with async-profiler for 30 seconds and save a CPU flame graph. List the 3 widest frames with their share of samples, fix the top one, then profile again under the same load and show both graphs side by side with the CPU time per request before and after. Vague vs precise prompt
Vague prompt
the PDF export uses too much CPU, optimize it Typical resultGuesses at the cause: rewrites the template loop and adds a cache for the database query. The CPU use barely moves, because the real cost was loading fonts on every request.
Precise prompt
Take a 30-second CPU profile of the PDF service under load (async-profiler, flame graph output). Name the 3 widest frames with their share of samples, fix the widest one, and profile again under the same load to show the CPU per request before and after. Typical resultThe flame graph shows loadFonts at 55% of the samples. Caching the fonts cuts CPU per request from about 400 ms to about 185 ms, and the second graph proves it.
Seen on
- Brendan Gregg: The page by the inventor of flame graphs: the wider a frame, the more often it was in the sampled stacks, and the x-axis is sorted by name, not time.
- Grafana Pyroscope: Its docs explain reading a flame graph: each bar is a function, its width is the time spent in it, and the bars stack in call order.
You might describe it as
- find which function eats the CPU
- a chart of where the server spends its time
- wide bars show the expensive code
Not to be confused with
- Distributed trace
A trace follows one request across services over time; a flame graph adds up where CPU time goes inside one program.