Skip to content
Prompt Words
History

p99 latency

Also called: tail latency, 99th percentile, percentile latency.

The response time that 99% of requests beat: sort the times, and the one 1% from the slow end is p99. An average hides a few very slow requests, but p99 shows them. p50 (the median) is what a typical user sees.

Each bar is one request: the taller, the slower. Send 100 requests and compare the average with p99.

Requests
0
Average
–
p50
–
p99
–

Nothing sent yet. Most requests take about 40 ms; a few call a slow tax service.

    Idle

    Say it in a prompt

    Add p50, p95 and p99 latency for every endpoint, measured with an OpenTelemetry histogram and shown per endpoint on the dashboard, not as an average. Alert when p99 of GET /orders is above 300 ms for 5 minutes, and log the trace id of any request slower than 1 s so we can see where it waited.

    Vague vs precise prompt

    Vague prompt

    check if the API is fast enough

    Typical resultAdds an average response time to the dashboard. It says 55 ms, so everything looks fine, while 2 users in 100 wait about 0.8 s.

    Precise prompt

    Record a latency histogram per endpoint and show p50, p95 and p99 on the dashboard instead of the average. Alert when p99 of GET /orders is over 300 ms for 5 minutes, and log the trace id of requests slower than 1 s.

    Typical resultThe dashboard shows the slow 1% (p99 790 ms) next to the typical 40 ms, an alert fires when the tail gets worse, and the slow requests come with a trace id to look into.

    Seen on

    • Google SRE book: Says an average hides a long tail of much slower requests, and suggests a high percentile like the 99th as a plausible worst case.
    • Amazon CloudWatch: Explains percentiles (the 95th percentile is the value 95% of data points are below) and that watching the average can hide unusual values. Load balancer, API Gateway and Lambda metrics support them.

    You might describe it as

    • the average looks fine but some users wait for ages
    • how slow the slowest requests are
    • fast for most people, slow for a few

    Not to be confused with

    • Request timeout

      A timeout cuts one slow request off; p99 measures how slow the slowest 1% of requests are.

    • Distributed trace

      p99 tells you that some requests are slow; a trace shows where one request spent its time.