Skip to content
Prompt Words
History

Distributed trace

Also called: tracing, trace, OpenTelemetry trace.

A record of one request as it moves through several services. Each step is a span with a start time and a duration, and every span carries the same trace id. Shown as a waterfall, it tells you where the request spent its time.

Send one checkout request. Each service it passes adds a span to the same trace, and the waterfall shows where the time went.

Nothing traced yet. The request will pass 4 services.

    Idle

    Say it in a prompt

    Add OpenTelemetry tracing to the gateway, the Orders API and the email service. Pass the trace id between services in the W3C traceparent header, add a span for every database query and outgoing HTTP call, send spans to Jaeger, and put the trace id in every log line so a slow checkout can be found from its logs.

    Vague vs precise prompt

    Vague prompt

    checkout is slow sometimes, find out why

    Typical resultAdds console.log timings in one service. The logs of the gateway, the API and the email service can't be matched up, so nobody sees that checkout waits for the email.

    Precise prompt

    Add OpenTelemetry tracing to the gateway, Orders API and email service, pass the trace id in the traceparent header, add spans for database queries and HTTP calls, export to Jaeger, and log the trace id with every line.

    Typical resultOne slow checkout opens as a waterfall of 4 spans: the email call takes 340 of the 420 ms, so the fix is to send the email in the background.

    Seen on

    • OpenTelemetry: Explains that a trace is made of spans, each with a start and end time, a parent span, and the trace id they all share.
    • W3C Trace Context: The standard traceparent HTTP header that carries the trace id from one service to the next, for example 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01.

    You might describe it as

    • follow one request through all the services
    • see which service made the page slow
    • a timeline of every call one click made

    Not to be confused with

    • p99 latency

      p99 tells you that some requests are slow; a trace shows where one request spent its time.

    • Flame graph

      A trace follows one request across services over time; a flame graph adds up where CPU time goes inside one program.