Skip to content
All articles

LLM Inference · Serving

Where a slow first token comes from

A three-second first token can come from the queue, a media fetch, the tokenizer, prefill, rank synchronization or a buffering proxy. A faster kernel fixes one of those. Measure every boundary first.

· 4 min read

The stages of one LLM request on a single request clock: gateway (auth, quota, parameters), route, queue, preparation (media fetch, tokenizer), prefill (prompt forward pass, KV cache), first decode (logits, sample one token) and transport (stream to the user). Causes of a slow first token sit beside them: a two-second queue, a slow fetch or overloaded tokenizer, a long prefill or rank synchronization, and a stream buffered at the proxy. A 10% faster matrix multiply would do nothing for five of those six causes.

A model API looks like one call. Behind it is a pipeline of stages, some deterministic and some not. A gateway authenticates the caller, enforces quotas, validates the generation parameters and chooses a route. A tokenizer turns the text into token IDs. The engine reserves memory, batches the request with compatible work, runs a forward pass over the prompt and stores the attention keys and values. Then it computes new tokens a step at a time, a sampler applies the decoding rules, a detokenizer rebuilds the text, and the transport streams events until a stop rule, a cancellation, an error or a token limit ends the request.

The model server is only one failure domain in that chain.

One symptom, six causes

Suppose users report a three-second first token. The same symptom can come from:

  • a two-second queue before any kernel launches,
  • a slow fetch of the request's media,
  • an overloaded tokenizer, which can also block an event loop,
  • a long prefill,
  • synchronization between distributed ranks,
  • a stream buffered at a proxy.

Making matrix multiplication ten percent faster would do nothing for five of them.

Decompose time to first token

Give each boundary its own timestamp, all on one monotonic request clock, and split the time to first token into its parts:

TTFT = gateway + route + queue + preparation + prefill + first decode + transport

The parts should add up closely to what the user saw. When they do not, there is a boundary you have not instrumented yet. Making that possible means production traces carry identifiers for the request, the route, the model artifact, the engine configuration, the queue, prefill, decode, tokens, cancellation and the response.

Engineering Insight

Record why each request ended as carefully as when: natural stop, length limit, policy block, client cancellation, deadline, worker failure or upstream disconnect. That is what separates a model that habitually runs out of output budget from a network that abandons healthy work, and it exposes the disconnected client whose expensive decode kept running after nobody was listening.

Senior performance work begins with this discipline: find where the time and the state are before changing capacity or code.

Get one diagram a week

A short article built around one engineering diagram, from the same library as these courses.

One diagram-led article a week on AI and systems engineering. We email you once to confirm, and every newsletter has an unsubscribe link. Privacy policy