A model API looks like one call. Behind it is a pipeline of stages, some deterministic and some not. A gateway authenticates the caller, enforces quotas, validates the generation parameters and chooses a route. A tokenizer turns the text into token IDs. The engine reserves memory, batches the request with compatible work, runs a forward pass over the prompt and stores the attention keys and values. Then it computes new tokens a step at a time, a sampler applies the decoding rules, a detokenizer rebuilds the text, and the transport streams events until a stop rule, a cancellation, an error or a token limit ends the request.
The model server is only one failure domain in that chain.
One symptom, six causes
Suppose users report a three-second first token. The same symptom can come from:
- a two-second queue before any kernel launches,
- a slow fetch of the request's media,
- an overloaded tokenizer, which can also block an event loop,
- a long prefill,
- synchronization between distributed ranks,
- a stream buffered at a proxy.
Making matrix multiplication ten percent faster would do nothing for five of them.
Decompose time to first token
Give each boundary its own timestamp, all on one monotonic request clock, and split the time to first token into its parts:
TTFT = gateway + route + queue + preparation + prefill + first decode + transport
The parts should add up closely to what the user saw. When they do not, there is a boundary you have not instrumented yet. Making that possible means production traces carry identifiers for the request, the route, the model artifact, the engine configuration, the queue, prefill, decode, tokens, cancellation and the response.
Engineering Insight
Record why each request ended as carefully as when: natural stop, length limit, policy block, client cancellation, deadline, worker failure or upstream disconnect. That is what separates a model that habitually runs out of output budget from a network that abandons healthy work, and it exposes the disconnected client whose expensive decode kept running after nobody was listening.
Senior performance work begins with this discipline: find where the time and the state are before changing capacity or code.


