Skip to content
All articles

Observability

One metric label can cost 648 GB of memory

Metrics, logs and traces are three shapes of data, not three tools. Put user_id in a metric label and a 1,080-series counter becomes 216 million series. Where each question belongs.

· 5 min read

Bar chart on a log scale for a request counter on the order service. Service, endpoint, status and region labels give 1,080 series and 3.2 MB of Prometheus memory. As an 8-bucket histogram: 10,800 series and 32 MB. With user_id added for 200,000 users: 216,000,000 series and 648 GB. As a log field or span attribute, user_id is free because it is indexed per event; as a metric label it is a series for every user, in every combination. A raw URL is the same mistake: label /orders/:id, not /orders/8871.

Metrics, logs and traces are usually presented as three products to buy. They are more usefully three shapes of data, and the shape decides where a given piece of information belongs.

  • Metrics are numbers aggregated over time, indexed by a small set of label combinations. They are cheap and compress well, and their cost is a product of label cardinality. They answer "how much, how many, how fast" for the whole system, and nothing about one particular request.
  • Logs are discrete events with arbitrary fields, indexed per event. They are expensive per unit of information, but they can carry anything, including high-cardinality identifiers.
  • Traces are causally linked spans across services, indexed by a trace id. They are the only signal that shows a request's path through a distributed system, and the only one that can attribute latency to a hop.

Cardinality, multiplied out

Take a request counter on the order service with four sensible labels: 6 services, 12 endpoints, 5 status codes and 3 regions. At about 3 KB of Prometheus memory per active series:

LabelsSeriesMemory
service × endpoint × status × region1,0803.2 MB
The same counter as a histogram with 8 buckets, plus sum and count10,80032 MB
The counter with user_id added, for 200,000 users216,000,000648 GB

The last row is not an exaggeration. It is one label, added by someone who wanted to know which customers were affected, which is an entirely reasonable thing to want. The same field in a log line or a span attribute costs nothing structural, because those are indexed per event rather than per label combination. The question was fine; the shape was wrong.

The middle row catches people who already know that rule. A histogram multiplies the series count by its number of buckets plus two, so the eight-bucket latency histogram an SLO needs carries ten times the cardinality of the counter beside it. That is still cheap at 10,800, but the multipliers compound: put user_id on the histogram instead of the counter and the last row becomes 2.16 billion series.

Which shape answers which question

You want to knowShape
How many orders failed in the last hourMetric
Which customer's orders failedLog field or span attribute
Why this one order took 2 secondsTrace

Common Mistake

Keeping user_id out of the labels but adding the full URL, query string and all. A raw path with ids and parameters in it is just as unbounded: /orders/8871?ref=email becomes its own series for good. Normalise to the route template, /orders/:id, in the instrumentation layer rather than trusting every caller to be tidy. The tell is a series count that keeps growing with traffic instead of levelling off.

Get one diagram a week

A short article built around one engineering diagram, from the same library as these courses.

One diagram-led article a week on AI and systems engineering. We email you once to confirm, and every newsletter has an unsubscribe link. Privacy policy