Skip to content
All articles

AWS · DynamoDB

A filtered DynamoDB Scan still reads the whole table

On a 40-million-item table, a Query evaluates 200 items to return 200 orders and a filtered Scan evaluates 40 million. Design DynamoDB keys around the questions you know you will ask.

· 5 min read

40 million orders of about 1 KB each, and a question that returns 200. A Query on the partition key evaluates 200 items, uses 24.5 read units and returns one page. A Scan with a filter expression evaluates 40,000,000 items, uses about 4.9 million read units and reads at least 38,147 pages, about 13 minutes: about 199,000 times the read capacity for the same 200 results. Design the key for the question: PK = CUSTOMER#4471, SK = ORDER#2026-09-09#8871 makes all orders for a customer, newest first, one Query.

DynamoDB is a distributed hash table with a sorted range, and nearly everything about designing for it follows from that. The partition key is hashed to choose a partition, so equality is the only operation it supports. The sort key orders items within a partition and supports ranges, prefixes and ordered reads. That is why so much DynamoDB modelling comes down to encoding hierarchy into the sort key, as in ORDER#2026-09-09#8871.

The practical consequence: know your access patterns before you design the table. A Scan does not turn an unindexed access pattern into a cheap lookup.

What a filtered Scan actually reads

Take a table of 40 million order items at about 1 KB each, 40 GB in all, and a question whose answer is 200 orders.

Query on the partition keyScan with a filter expression
Items evaluated20040,000,000
Read units24.5About 4.9 million
PagesOneAt least 38,147 of up to 1 MiB: about 13 minutes at 20 ms a page

The Scan spends roughly 199,000 times the read capacity to return the same 200 items. A filter is applied after items are read, so it shrinks what comes back without shrinking the capacity spent evaluating them. That capacity also competes with your operational reads, and a throttled read can stall a flow that has nothing to do with the report.

Design the key for the question

The questionKey design
Has this idempotency key been used?Partition key = the key itself, claimed atomically with a conditional write
All orders for a customer, newest firstPK = CUSTOMER#4471, SK = ORDER#2026-09-09#8871, queried in descending order
Orders in a date range, across all customersA global secondary index keyed on a coarse bucket, such as the month
Failed orders above a valueAn index with status as the partition key and the value as a numeric sort key

Every one of these has a cost. A heavily used customer becomes a hot key. Month buckets make the current month a hot partition, while day buckets need 30 queries to cover a month. Each index adds write and storage cost, and its results are eventually consistent.

Engineering Insight

Choose DynamoDB when the access patterns are known, stable and key-based: sessions, idempotency keys, event stores, user profiles by id. Choose a relational database when someone will eventually ask a question you did not anticipate, which describes most product data. "We might need to scale" is not a reason. "Every read is by this key, forever" is.

A nightly Scan-and-filter report can still be acceptable if its load is bounded and the table has read headroom to spare. Measure consumed read capacity and read throttling before deciding, or move the reporting to an export or a suitable index.

Get one diagram a week

A short article built around one engineering diagram, from the same library as these courses.

One diagram-led article a week on AI and systems engineering. We email you once to confirm, and every newsletter has an unsubscribe link. Privacy policy