Articles
Engineering, one diagram at a time
Short, free articles on AI and systems engineering. Each one explains a single diagram and links to the full lesson behind it.

Containers · Kubernetes
A container is a process, not a small VM
A container is an ordinary Linux process with namespaces and cgroups around it. That one fact explains the shared kernel, the PID 1 shutdown bug and the heap sized for the whole node.
· 4 min read

AWS · DynamoDB
A filtered DynamoDB Scan still reads the whole table
On a 40-million-item table, a Query evaluates 200 items to return 200 orders and a filtered Scan evaluates 40 million. Design DynamoDB keys around the questions you know you will ask.
· 5 min read

Observability
One metric label can cost 648 GB of memory
Metrics, logs and traces are three shapes of data, not three tools. Put user_id in a metric label and a 1,080-series counter becomes 216 million series. Where each question belongs.
· 5 min read

MCP · Developer Tools
Read an MCP server before you trust it
List the tools of the MCP reference server and you find a read-only tool that hands over your secrets and another that fetches any URL from your machine. How to pin, inspect and allow.
· 5 min read

AI Security · Governance
Security, safety, robustness and privacy are four different problems
One AI incident can touch all four, but each needs its own evidence and its own owner. A clinical summarizer shows why no single guardrail can prove them all.
· 4 min read

AI Agents · Agentic RAG
What an AI agent actually is
The same model becomes an agent when a loop lets it call tools, read the results and choose its next step, while your runtime keeps the limits. A worked example, and where autonomy should stop.
· 5 min read

AWS Architecture
What multi-AZ buys, and what it doesn't
Availability zones are a failure-domain definition. Multi-AZ survives a power or fibre loss, not a bad deploy, an expired certificate or a corrupt migration, and only with spare capacity.
· 4 min read

LLM Inference · Serving
Where a slow first token comes from
A three-second first token can come from the queue, a media fetch, the tokenizer, prefill, rank synchronization or a buffering proxy. A faster kernel fixes one of those. Measure every boundary first.
· 4 min read

Agentic AI · Context
Your context window is a budget
Every token in the window costs money, latency and the model's attention. Reserve room for the answer first, then spend what is left on whole bundles of evidence, never on half of one.
· 5 min read

Databases at Scale
Your database reads pages, not rows
Ask for one 120-byte row and the engine reads an 8 KiB page. The same index can find 200 rows in four page reads or in two hundred, depending on how the rows were stored.
· 5 min read

Sharding · System Design
A good hash can't fix a hot tenant
Hash partitioning on tenant_id spreads tenants evenly, but not load. When one merchant is 40% of writes, one shard sets your capacity. Three repairs, and what each one costs.
· 4 min read

Realtime Voice AI
A voice agent is a stack, not a model
A demo voice bot is speech-to-text, a language model and text-to-speech. A production voice agent adds media, state, tools, people and controls, and keeps media apart from business actions.
· 4 min read

Claude Code · Codex
Build the safety net before your coding agent's first edit
Coding agents will eventually do something you did not intend. Four layers limit the damage: permissions, a sandbox, checkpoints and Git. Each catches what the one above misses.
· 5 min read

AI Coding Agents
Five ways to code with AI, five different risks
Autocomplete, chat, interactive agents, asynchronous agents and agentic workflows differ in what they can touch and how they fail. Choose by reversibility and consequence, not by maturity.
· 4 min read

Multimodal AI
Multimodal AI starts with evidence, not modality
Sending every photo, recording and document to the largest model is the common multimodal mistake. Start from the decision, treat each signal as evidence with an origin, and let policy decide.
· 3 min read

ChatGPT · Claude · Developers
Pick the place before you pick the model
Most disappointing AI answers come from asking in the wrong place, not from the wrong model. Six kinds of developer work, and where each one belongs.
· 5 min read

Observability · Latency
Same average, very different users
Two services with an identical 153 ms mean latency: one where nobody waits more than 161 ms, one where a thousand requests take over a second. Why the mean cannot tell them apart.
· 4 min read

Distributed Systems · Payments
Why a retry charges the customer twice
A correct payment service, a correct client and one lost reply are enough to charge a customer twice. Why the client cannot fix it, and what an idempotency key changes.
· 4 min read

Kafka · Event-Driven Systems
Why Kafka is a log, not a queue
A queue deletes a message once it is delivered. A Kafka topic keeps it and remembers one position per consumer group. That one choice explains fan-out, replay, and the head-of-line problem.
· 4 min read

Databases at Scale
Your primary key is a storage decision
The same 572 MiB of rows can cost 1.1 GiB or 38.7 GiB of disk writes, depending on the order keys arrive in and how the storage engine compacts. A lab model of write amplification.
· 4 min read
Get one diagram a week
A short article built around one engineering diagram, from the same library as these courses.
One diagram-led article a week on AI and systems engineering. We email you once to confirm, and every newsletter has an unsubscribe link. Privacy policy