Skip to content
Handbook navigation

Full-access lesson

Model Serving

> Increasing batch size from 4 to 16 improves throughput by 33% — and makes average latency 20% worse. Throughput and latency are not the same goal, and this chapter measures exactly where they pull apart.

The complete curriculum

All 12 volumes, companion resources, interviews, and architecture reviews.

Executable engineering practice

Subscriber-only Python, Java, TypeScript, and SQL labs in the isolated runner.

Model Serving | KnowledgeOS