Skip to content

ML Systems Track

Distributed AI Training and ML Systems Engineering

Scale a verified training loop through DDP, FSDP, ZeRO, model parallelism, high-throughput data, atomic checkpoints, GPU scheduling, observability, recovery, and outcome-based economics.

ML Engineer to Distributed AI Systems Architect
About 155 min · 3 labs
52 authored sections

Open before purchase

A useful preview, not a teaser

Read these complete sections to judge the technical depth, teaching style, and architecture standard before buying.

Preview 1

Course Contract: Preserve the Optimization While Scaling the System

More GPUs create more ways to run quickly, expensively, and incorrectly.

Preview 2

Trace the Training Loop Before Distributing It

A framework call hides the state that must remain consistent across ranks.

Preview 3

Separate Training Data, Execution, and Control Planes

The training script is one participant in a larger artifact-producing system.

Start the free preview

Complete curriculum

52 sections · 5 parts + training systems studio

Course orientation

  1. Course ContractFree

Part 1 · Understand the Work

  1. Training LoopFree
  2. Units and ArithmeticCourse Pass
  3. Accelerator HardwareCourse Pass
  4. TopologyCourse Pass
  5. CollectivesCourse Pass
  6. Reference ArchitectureFree
  7. Single-GPU BaselineCourse Pass
  8. Data ParallelismCourse Pass
  9. Distributed SamplingCourse Pass
  10. Scaling EfficiencyCourse Pass

Part 2 · Fit and Parallelize

  1. Memory AccountingCourse Pass
  2. Mixed PrecisionCourse Pass
  3. Activation CheckpointingCourse Pass
  4. ZeRO and FSDPCourse Pass
  5. Tensor ParallelismCourse Pass
  6. Pipeline ParallelismCourse Pass
  7. Context ParallelismCourse Pass
  8. Expert ParallelismCourse Pass
  9. Compose ParallelismCourse Pass

Part 3 · Feed and Optimize

  1. Training CorpusCourse Pass
  2. Tokenization and PackingCourse Pass
  3. Data LoadersCourse Pass
  4. DeterminismCourse Pass
  5. Optimizer StateCourse Pass
  6. Global BatchCourse Pass
  7. Kernels and CompilationCourse Pass
  8. ProfilingCourse Pass
  9. Communication OverlapCourse Pass

Part 4 · Operate the Cluster

  1. CheckpointingCourse Pass
  2. Elastic RecoveryCourse Pass
  3. SchedulingCourse Pass
  4. Kubernetes GPUsCourse Pass
  5. Quotas and PriorityCourse Pass
  6. ObservabilityCourse Pass
  7. StragglersCourse Pass
  8. NetworkCourse Pass
  9. StorageCourse Pass
  10. Reliability and ReleaseCourse Pass
  11. Model ValidationCourse Pass
  12. SecurityCourse Pass
  13. Cost and EnergyCourse Pass
  14. Incident ReviewCourse Pass
  15. Worked CaseCourse Pass

Part 5 · Architect and Prove

  1. CapstoneCourse Pass

Training Systems Studio

  1. Capacity LabCourse Pass
  2. Recovery LabCourse Pass
  3. Release Gate LabCourse Pass
  4. AssessmentCourse Pass
  5. Interview QuestionsCourse Pass
  6. Cheat SheetCourse Pass
  7. ReferencesCourse Pass