Skip to content
Handbook navigation

Full-access lesson

Recurrent Neural Networks

> A gradient flowing back 5 steps through a plain RNN shrinks to 0.0000233. At 20 steps, it has shrunk to 0.000000000000000000021 -- eighteen orders of magnitude smaller, using nothing but repeated multiplication by numbers already less than 1.

The complete curriculum

All 12 volumes, companion resources, interviews, and architecture reviews.

Executable engineering practice

Subscriber-only Python, Java, TypeScript, and SQL labs in the isolated runner.

Recurrent Neural Networks | KnowledgeOS