Chapter 25
25Reasoning and Test-Time Compute
25.1Spending compute at inference
Planned: The shift from bigger models to longer thinking; the o1/R1 turn.
25.2RL on verifiable rewards
Planned: Training reasoning with automatically checkable answers (math, code).
25.3Search, self-consistency, and verification
Planned: Sampling many chains and selecting; process vs outcome reward.
25.4What this changes
Planned: New scaling axis, new costs, and the blurring of train/inference.