Foundations of
Large Language ModelsTraining, serving, and shipping LLMs — everything for a 2026 engineering interview.
  • •Preface
I · Foundations
  • 1What Is a Large Language Model?
  • 2Deep Learning, Just Enough
  • 3Tokenization
  • 4The Transformer
  • 5How Modern Architectures Differ
II · Pretraining
  • 6The Pretraining Objective and the Data
  • 7Training Dynamics at Scale
  • 8Distributed Training
  • 9Scaling Laws
III · Post-training and Alignment
  • 10Supervised Fine-Tuning
  • 11RLHF and Reward Modeling
  • 12Preference Optimization Without RL
  • 13Parameter-Efficient Fine-Tuning
IV · Inference and Serving
  • 14Decoding and Sampling
  • 15Making Inference Fast
  • 16Quantization
  • 17Serving Systems in Production
V · The Harness: Making Models Behave
  • 18Prompting and System Prompts
  • 19Tool Use and Function Calling
  • 20Structured Output
  • 21Retrieval-Augmented Generation
  • 22Agents
  • 23Safety, Guardrails, and Moderation
VI · Evaluation
  • 24Evaluating Language Models
VII · The Frontier
  • 25Reasoning and Test-Time Compute
  • 26Open Problems (Conclusion)
Appendices
  • ARunning LLMs Locally
  • BMath and Notation Reference
  • CGlossary
The book cover: a row of tokens with attention arcs converging on the empty next position.

Foundations of Large Language Models

Training, serving, and shipping LLMs — everything for a 2026 engineering interview.

•Preface

I · Foundations

1What Is a Large Language Model? 2Deep Learning, Just Enough 3Tokenization 4The Transformer 5How Modern Architectures Differ

II · Pretraining

6The Pretraining Objective and the Data 7Training Dynamics at Scale 8Distributed Training 9Scaling Laws

III · Post-training and Alignment

10Supervised Fine-Tuning 11RLHF and Reward Modeling 12Preference Optimization Without RL 13Parameter-Efficient Fine-Tuning

IV · Inference and Serving

14Decoding and Sampling 15Making Inference Fast 16Quantization 17Serving Systems in Production

V · The Harness: Making Models Behave

18Prompting and System Prompts 19Tool Use and Function Calling 20Structured Output 21Retrieval-Augmented Generation 22Agents 23Safety, Guardrails, and Moderation

VI · Evaluation

24Evaluating Language Models

VII · The Frontier

25Reasoning and Test-Time Compute 26Open Problems (Conclusion)

Appendices

ARunning LLMs Locally BMath and Notation Reference CGlossary