Chapter 23
23Safety, Guardrails, and Moderation
23.1Layers of defense
Planned: Alignment training, system prompts, input/output classifiers, and monitoring.
23.2Jailbreaks and red-teaming
Planned: The adversarial cat-and-mouse; common attack shapes.
23.3Content moderation classifiers
Planned: Separate models guarding the main one; precision/recall tradeoffs.
23.4Governance and the harness's responsibility
Planned: Where provider policy lives; refusal behavior and its costs.