Chapter 23

23Safety, Guardrails, and Moderation

Draft pending — this page is an outline of planned content. Write it by creating content/safety-guardrails.md.

23.1Layers of defense

Planned: Alignment training, system prompts, input/output classifiers, and monitoring.

23.2Jailbreaks and red-teaming

Planned: The adversarial cat-and-mouse; common attack shapes.

23.3Content moderation classifiers

Planned: Separate models guarding the main one; precision/recall tradeoffs.

23.4Governance and the harness's responsibility

Planned: Where provider policy lives; refusal behavior and its costs.