Deliberative alignment: reasoning enables safer language models

Practical AI: Tools, Models & Frameworks

What changed

Deliberative alignment: reasoning enables safer language models Introducing our new alignment strategy for o1 models, which are directly taught safety specifications and how to reason over them.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Introducing OpenAI o1
  • Introducing HELMET: Holistically Evaluating Long-context Language Models
  • Reasoning models struggle to control their chains of thought, and that’s good

Sources