Deliberative alignment: reasoning enables safer language models
What changed
Deliberative alignment: reasoning enables safer language models Introducing our new alignment strategy for o1 models, which are directly taught safety specifications and how to reason over them.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Introducing OpenAI o1
- Introducing HELMET: Holistically Evaluating Long-context Language Models
- Reasoning models struggle to control their chains of thought, and that’s good
Sources
- Deliberative alignment: reasoning enables safer language models (openai-blog)primary