How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
What changed
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Measuring progress toward AGI: A cognitive framework
- Introducing GPT-5.1 for developers
- How GPT-5.6 fuses frontier intelligence with frontier efficiency
Sources
- How enabling two settings tripled our scores on the ARC-AGI-3 benchmark (openai-blog)primary