Measuring the performance of our models on real-world tasks
What changed
OpenAI introduces GDPval, a new evaluation that measures model performance on real-world economically valuable tasks across 44 occupations.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Introducing Real World VoiceEQ: Measuring the human quality of voice AI
- Introducing GPT-5 for developers
- Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
Sources
- Measuring the performance of our models on real-world tasks (openai-blog)primary