BenchMIRT: What are LLM benchmarks actually measuring?

Practical AI: Tools, Models & Frameworksllm

What changed

Today we’re introducing BenchMIRT, a new method for auditing LLM benchmarks at the level of individual prompts—the questions and tasks a model is scored on. A benchmark is usually designed to measure a particular ability, such as safety, general reasoning, or instruction following. The models we used to train and evaluate BenchMIRT were all released by March 2025, so our analysis doesn’t capture how BenchMIRT behaves on newer generations of LLMs.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Measuring benchmark optimization in speech recognition
  • Introducing Real World VoiceEQ: Measuring the human quality of voice AI
  • Introducing the Open Ko-LLM Leaderboard: Leading the Korean LLM Evaluation Ecosystem

Sources