Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers
What changed
As enterprise generative AI applications move to production, platform engineers face a key challenge: balancing the flexibility of LLM-as-a-judge guardrails with the reliability and portability of traditional classifiers that require custom training data. The recent emergence of "decision models"—highlighted by TypeSafe AI's recent announcement of Jev and "System One" models—promises a flexible middle ground by producing fixed "decisions" given a state and a list of questions rather than generating text. A trivial example of using a decision model (adapted from John Berryman of Arcturus Lab's blog post) might look like the following.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Clef: Open-weight decision models, and new RL fine-tuning platform
- Web Search API
- Running a 28.9M parameter LLM on an $8 microcontroller
Sources
- Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers (hn-frontpage)primary