Multimodal open d1 decision models for the edge
What changed
- Best decision model under 10B on the Decision Index 0.2.1: d1-3B scores 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11). Unlike our generative models, decision models don’t produce tokens but answer in a single forward pass. It adds vision and audio encoders to handle all three modalities.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Ollama now supports Jev-style decision models
- Ollama's new engine for multimodal models
- Fine-tuning Florence-2 - Microsoft's Cutting-edge Vision Language Models
Sources
- Multimodal open d1 decision models for the edge (huggingface-blog)primary