Multimodal open d1 decision models for the edge

Practical AI: Tools, Models & Frameworksmultimodal

What changed

  • Best decision model under 10B on the Decision Index 0.2.1: d1-3B scores 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11). Unlike our generative models, decision models don’t produce tokens but answer in a single forward pass. It adds vision and audio encoders to handle all three modalities.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Ollama now supports Jev-style decision models
  • Ollama's new engine for multimodal models
  • Fine-tuning Florence-2 - Microsoft's Cutting-edge Vision Language Models

Sources