Faster Gemma 4 on MLX with multi-token prediction
What changed
Faster Gemma 4 on MLX with multi-token prediction June 29, 2026 Gemma 4 is now significantly faster in Ollama 0.31. On Apple Silicon, it generates tokens nearly 90% faster on average across a coding-agent benchmark. Gemma 4 ships with a small, fast draft model that runs alongside the main model and proposes the next several tokens.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Ollama is now powered by MLX on Apple Silicon in preview
- Improved performance and model support with GGUF
- Google releases Gemma 2 2B, ShieldGemma and Gemma Scope
Sources
- Faster Gemma 4 on MLX with multi-token prediction (ollama-blog)primary