Ollama is now powered by MLX on Apple Silicon in preview
What changed
Ollama is now powered by MLX on Apple Silicon in preview March 30, 2026 Today, we’re previewing the fastest way to run Ollama on Apple silicon, powered by MLX, Apple’s machine learning framework. On Apple’s M5, M5 Pro and M5 Max chips, Ollama leverages the new GPU Neural Accelerators to accelerate both time to first token (TTFT) and generation speed (tokens per second). Get started This preview release of Ollama accelerates the new Qwen3.5-35B-A3B model, with sampling parameters tuned for coding tasks.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Faster Gemma 4 on MLX with multi-token prediction
- Improved performance and model support with GGUF
- Introducing AnyLanguageModel: One API for Local and Remote LLMs on Apple Platforms
Sources
- Ollama is now powered by MLX on Apple Silicon in preview (ollama-blog)primary