Improved performance and model support with GGUF
What changed
Improved performance and model support with GGUF June 5, 2026 Ollama 0.30 is now available with improved performance and GGUF model compatibility through llama.cpp. This augments Ollama’s MLX engine on Apple silicon, bringing support to more models on a wider range of hardware. Performance across more GPUs Faster throughput on NVIDIA hardware With Ollama 0.30, performance on NVIDIA hardware is now up to 20% faster, leveraging optimizations contributed by the NVIDIA and llama.cpp teams.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Ollama is now powered by MLX on Apple Silicon in preview
- Ollama's new engine for multimodal models
- Llama 3.2 goes small and multimodal
Sources
- Improved performance and model support with GGUF (ollama-blog)primary