Improved performance and model support with GGUF

Practical AI: Tools, Models & Frameworks

What changed

Improved performance and model support with GGUF June 5, 2026 Ollama 0.30 is now available with improved performance and GGUF model compatibility through llama.cpp. This augments Ollama’s MLX engine on Apple silicon, bringing support to more models on a wider range of hardware. Performance across more GPUs Faster throughput on NVIDIA hardware With Ollama 0.30, performance on NVIDIA hardware is now up to 20% faster, leveraging optimizations contributed by the NVIDIA and llama.cpp teams.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Ollama is now powered by MLX on Apple Silicon in preview
  • Ollama's new engine for multimodal models
  • Llama 3.2 goes small and multimodal

Sources