Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

Practical AI: Tools, Models & Frameworksinference

What changed

Run open models as fast as your hardware allows Magnitude is an open source inference engine for agents that optimizes itself for your exact hardware. It compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. One click connects the agent you already use (Pi, OpenCode, Hermes, Codex, and more).

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
  • Improved performance and model support with GGUF
  • OpenJarvis: a local-first personal AI is now available to run with Ollama

Sources