LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation

Practical AI: Tools, Models & Frameworksquantization

What changed

The BF16 GGUF serves as the in-format ceiling. We also add one scale-appropriate math evaluation: GSM8K for LFM2.5-230M and LFM2.5-350M, and AIME25 for LFM2.5-1.2B-Instruct and LFM2.5-2.6B. Across all four models, QAD substantially improves the Q4_0 checkpoint.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • LFM2.5-Encoders for Fast Long-Context Inference on CPU
  • Exploring Quantization Backends in Diffusers
  • Model Distillation in the API

Sources