LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation
What changed
The BF16 GGUF serves as the in-format ceiling. We also add one scale-appropriate math evaluation: GSM8K for LFM2.5-230M and LFM2.5-350M, and AIME25 for LFM2.5-1.2B-Instruct and LFM2.5-2.6B. Across all four models, QAD substantially improves the Q4_0 checkpoint.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- LFM2.5-Encoders for Fast Long-Context Inference on CPU
- Exploring Quantization Backends in Diffusers
- Model Distillation in the API
Sources
- LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation (huggingface-blog)primary