Qwen 3.8 27B available on Cerebras at 1500 tokens/s
What changed
Available Models Model Compression This section provides transparency about the compression state of each model available on our platform. We host a variety of open-source models from the community. We do not currently host pruned models on our public endpoints.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Running a 28.9M parameter LLM on an $8 microcontroller
- The efficient frontier of LLM inference
- AirLLM 70B inference with single 4GB GPU
Sources
- Qwen 3.8 27B available on Cerebras at 1500 tokens/s (hn-frontpage)primary