Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains
What changed
- Mellum2 is a 12B-parameter Mixture-of-Experts model trained from scratch on natural language and code. - The model activates only 2.5B parameters per token, making it efficient for high-throughput, low-latency inference. - It is released under the Apache 2.0 license.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Introducing Gemma 4 12B: a unified, encoder-free multimodal model
- Embedding AI into developer software
- Introducing the Model Spec
Sources
- Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains (huggingface-blog)primary