Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
What changed
Today we’re releasing Olmo-core 3, a significant upgrade to our framework for developing large language models featuring a redesigned open mixture-of-experts (MoE) training system. Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency. It’s one of the core systems behind the next generation of Olmo, and part of our ongoing commitment to open up the tools and training infrastructure behind each new model.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Introducing Training Cluster as a Service - a new collaboration with NVIDIA
- Stable Diffusion XL on Mac with Advanced Core ML Quantization
- Unlocking large scale AI training networks with MRC (Multipath Reliable Connection)
Sources
- Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs (huggingface-blog)primary