AI News Feed
Market watch
Large Language Models

Hugging Face Releases Olmo-core 3 for Trillion-Parameter MoE Training

Hugging Face released Olmo-core 3 on Oct. 1, 2026, an open training framework for large mixture-of-experts models. The system is designed to scale MoE training to more than a trillion parameters while reducing communication and memory costs, the company said.

The release addresses a practical constraint in large-model development. Training large AI models requires substantial compute, raising costs and energy use and putting advanced model development out of reach for many academic researchers and smaller labs. MoE models offer a more efficient approach because they can contain many more learned components, or parameters, without requiring every input to use all of them. But the full model still has to be stored across GPU memory and updated during training. Routing inputs to the right experts across a cluster creates communication and coordination costs. As MoEs grow, those costs can erode much of the computational advantage of using only part of the model for each input.

Olmo-core 3 is built to close that gap, according to the blog post. In one benchmark, the expert pool increased from 8 to 128 while the system still selected only four experts per token, the small units of text a language model processes. The number of active parameters per token stayed roughly fixed at about 3.2 billion. Total parameter capacity grew from 4.6 billion to 47 billion, while training throughput fell by less than 5 percent. The same infrastructure has been benchmarked at over one trillion total parameters.

The framework has evolved with each generation of Olmo. Work on sparse models goes back to OlmoE, which used an MoE architecture with 64 routed experts. Olmo 3 used a dense architecture, meaning nearly all of the model was active for every token, and its training stack was built around that design. Olmo-core 3 extends the framework with a training system designed for much larger MoE models.

The earlier MoE implementation in Olmo-core used fully sharded data parallelism, configured to gather and reshard model weights for each small batch of training data. Olmo-core 3 switches to a system based on distributed data parallelism. It keeps experts resident on GPUs and routes the relevant data to them, avoiding repeated weight gathering.

NVIDIA's Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over the earlier FSDP-based implementation. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using the earlier implementation, about 2.7 times the throughput.

Olmo-core 3 combines several techniques for distributing large MoEs across GPU clusters with optimizations that make routing and computation more efficient. Three techniques determine how the model and its training state are split across hardware. Expert parallelism spreads the experts across GPUs, so each GPU stores only part of the full expert pool. Pipeline parallelism splits the model's layers, the successive stages that transform an input, across groups of GPUs, reducing how much of the model each GPU needs to keep in memory. A distributed optimizer spreads the optimizer state, the additional data used to calculate and apply updates during training, across GPUs instead of storing a full copy on every GPU. Together, these techniques allow an MoE to scale without requiring every GPU to keep the entire model and its training state in memory.

Olmo-core 3 also reduces the cost of routing data to the right experts and running their computations. Rowwise expert parallelism places routed data directly into expert input buffers, minimizing the extra work needed to rearrange it. GPU-resident routing keeps routing metadata on the GPUs, so the CPU can queue work without waiting for that information to be copied back. Grouped GEMM combines many small expert computations so GPUs can execute them more efficiently.

Finally, Olmo-core 3 supports MXFP8, a lower-precision number format that represents some values with fewer bits. This can reduce computation and the amount of data moved between GPUs, as long as those savings outweigh the cost of converting between number formats. The blog post said a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts, measured MXFP8's effect on end-to-end training throughput. With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21 percent higher than with BF16, the higher-precision format used as the baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.

Editor's Summary

Hugging Face released Olmo-core 3, an open MoE training stack designed to scale models into the trillion-parameter range while limiting throughput loss. The company reported a 2.7x throughput improvement over its earlier FSDP-based implementation in one eight-GPU test and about 21% higher throughput with MXFP8 in a four-GPU benchmark. The release extends the Olmo framework's focus on open training infrastructure for large sparse models.