Ai2 Releases Olmo-core 3 to Improve Trillion-Parameter MoE LLM Training
Ai2 released Olmo-core 3, a framework for training trillion-parameter mixture-of-experts LLMs more efficiently.
MoE models differ from dense AI models in how they use computation. They split computation across specialist portions of the model each time a token is generated, while dense models activate the entire model. A MoE model may contain far more total parameters, the individual dials that tune model behavior, but it computes with only a small number of them as it generates each part of an answer. Training a MoE model likewise allows only portions of it to be activated while learning each token, a small piece of text that an AI reads and generates. That lowers overall compute, but the full model still must be stored across GPU memory, and coordinating networking between experts during training adds costs.
Ai2 said Olmo-core 3 was built to bridge that gap. The framework allows the expert pool to grow from eight to 128 while still selecting only four experts per token. Using the same infrastructure, Ai2 said, large language models can scale to more than one trillion parameters. According to SiliconANGLE, the release is part of Ai2's effort to give researchers tools to build and train larger models, which are often beyond the reach of those without access to state and enterprise infrastructure.
In benchmarks, Ai2 said Olmo-core 3 processed 52,000 tokens per second on Nvidia B3000 GPUs for a 47-billion-parameter model. Nvidia's Megatron-core training architecture, an established option for training large MoEs, topped out around 19,400 tokens per second in the comparison. Ai2 said that represented a jump of about 2.7 times the throughput.
Ai2 described several architectural choices in a whitepaper on the project. The new architecture uses expert parallelism to spread experts across multiple GPUs, allowing each card to store only part of the full expert pool. It also splits the model's layers, the successive stages that generate inputs, across groups of GPUs to reduce how much of the model each GPU needs to keep in memory. A distributed optimizer spreads the optimizer state, additional data used to calculate and apply updates during training, across multiple GPUs instead of storing full copies on every GPU. Ai2 said that reduces memory overhead as models scale because the entire model and its training state do not need to be stored in memory all at once.
The framework also supports MXFP8, a number format for LLMs that represents some values with fewer bits. According to Ai2, it can reduce computation and the amount of data moved between GPUs. The company added that Olmo-core 3 would allow researchers to adapt MoE training to different hardware and experiment with routing, parallelism and other parts of the system to build out a new ecosystem.
The project and related systems are currently available for developers and the open-source community on GitHub, according to SiliconANGLE.
Editor's Summary
Ai2's Olmo-core 3 is an open-source framework designed to make mixture-of-experts LLM training more efficient, with reported throughput roughly 2.7 times that of Nvidia's Megatron-core on a 47-billion-parameter model. The framework uses expert parallelism, layer splitting, a distributed optimizer and MXFP8 support to reduce memory and compute costs while scaling to more than one trillion parameters. It is available to developers and the open-source community on GitHub.