AI News Feed
Market watch
Large Language Models

JetBrains Releases Mellum2.1, a 12B Open MoE Model for Coding Agents

JetBrains has released Mellum2.1, a 12B open mixture-of-experts reasoning model with 2.5B active parameters for coding agents and self-hosting. It beats Mellum2 on 15 of 17 benchmarks but still trails Qwen3.5-9B on several hard agentic and knowledge tasks.

The architecture is unchanged from Mellum2. Mellum2.1 has 28 layers and 64 experts, with a router activating eight experts per token. Attention uses grouped-query attention with 32 query heads and four KV heads, and three of every four layers use a 1,024-token sliding window. Context length is 131,072 tokens and the vocabulary has 98,304 tokens. Weights ship in bfloat16. JetBrains targets three uses: agent worker, general reasoning assistant, and private self-hosted deployment.

JetBrains said almost all new work went into post-training. Reinforcement learning moved from a short final stage to the main part of training. The company added RL tasks in math, competitive programming, science, tool use, and software engineering, and filtered open RL datasets for broken tests, unverifiable answers, and tasks that were too easy or impossible. For software engineering, the model trains in real repositories with a shell and file-editing tools and is rewarded when tests pass. Training launched millions of sandboxes across thousands of environments.

JetBrains evaluated Mellum2.1, Mellum2, Qwen3.5-9B and Gemma 4 E4B with one pipeline in thinking mode. All scores are self-reported by JetBrains. The largest jump is agentic coding. SWE-bench Verified rose from 2.0 to 47.0. SWE-bench Pro rose from 0.0 to 28.0. Terminal-Bench 2.1 rose from 0.6 to 17.4. Agentic runs used the open-source Pi v0.73.1 harness with a 114K-token context. Mellum2.1 leads the group on LiveCodeBench v6 at 82.0, HumanEval+ at 91.5, MBPP+ at 79.4 and BFCL v4 at 62.3. Qwen3.5-9B still leads on SWE-bench Verified at 50.0, SWE-bench Pro at 38.0, AIME 25/26 at 86.7 and GPQA Diamond at 77.8. Safety also improved: HarmBench fell from 21.5 to 8.5, where lower is better.

JetBrains said pipelines matter. Qwen's own model card lists 65.6 on LiveCodeBench v6 and 81.7 on GPQA Diamond. JetBrains measured 75.4 and 77.8 for the same model. Mellum2.1 beats Mellum2 on 15 of 17 listed benchmarks and wins 5 of 17 against Qwen3.5-9B, according to the released comparison. Its best listed result is LiveCodeBench v6, ahead of Qwen3.5-9B at 75.4 and Gemma 4 E4B at 69.4; its weakest is Terminal-Bench 2.1 at 17.4, below Qwen3.5-9B at 21.7.

Post-training left the architecture untouched, so speed matches Mellum2. On one NVIDIA H200 under heavy load, JetBrains says Mellum2.1 serves almost twice the tokens of Qwen3.5-9B. For a single request, multi-token prediction makes it about 1.6 times faster. The MTP head for vLLM speculative decoding is listed as coming soon. The full model serves on vLLM with --reasoning-parser qwen3; tool calling adds --enable-auto-tool-choice --tool-call-parser hermes. JetBrains recommends temperature 0.6, top_p 0.95 and top_k 20.

GGUF builds are in progress. A GGUF repository already lists five files: BF16 at 24.3 GB as reference, Q8_0 at 12.9 GB with 96.1% top-token match, Q6_K at 10.9 GB with 93.9%, Q4_K_M at 8.1 GB with 88.0% and recommended, and MXFP4_MOE at 7.0 GB with 85.6%. The smaller builds are intended for llama.cpp, Ollama and LM Studio, with GGUF starting at 7.0 GB. The model also runs on vLLM or SGLang on your own GPUs.

In the comparison provided, Mellum2.1 and Mellum2 Thinking both use a 12B MoE architecture with 2.5B active parameters and a 131,072-token context. Qwen3.5-9B is a 9B dense language model with a hybrid Gated DeltaNet plus gated attention design, 262,144 native context and up to 1,010,000 tokens, and text, image and video modalities. Gemma 4 E4B is dense with per-layer embeddings, 8B parameters with embeddings, 4.5B effective parameters, 128K context, and text, image and audio modalities. All four are listed under Apache 2.0.