Mellum2.1 benchmarks
Mellum2.1 benchmarks Open source material ↗

Evidence at a glance

12B parameters,2.5B tokenEvidence
64 , token 8Evidence
2.0 47.0SWE-bench Verified
0.0 28.0SWE-bench Pro
131,072 tokenEvidence
28 ; 98,304Evidence

The mechanism in one line

InputReduce the input to a workable scale

Compress the visual or contextual input before the main reasoning path.

MechanismSpend compute where it matters

Route or verify the expensive step instead of repeating the full path.

OutcomeEnd with a measurable workflow result

Translate the mechanism into a bounded deployment or evaluation check.

The Architecture Stayed Put; the Task Changed

JetBrains released Mellum2.1 on October 8, 2026, as an open model for coding agents and fast sub-agents, under the Apache 2.0 license. It has 12 billion parameters in total but activates about 2.5 billion per token. JetBrains positions it as a model that can explore a repository, edit files, and check its changes—not merely answer isolated programming questions.

The update is notable not because it introduces a larger model or a new architecture. Mellum2.1 keeps Mellum2’s design, while most of the work went into post-training: JetBrains made reinforcement learning a central phase rather than a short final stage, and reports that SWE-bench Verified rose from 2.0 to 47.0. For technical leaders, the practical question is whether a model can move beyond producing plausible code and complete repository tasks under tool and test constraints.

Passing Tests Became a Training Signal

JetBrains describes training the model in real software repositories, where it can use shell and file-editing tools and receive rewards based on task outcomes. For software engineering tasks, whether tests pass is one key signal. The training mix also includes mathematics, competitive programming, science, and tool use. JetBrains says it filtered open reinforcement-learning datasets to remove broken tests, unverifiable answers, and tasks that were either too easy or impossible.

This changes the route by which the model learns what to do. Code examples can teach it to imitate local patterns, but may not teach it to find a cross-file issue, apply a change, and respond to a failed result; in an executable environment, test outcomes provide at least some checkable feedback on those actions. JetBrains says it launched millions of sandbox tasks across thousands of environments, but has not published the task distribution, dataset proportions, or training compute. It is therefore not possible to tell which tasks contributed most.

A Large Jump Is Not the Same as a Broad Lead

The largest reported gains are on agentic software tasks: SWE-bench Verified rises from 2.0 to 47.0, SWE-bench Pro from 0.0 to 28.0, and Terminal-Bench 2.1 from 0.6 to 17.4. Mellum2.1 also scores 82.0 on LiveCodeBench v6, 91.5 on HumanEval+, and 79.4 on MBPP+, suggesting strength across several code-generation and programming evaluations.

But the comparison is not a clean sweep. In JetBrains’ shared evaluation pipeline, Qwen3.5-9B scores 50.0 on SWE-bench Verified and 38.0 on SWE-bench Pro, and its 21.7 on Terminal-Bench 2.1 is also higher than Mellum2.1’s 17.4. Qwen leads on AIME 25/26 and GPQA Diamond as well. On LiveCodeBench, by contrast, Mellum2.1 scores 82.0 against Qwen’s 75.4. The results need to be read by task, rather than compressed into a single verdict about which model is stronger.

Sparse Activation Helps Deployment, but Speed Claims Need Checking

Mellum2.1 uses a mixture-of-experts design: a router activates eight of 64 experts for each token. The model has 28 layers, supports a 131,072-token context, and ships with bfloat16 weights. Its 2.5 billion active parameters describe how much participates in computation per token; they do not mean a deployment only needs to store that portion of the weights. For self-hosting teams, total model size, inference framework, and quantized builds still affect hardware and serving costs.

JetBrains says that under heavy load on one NVIDIA H200, Mellum2.1 delivers nearly twice Qwen3.5-9B’s throughput; for a single request, multi-token prediction (MTP) makes it about 1.6 times faster. The latter claim deserves particular caution: JetBrains also says the MTP head for vLLM speculative decoding is coming soon, and the public materials do not yet provide complete conditions for independent reproduction. GGUF builds are also available for llama.cpp, Ollama, and LM Studio, starting at 7.0 GB. File size alone, however, cannot substitute for latency, throughput, and quality tests on the target hardware.

Treat It as a Specialist, Not a Leader by One Score

For teams building coding agents, Mellum2.1 merits a place on the shortlist—especially for self-hosted deployments, cost-sensitive inference, or a supporting role alongside a primary agent, such as code search and localized edits. Its appeal lies in combining sparse activation with post-training aimed at repository work, not in beating larger competitors across every knowledge, reasoning, and difficult agent task.

The evaluation has limits. The model card labels the scores as JetBrains’ own results from a shared pipeline, and agent tasks used Pi v0.73.1 with a 114K-token context. Different pipelines can change comparisons, and some Qwen scores measured by JetBrains differ from those on Qwen’s own model card. Public materials do not include per-task logs, statistical uncertainty, or independent replications. Teams should evaluate completion rates, recovery from failure, and serving cost on their own repositories, tools, and tests before deciding what work to assign. A score of 47.0 is not, by itself, proof of production reliability.