Evidence at a glance

IOI 2026 535.4/600,Evidence
IOI 2026 361.12Evidence
IOI 2026 498.27Evidence
IOI 2025 Nano 130Evidence
Nano SFT 280Evidence
Nano RL 291Evidence

One Model Family, Not One Contest Model

NVIDIA’s Nemotron results, published on Hugging Face, describe two specialist systems built from Nemotron 3 for the International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO). IOI requires algorithm design and code that passes hidden tests, while IMO requires rigorous proofs in natural language. The failure modes differ: a program can have the right idea and still time out, while a proof can reach the right conclusion and still omit a necessary argument.

So “one model family” reaching gold-level results in both contests does not mean one general-purpose checkpoint simply handled both. IOI used programming specialists such as Nano-CC and Ultra-CC. The IMO project trained proof-focused checkpoints from Nemotron 3 Ultra and combined the general model with two specialists in its final workflow. What transfers is the foundation model and the method of adapting it, not an undifferentiated capability belonging to a single checkpoint.

IOI Scores Separate Training from Inference-Time Gains

The IOI 2025 experiments provide a useful slice of the process. Nano-CC scored 130 before post-training, rose to 280 after supervised fine-tuning (SFT), and reached 291 after reinforcement learning (RL). The progression suggests that SFT accounted for the largest jump, while RL added a smaller but still positive gain. The stages did not contribute equally.

After GenCorrect was added, the score rose to 468, above that year’s gold threshold of 438.3. GenCorrect is not another training round. At inference time, it repeatedly generates candidate code, evaluates it, and revises promising solutions. Ultra-CC reached 502 with the same strategy. The careful conclusion is that specialist training and test-time generation, selection, and correction all contributed. The scores cannot be reduced to “RL won gold” or “model size decides everything.”

Model Scale Changed the Adaptation Recipe

The programming work used 22,000 curated problems and synthetic reasoning traces, but the two models followed different training recipes. Nano-CC has 30 billion total parameters, with 3 billion active, and received both SFT and RL. Ultra-CC has 550 billion total parameters, with 55 billion active, and received SFT only. NVIDIA also reports that one SFT epoch on the stronger Ultra model outperformed the fully post-trained Nano model on IOI, ICPC, and LiveCodeBench Pro.

This comparison does not show that larger models need no adaptation, nor does it prove that RL has no value. It points to a more specific possibility: the starting capability of a foundation model changes the marginal value of further training. Nano gained most from SFT and then improved somewhat with RL, while Ultra was stronger after SFT. For engineering teams, the recipe should not be treated as a fixed pipeline. Task-specific evaluation should determine whether more RL is justified and whether inference-time search can deliver more useful gains than additional training.

IMO Put Critique and Revision Inside the Proof Workflow

The mathematics training targeted more than final answers. The SFT corpus contained 414,890 quality-filtered examples across 15,818 distinct proof problems. Its tasks included generating proofs, refining them, verifying arguments, and checking that verification was complete. RL used 9,597 proof problems selected near the model’s capability frontier. This data design treated finding gaps in an argument and judging whether a proof is complete as abilities the model needs to learn.

The final system did not ask one checkpoint to write a proof in a single pass. The general model and the SFT and RL specialists generated candidates, scored them, critiqued them, and revised promising attempts. A separate high-compute stage selected the submission. NVIDIA reported a score of 30 out of 42, with full marks on four of six problems. The stated gold threshold was 29, and official IMO graders scored the submitted proofs. The important design choice is not that this process guarantees success, but that candidate generation, quality judgment, and revision form a loop.

Evaluation Status and Reproduction Costs Both Matter

The two results have different evidentiary status. IMO submissions were scored by official graders, so the reported 30 points can be compared directly with the stated gold threshold of 29. The IOI 2026 score of 535.4 out of 600 came from a prospective test under contest time, internet-access, and submission constraints. It exceeded the stated gold threshold of 361.12 and the highest human score of 498.27, but NVIDIA explicitly describes the test as unofficial and unsupervised, and says it was not included in the official IOI ranking. Calling both results gold-level is reasonable. Treating them as equivalent official medal certifications is not.

For technical leaders, the reusable lesson is the system decomposition: domain data, specialist post-training, and feedback-driven inference. The scores themselves are not a promise. NVIDIA says both training and inference runs were substantial, but the available material gives no concrete figures from which to estimate a compute budget. The IMO project reportedly released training data, specialist checkpoints, inference code, and Nemotron-IMO-Bench, a set of 200 new olympiad-level problems. These assets offer entry points for inspection and attempted replication, but they do not substitute for independent validation. A practical decision should begin with the target task’s scoring criteria and budget, then measure separately what data, training, and the inference loop contribute.