🔍 Read the full analysis: One Model Family, Two Challenges: Nemotron’s IOI And IMO Results on ThorstenMeyerAI.com
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Hugging Face says specialized systems built from its Nemotron 3 family scored 535.4 out of 600 in an unofficial 2026 IOI run and 30 out of 42 at the 2026 IMO. The IMO proofs were graded by official graders; the IOI result was outside the competition’s official ranking. The company says it plans to provide IMO checkpoints, datasets and a 200-problem benchmark.
As detailed in the original analysis, Hugging Face says two specialist systems built from its Nemotron 3 model family reached gold-threshold scores in programming and mathematics at the 2026 International Olympiad in Informatics (IOI) and International Mathematical Olympiad (IMO). The IMO proofs received official grading and earned 30 of 42 points; the IOI system scored 535.4 of 600 in a competition-style run that was unofficial and excluded from the official ranking.
For the IOI, Hugging Face used a competition-specific version of Nemotron-3-Ultra-CC, trained with supervised fine-tuning, and paired it with GenCorrect, an iterative method for generating, evaluating and refining code solutions. The company says the system ran prospectively under the same time, internet-access and submission constraints as human contestants. Its reported score was above the stated 361.12 gold threshold and the top human score of 498.27. Because the run was unofficial, those comparisons do not make it an official medal or ranking result.
For the IMO, the team combined the general Nemotron 3 Ultra model with supervised fine-tuning and reinforcement-learning checkpoints in a natural-language proof system. It generated candidate proofs, scored and critiqued them, and revised selected attempts. According to Hugging Face, official graders awarded the submitted work 30 points out of 42, above the stated gold threshold of 29, with full credit on four of six problems. The company says the system used no formal prover, external tools or internet access.
The training setups differed by subject. Hugging Face reports that the IOI work used 22,000 programming problems and synthetic reasoning traces. For IMO, its supervised fine-tuning data included 414,890 quality-filtered examples from 15,818 proof problems, while the reinforcement-learning model was trained on 9,597 problems selected near the model’s capability frontier. These details and scores come from the company’s account.
What the Two Scores Show
The results show reported performance by two systems adapted from the same model family for different tasks. They are not directly equivalent: the IMO result reflects official grading of submitted proofs, while the IOI score comes from a run outside the official ranking. The distinction is relevant when comparing the scores or describing the systems’ competition results.
Hugging Face says its approach combines domain-specific training with inference-time checking and revision. The reported workflows generated multiple candidates and evaluated or refined selected outputs. The results document the systems’ performance on the reported problems; they do not establish performance across mathematics, programming or real-world workloads more broadly.
The competitions assess different capabilities. IOI tasks require executable programs to pass tests, while IMO problems require written mathematical proofs. The reported scores therefore concern separate evaluations, and neither result establishes how well a system would perform beyond its respective olympiad problems.
As an affiliate, we earn on qualifying purchases.
From IOI Experiments to Proofs
Hugging Face describes the 2026 work as an extension of its IOI 2025 experiments, where it studied how post-training and additional computation at inference could improve coding scores. The company reported that a Nemotron-3-Nano-CC model rose from 130 points before post-training to 280 after supervised fine-tuning and 291 after reinforcement learning. GenCorrect then raised the reported score to 468, above that year’s stated gold threshold of 438.3; an Ultra-CC version scored 502 with the same test-time strategy.
The 2026 projects applied related ideas to two different formats: competition programming and written mathematical proof. For IMO, the company says supervised fine-tuning and reinforcement-learning checkpoints had complementary strengths, leading the team to combine them with the general model. The current report concerns the outcomes and methods described by Hugging Face; the supplied material does not include independent replication or a separate external assessment of the IOI run.
“Success at both points to something broader.”
— Hugging Face
As an affiliate, we earn on qualifying purchases.
Validation and Generalization Questions
The IOI score was unofficial, and the provided account does not describe independent verification, a detailed audit of the run or replication by another group. Although Hugging Face says the system followed competition-like constraints, it was not included in the official ranking. The reported score should therefore be treated as a company-reported benchmark result, not an official competition placement.
It is also unclear how reliably either system would perform on other competitions, unfamiliar proof styles or practical workloads. The supplied information does not establish the effect of differences in data selection, compute budgets or evaluation setup on the results. The IMO proofs did receive official grading, but that does not by itself establish general capability beyond the submitted problems.
Hugging Face says it is releasing IMO checkpoints and datasets, but the source material’s final description of the associated repository is incomplete. It does not provide a timetable for further releases or say whether the IOI evaluation will receive external scrutiny.
mathematics Olympiad preparation books
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Checkpoints and Benchmarks to Come
Hugging Face says its Nemotron Labs IMO 2026 collection includes supervised fine-tuning and reinforcement-learning checkpoints, both training datasets, and Nemotron-IMO-Bench, a benchmark of 200 olympiad-level problems. The company also points to an IMO paper describing the training and generate-verify-refine system, alongside a NeMo-Skills repository.
The materials are intended to let researchers examine the methods and evaluate the models on additional problems. Further evidence could include independent evaluations, replication of the reported scores and more detail about the IOI run’s verification. No release timetable or independent IOI assessment is specified in the supplied account.
As an affiliate, we earn on qualifying purchases.
Key Questions
Did a Nemotron system officially win an IOI medal?
No. Hugging Face reported a 535.4 out of 600 score from an unofficial run. The result was not entered in the IOI’s official ranking, so it is not an official medal or placement.
Was the IMO result officially graded?
Yes. Hugging Face says official IMO graders awarded the submitted proofs 30 of 42 points, including full credit on four of six problems. The company states that the score was above the gold threshold of 29.
How did the systems generate their answers?
The IOI system paired a fine-tuned Nemotron-3-Ultra-CC model with GenCorrect, which iteratively generated, evaluated and refined code. The IMO system generated candidate proofs, scored and critiqued them, then revised selected attempts.
Are the results independently verified?
The supplied account does not report independent verification or replication of the IOI run. It says the IMO submissions were officially graded, but that is distinct from an independent replication of the training and evaluation.
What materials does Hugging Face say it will make available?
The company says its IMO collection includes SFT and RL checkpoints, training datasets, and a 200-problem benchmark. The source material gives no release timetable and leaves some repository details incomplete.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
