Presents an end-to-end specialization pipeline (22k curated problems + SFT/RL) that trains Nemotron-3 models to competitive-programming level on IOI/ICPC tasks.

Topological visualization of Post-Training Language Models for Gold-Medal Performance in Coding Competitions
Brave API

Nemotron-3-Nano-CC (30B total / 3B active parameters) and Nemotron-3-Ultra-CC (550B total / 55B active parameters) are the primary models developed using this pipeline. Nemotron-3-Nano-CC utilizes both Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), while Nemotron-3-Ultra-CC relies solely on SFT. Both models employ GenCorrect, an iterative test-time refinement strategy that generates, evaluates, and corrects diverse solutions under submission constraints.

The pipeline curates 22,000 problems and generates over 1.2 million reasoning traces using DeepSeek-V4-Flash. On the IOI 2025, Nemotron-3-Nano-CC achieved a score of 468/600, exceeding the gold threshold of 438.3, while Nemotron-3-Ultra-CC reached 502/600. In a live prospective evaluation during IOI 2026, the competition-specific Ultra-CC system scored 535.4/600, surpassing both the gold threshold and the highest-scoring human contestant.

Generated Sep 3, 2026
Open-Weights Reasoning

The material presents an end-to-end post-training pipeline for specializing Nemotron-3 language models on competitive programming. At its core is a curated dataset of roughly 22,000 coding problems, followed by supervised fine-tuning and reinforcement learning. The pipeline is designed around competition-style tasks, with training signals likely grounded in executable correctness, test-case validation, and difficulty-calibrated problem selection. Evaluation is framed in terms of IOI/ICPC-style benchmarks, where the resulting models are reported to reach gold-medal-level performance.

A key contribution is the systematic specialization stack rather than a single model architecture or prompting trick. The work emphasizes that high-quality problem curation, SFT, and RL can jointly transform a general-purpose model into a system capable of the full competitive-programming workflow: parsing formal constraints, selecting appropriate algorithms, managing edge cases, optimizing complexity, and producing correct code under judge-like evaluation. This suggests that the distribution of training problems and the verifiability of the reward signal are central to achieving expert-level coding ability.

The result matters because competitive programming is a stringent proxy for algorithmic reasoning, constraint satisfaction, and reliable code generation under hidden tests. Reaching gold-medal performance indicates that post-training can push language models well beyond generic coding assistance into a regime where they can solve high-difficulty, verifiable programming tasks at an expert level. More broadly, the paper provides a useful template for building domain-specialized coding models and highlights the value of competition-derived, execution-verified training environments.

Generated Sep 3, 2026
Sources