One-shot on-policy distillation on a single query recovers most full-data OPD gains across domains and models by exploiting states visited during training.

Topological visualization of Rethinking On-Policy Distillation of Large Language Models II: One Training Example
Brave API

One-shot On-Policy Distillation (OPD) trains a student model on a single query, yet it recovers most of the performance gains achieved by full-data OPD across various task domains and model families. This effectiveness stems from the fact that a single query’s rollouts visit 71.5% of the states typically covered by full-data training, with most coverage occurring within the first 100 steps.

The process is characterized as data-overfed but algorithm-starved, meaning the student absorbs supervision increasingly slowly regardless of the dataset size. While adding semantically distinct queries improves state coverage—reaching 98.9% with just 16 queries to match full-data training—the primary bottleneck is the rate of alignment rather than data scarcity.

Generated 28d ago
Cerebras Thinking

This paper re-examines on-policy distillation (OPD) for large language models by asking how little task data is needed when the student’s own generated states are used for training. Rather than treating OPD as a method that requires large corpora of teacher-labeled trajectories, it studies a minimal setting: a single training query, from which the student produces on-policy rollouts and receives dense teacher feedback. The central empirical claim is that this one-shot, on-policy setup can recover a large fraction of the gains normally obtained from full-data OPD, across different domains and model configurations.

The key insight is that on-policy distillation is not simply “learning from more examples,” but learning from states that the student actually visits during generation. A single query can still expose a rich set of partial sequences, decision points, and high-uncertainty contexts, especially under stochastic decoding or long-horizon generation. When the teacher corrects these states, the student updates its policy in regions that are directly relevant to its current behavior, producing changes that generalize beyond the surface content of the prompt. In other words, the value of OPD may lie less in coverage of many input distributions and more in targeted correction of the student’s own distribution.

This matters because it suggests that LLM distillation can be far more data-efficient than commonly assumed, particularly when the goal is to improve a student model’s policy rather than merely imitate a fixed dataset. If a small number of self-generated trajectories can capture much of the benefit of large-scale distillation, OPD becomes more practical for continual adaptation, domain-specific tuning, privacy-sensitive settings, and rapid deployment. The paper therefore reframes OPD as a state-driven learning process and raises an important question for future work: how much apparent data diversity is truly necessary, versus how much can be substituted by on-policy exploration and teacher-guided correction?

Generated 28d ago
Open-Weights Reasoning

Rethinking On-Policy Distillation of Large Language Models II: One Training Example examines a surprising data-efficiency question in on-policy distillation (OPD): how much training data is actually needed for a student LLM to inherit teacher behavior? Rather than assuming that effective distillation requires a broad, domain-specific query set, the paper studies the extreme case of one-shot OPD, where the student is trained using only a single query or training example. The central finding is that this minimal setup can recover a large fraction of the gains typically obtained from full-data OPD, across multiple domains and model families.

The key insight is that OPD does not primarily benefit from prompt diversity in the naive sense; it benefits from supervision over the states the student actually visits. When the student rolls out from a single query, it exposes a distribution of internal decision points, partial generations, and error-prone continuations. Teacher feedback on those visited states can therefore correct the student in regions that are already relevant to its own behavior. In other words, the single example acts as a probe into the student’s state space, and the distillation signal is most valuable where the student is likely to operate. This reframes the role of training data in OPD: coverage of useful student-generated states may matter more than surface-level coverage of many different prompts.

The result matters because it substantially lowers the practical cost of on-policy distillation. If a single well-chosen query can yield most of the benefit of a larger dataset, OPD becomes more attractive for domain adaptation, continual learning, privacy-sensitive settings, and low-budget fine-tuning. It also shifts the research focus from simply collecting more distillation prompts to understanding which prompts or tasks induce the most informative state coverage. More broadly, the paper challenges the common assumption that teacher-student alignment requires large-scale data, suggesting that the structure of on-policy rollouts can make distillation far more sample-efficient than previously thought.

Generated 28d ago
Sources