On-policy distillation rises—and faces growing pains

The gist
On-policy distillation is shaking up AI training, promising efficiency over RLHF—but now faces fierce debate over collapse and teacher-student flaws.
What to know
- By September 2026, OPD let student models generate their own outputs while teacher models scored each token, boosting unlabeled scores by 5.88 points (+12.3%) in key setups.
- The shift from RLHF was gradual, with models like Llama 3 and DeepSeek R1 evolving toward multi-stage, distillation-heavy pipelines instead of classic reward loops.
- Despite efficiency gains, recent research warns that OPD can still inherit teacher mistakes or collapse entirely—one case saw accuracy spike from 25 to 46 before crashing to 11.
Token-Level Supervision Takes Over
Teacher models now guide student outputs at each token, enabling faster, cheaper, and more stable training than classic RL loops—while new routing methods squeeze extra gains from unlabeled data.
By late September, the mechanism had become unusually concrete: according to Into AI, OPD has the student generate its own responses, then a stronger teacher scores them token by token as the student minimizes reverse KL against the teacher’s next-token distribution; in the reported setup, the paper used a teacher model called QEN330B A3B instruct 2507 to guide the distillation process, and it compared different KL divergences, forward KL and reverse kl. That same token-level supervision was pushed further in Hugging Face Daily Papers’ MOPD-Router, which routes over the teacher pool at each token and, crucially, showed that “On unlabeled data, it improves the overall score by 5.88 (+12.3%) points over Mean aggregation; on domain-labeled data, it outperforms standard MOPD by 3.95 (+7.8%) points without using available domain labels.”
What made the shift feel inevitable was that the same teacher-scored setup also looked cheaper and more operationally stable than reward-heavy alternatives: Direct On-Policy Distillation Boosts Large Model Performance Efficiently reported that for the larger model “Quen31.7B,” it “increased its accuracy from 48.3% to 58.3% on the AIME 2024 benchmark… in just hours on eight A100 GPUs,” whereas “running RL directly on Quen 31.7 billion would have required using 32 A100 GPUs for at least a week.” Daily Paper Cast likewise wrote that “OPD2 represents a promising advancement in leveraging the training trajectory of teacher models to enhance reasoning capabilities in student models,” while Async OPD Experiments Show Staleness Resilience and Throughput Gains found a teacher-guided KL setup could scale in practice, with “Async OPD consistently achieved the highest throughput and overlap… up to 2.7 times the Strix sync baseline while achieving the best or tied for best final accuracy,” and reported that forward kl remained robust to stale rollouts.
Distillation Pipelines Eclipse RLHF
The once-standard RLHF loop has been quietly replaced by multi-stage, iterative distillation recipes that prioritize verifiability and consolidation, reshaping how leading models are trained.
What changed was not a sudden abandonment of RLHF so much as a long drift away from its original, neat loop. Interconnects AI traces the arc from “InstructGPT (Mar. 2022) — the canonical 3 steps” of SFT, reward-model training, and PPO, to “Llama 2 (Jul. 2023) — multi-stage RLHF” where each round became “rejection sampling → PPO,” and then to “Llama 3 (Jul. 2024) — a complex multi-stage recipe” of “reward model → sample K per prompt → rejection sampling → SFT → DPO” over “6 rounds, best models seed the next,” showing recipes becoming iterative, staged, and less like one static RL pass.
By 2024 and 2025, those recipes were already being bent toward verifiability, filtering, and consolidation, which is why OPD later looked like a cleaner continuation rather than a rupture. Interconnects AI notes that Llama 3 used “No online RL — the RM only filters; run over 6 rounds, best models seed the next,” that Tülu 3 was “Curated prompts → SFT → DPO → RLVR,” and that DeepSeek R1 evolved from “R1-Zero — pure RL (GRPO) on the base, no SFT” to “cold-start SFT → reasoning RL → rejection-sampling SFT → final RL → distill to dense,” before late-2025/2026 recipes emphasized consolidation via distillation from multiple RL-trained components.
Instability Shadows OPD's Rise
Despite early accuracy surges, OPD models face collapse risks and instability as teacher-student discrepancies and latent supervision failures reveal the method's unresolved vulnerabilities.
By mid-September, the push to treat on-policy distillation as the settled post-training default was already meeting organized resistance. Hugging Face Daily Papers’ September 8 review said “one failure mode now dominates the field: collapse,” and drew “a clear line between what is settled and what is still disputed,” while Calibrating Teacher--Student Discrepancy for On-Policy Distillation argued standard OPD can absorb teacher-side mistakes; crucially, “while retaining only about 52--65% of the original teacher--student discrepancy as the optimization signal, Cal-OPD consistently outperforms standard OPD and its variants across model scales.”
The end of the month made those objections concrete. LastOPD: Taming Collapse in Latent On-Policy Distillation showed “Early gain, late collapse: latent supervision alone lifts MATH-500 accuracy from 25 to 46 in 10 steps, but subsequent training degrades,” and when distilling Qwen3-4B and Qwen3-8B into Qwen3-1.7B-Base the score fell “down to 11 with no recovery”; meanwhile, later September 2026 variants directly targeted “privilege illusion,” reporting “average improvements of 7.5 points for LLM based setups and 6.0 points for VLM based setups across eight different benchmarks,” even as other work explicitly framed teacher-student divergence as a source of instability.



