Daily AI Paper

Daily AI Paper

2026-09-06

Archived
2026-09-06

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Zixuan Fu, Bingxiang He, Yuxin Zuo, Haohuan Huang, Jinqian Zhang, Ruhang Xiao, et al.

arXiv:2609.04172v1

Large language models often get better with more data, but this paper asks a surprising question: how little data can you use and still get most of the benefit? The authors study on-policy distillation, where a student model generates its own rollouts and a teacher model provides detailed token-level guidance. Instead of training on a full dataset, they try the extreme case of just one training example. Surprisingly, that single example can keep improving for hundreds of steps and recover much of the performance of full-data training across different tasks and model families. The key insight is state coverage: even one query can drive the student through a wide range of internal states, creating rich supervision. Adding a few semantically different examples expands that coverage further, until a small set can match full-data results. The bigger message is that this method is not starved for data so much as slowed by the learning process itself, which points to a new direction for making post-training far more efficient.

Previous daily papers

2026-09-07WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data2026-09-07UniMate: One Unified Model to Animate Diverse Skeletons2026-09-06Compile by Training: Turning Natural-Language Specifications into Local Neural Functions2026-09-06SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents2026-09-05From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research2026-09-05Last Translation Benchmark2026-09-03A Common Measure of Communication for Speech Brain-Computer Interfaces2026-09-03Discriminative World Models for Web Agents2026-09-03Post-Training Language Models for Gold-Medal Performance in Coding Competitions2026-09-02Mechanism Design for Alignment and Control2026-09-02Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation2026-09-02The Rise of Verbal Reinforcement Learning2026-09-01Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions2026-09-01Aspire: Can Models Self-Evolve from Vague Goals?2026-09-01PaperGym: Rubric-Centered Evolution for Research-Plan Generation2026-08-31DARTS: Decoder-Aware Representation Tuning via Surgery for Model Merging2026-08-31Video Generative Models as Geometry Learner2026-08-31When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI2026-08-30CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes2026-08-30Boosting LLM Exploration via Weak-Model Guidance in RLVR2026-08-30CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators2026-08-30TTPO: Test-Time Policy Optimization2026-08-29TTPO: Test-Time Policy Optimization2026-08-28TTPO: Test-Time Policy Optimization2026-08-27Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings2026-08-26BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes2026-08-25ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings2026-08-24VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences2026-08-23Inducing Task Models from Computer-Use Traces2026-08-22Inducing Task Models from Computer-Use Traces2026-08-21Inducing Task Models from Computer-Use Traces2026-08-20SPADE: Self-Play in Adaptive Synthetic Executable Environments2026-08-19From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation2026-08-18Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text2026-08-17Universal Thermodynamic Interatomic Potentials for Crystalline Materials2026-08-16OmniScientist: An Omni-Modal Omni-Discipline AI Scientist