On-Policy Self-Adaptation Collection Checkpoints of OPSA on different base models • 5 items • Updated about 21 hours ago • 1
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 4 days ago • 135
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 4 days ago • 135
On-Policy Self-Adaptation Collection Checkpoints of OPSA on different base models • 5 items • Updated about 21 hours ago • 1
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 4 days ago • 135
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control Paper • 2604.26326 • Published May 10 • 15
Learning Self-Correction in Vision-Language Models via Rollout Augmentation Paper • 2602.08503 • Published Feb 9 • 3
On-Policy Self-Adaptation Collection Checkpoints of OPSA on different base models • 5 items • Updated about 21 hours ago • 1
On-Policy Self-Adaptation Collection Checkpoints of OPSA on different base models • 5 items • Updated about 21 hours ago • 1