view article Article Distribution Matching Prevents Mode Collapse in Training Reasoning Models about 1 month ago • 2