WHALE: A Simple Recipe for Joint Harness-Weight Optimization Paper • 2609.00196 • Published 7 days ago • 30
Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks Paper • 1810.00825 • Published Oct 1, 2018 • 1
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems Paper • 2510.02263 • Published Oct 2, 2025 • 9
Discrete Infomax Codes for Supervised Representation Learning Paper • 1905.11656 • Published May 28, 2019
Learning Dynamics of Attention: Human Prior for Interpretable Machine Reasoning Paper • 1905.11666 • Published May 28, 2019
DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature Paper • 2301.11305 • Published Jan 26, 2023 • 2
Self-Guided Masked Autoencoders for Domain-Agnostic Self-Supervised Learning Paper • 2402.14789 • Published Feb 22, 2024
Personalized Preference Fine-tuning of Diffusion Models Paper • 2501.06655 • Published Jan 11, 2025 • 1
Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models Paper • 2502.17387 • Published Feb 24, 2025 • 8
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems Paper • 2510.02263 • Published Oct 2, 2025 • 9
Recursive Introspection: Teaching Language Model Agents How to Self-Improve Paper • 2407.18219 • Published Jul 25, 2024 • 3
Guided Data Augmentation for Offline Reinforcement Learning and Imitation Learning Paper • 2310.18247 • Published Oct 27, 2023
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning Paper • 2503.07572 • Published Mar 10, 2025 • 48
Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs Paper • 2503.01307 • Published Mar 3, 2025 • 39