S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published 4 days ago • 25
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 3 days ago • 119
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published 1 day ago • 357
MolmoWeb-Data Collection This is the collection of all datasets in MolmoWebMix. • 6 items • Updated Mar 24 • 33
yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K Viewer • Updated Mar 13 • 260k • 5.17k • 52
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published 20 days ago • 445
Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report Paper • 2608.15763 • Published 13 days ago • 50
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published 11 days ago • 64
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses Paper • 2608.24876 • Published 10 days ago • 28
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks Paper • 2608.23035 • Published 11 days ago • 41
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 11 days ago • 205
Fara-1.5: Scalable Learning Environments for Computer Use Agents Paper • 2606.20785 • Published Jun 18 • 9
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents Paper • 2608.17393 • Published 17 days ago • 24