FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 10 days ago • 147
Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning Paper • 2604.16890 • Published Apr 18 • 2
Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning Paper • 2604.16890 • Published Apr 18 • 2
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning Paper • 2607.07708 • Published Jul 8 • 89
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? Paper • 2606.24530 • Published Jun 23 • 66
view article Article PhysicsIntern: from an Autonomous Benchmark-runner to a Research Sidekick dlouapre • Jun 11 • 7
Running 57 physics-intern: an Autonomous Agent for Physics Research 📝 57 Explore an autonomous AI workflow for physics research
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research Paper • 2606.07591 • Published May 28 • 103
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation Paper • 2605.31264 • Published May 29 • 132
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration Paper • 2605.20025 • Published May 19 • 192
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning Paper • 2605.06326 • Published May 7 • 26
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning Paper • 2510.01833 • Published Oct 2, 2025
QCBench: Evaluating Large Language Models on Domain-Specific Quantitative Chemistry Paper • 2508.01670 • Published Aug 3, 2025
PolyReal: A Benchmark for Real-World Polymer Science Workflows Paper • 2604.02934 • Published Apr 3 • 1
$δ$-mem: Efficient Online Memory for Large Language Models Paper • 2605.12357 • Published May 12 • 133