FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 3 days ago • 130
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 4 days ago • 196
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence Paper • 2608.11341 • Published 17 days ago • 64
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published 16 days ago • 113
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published 21 days ago • 109
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents Paper • 2607.18754 • Published Jul 21 • 25
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents Paper • 2607.18754 • Published Jul 21 • 25
BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discovery Paper • 2606.20997 • Published Jun 19 • 13
Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Paper • 2504.01990 • Published Mar 31, 2025 • 305