H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models Paper • 2608.13049 • Published 4 days ago • 16
Intern-S2-Preview: Scientific Agentic Foundation Model Paper • 2608.13505 • Published 4 days ago • 56
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published 16 days ago • 255
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents Paper • 2608.08389 • Published 8 days ago • 12
OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching Paper • 2608.08097 • Published 9 days ago • 24
FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds Paper • 2608.01049 • Published 15 days ago • 13
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents Paper • 2608.04003 • Published 13 days ago • 33
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning Paper • 2608.04007 • Published 13 days ago • 19
N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Paper • 2607.23782 • Published 22 days ago • 77
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published 18 days ago • 184