V-JEPA 2 Collection A frontier video understanding model developed by FAIR, Meta, which extends the pretraining objectives of https://ai.meta.com/blog/v-jepa-yann • 8 items • Updated Jun 13, 2025 • 229
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Paper • 2607.14187 • Published Jul 15 • 31
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Paper • 2607.14187 • Published Jul 15 • 31
Running on Zero MCP 3 RxBrain Embodied Cognition 🔮 3 Embodied reasoning + visual imagination (Hy-RxBrain 1.0)
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Paper • 2605.12500 • Published May 12 • 199
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Paper • 2604.23781 • Published Apr 26 • 34
HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System Paper • 2604.14125 • Published Apr 15 • 21
Beyond Language Modeling: An Exploration of Multimodal Pretraining Paper • 2603.03276 • Published Mar 3 • 107
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos Paper • 2602.06949 • Published Feb 6 • 37
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation Paper • 2601.02204 • Published Jan 5 • 64