World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation Paper • 2608.05369 • Published 5 days ago • 24
Invisible Shortcuts: Why Vision Encoders Know Your Camera Paper • 2608.05424 • Published 5 days ago • 16
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published 4 days ago • 42
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published 5 days ago • 55
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published 6 days ago • 89
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing Paper • 2607.29025 • Published 10 days ago • 15
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published 11 days ago • 181