Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 14 days ago • 38
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published 19 days ago • 105
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Paper • 2607.27380 • Published 16 days ago • 71
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Paper • 2607.24027 • Published 18 days ago • 37
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report Paper • 2607.18367 • Published 24 days ago • 60
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published 28 days ago • 44
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published 27 days ago • 139
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 87
TurboServe: Serving Streaming Video Generation Efficiently and Economically Paper • 2606.19271 • Published Jun 17 • 38