Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games Paper β’ 2606.19338 β’ Published Jun 17 β’ 50
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence Paper β’ 2606.14777 β’ Published Jun 10 β’ 216
DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing Paper β’ 2602.12205 β’ Published Feb 13 β’ 83
Running Agents Featured 182 HunyuanImage-3.0 π 182 Generate images from text prompts (PRO users only)
Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity Paper β’ 2508.05609 β’ Published Aug 7, 2025 β’ 29
Running Agents Featured 857 Qwen3 Demo π 857 Chat with an AI assistant that thinks before answering
mikeyandfriends/PixelWave_FLUX.1-dev_03 Text-to-Image β’ 12B β’ Updated Nov 5, 2024 β’ 1.23k β’ 196
Light-A-Video: Training-free Video Relighting via Progressive Light Fusion Paper β’ 2502.08590 β’ Published Feb 12, 2025 β’ 43
VideoRoPE: What Makes for Good Video Rotary Position Embedding? Paper β’ 2502.05173 β’ Published Feb 7, 2025 β’ 64
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding? Paper β’ 2501.05510 β’ Published Jan 9, 2025 β’ 44
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning Paper β’ 2501.03226 β’ Published Jan 6, 2025 β’ 43
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Paper β’ 2412.09596 β’ Published Dec 12, 2024 β’ 97