MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models Paper • 2607.07673 • Published 12 days ago • 14
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published 4 days ago • 178
KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill Paper • 2607.12625 • Published 5 days ago • 56
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras Paper • 2607.12993 • Published 6 days ago • 131
Phone Segmentation and Recognition through Phonological Activation Mapping Paper • 2607.09020 • Published 10 days ago • 9
mxcui/maxmin-imdb-ppo-prop0.2-alpha1.0-seed42-mean_kl0.1-Qwen-Qwen3-4B-Base Text Generation • 4B • Updated 8 days ago • 492 • 1
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation Paper • 2607.06559 • Published 13 days ago • 94
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams Paper • 2606.21337 • Published Jun 19 • 75
OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources Paper • 2605.29250 • Published May 28 • 81
Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents Paper • 2605.30723 • Published May 29 • 17
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 252
Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models Paper • 2605.31603 • Published May 29 • 8
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Paper • 2605.18018 • Published May 18 • 33
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? Paper • 2605.22109 • Published May 21 • 171