MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes Paper • 2509.24945 • Published Sep 29, 2025 • 7 • 1
Source-Grounded Semantic Reinforcement Learning for Low-Resource Target-Language Generation Paper • 2605.29502 • Published May 28 • 1
WTF GENIUS PAPERS Collection Papers that made me appreciate my major and my life a little more. obs=Observation, innov=Innovation. Most papers are abt improving tiny models. • 218 items • Updated about 5 hours ago • 50
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes Paper • 2509.24945 • Published Sep 29, 2025 • 7
WTF GENIUS PAPERS Collection Papers that made me appreciate my major and my life a little more. obs=Observation, innov=Innovation. Most papers are abt improving tiny models. • 218 items • Updated about 5 hours ago • 50
Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training Paper • 2607.19058 • Published 2 days ago • 4
WTF GENIUS PAPERS Collection Papers that made me appreciate my major and my life a little more. obs=Observation, innov=Innovation. Most papers are abt improving tiny models. • 218 items • Updated about 5 hours ago • 50
WTF GENIUS PAPERS Collection Papers that made me appreciate my major and my life a little more. obs=Observation, innov=Innovation. Most papers are abt improving tiny models. • 218 items • Updated about 5 hours ago • 50
LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Paper • 2607.18110 • Published 3 days ago • 11
WTF GENIUS PAPERS Collection Papers that made me appreciate my major and my life a little more. obs=Observation, innov=Innovation. Most papers are abt improving tiny models. • 218 items • Updated about 5 hours ago • 50
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Paper • 2607.17524 • Published 3 days ago • 5
WTF GENIUS PAPERS Collection Papers that made me appreciate my major and my life a little more. obs=Observation, innov=Innovation. Most papers are abt improving tiny models. • 218 items • Updated about 5 hours ago • 50
Distilled Reinforcement Learning for LLM Post-training Paper • 2607.17247 • Published 4 days ago • 8
Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel Paper • 2607.14431 • Published 8 days ago • 11
WTF GENIUS PAPERS Collection Papers that made me appreciate my major and my life a little more. obs=Observation, innov=Innovation. Most papers are abt improving tiny models. • 218 items • Updated about 5 hours ago • 50