8 12 9

Richard Zhuang PRO

RZ412

https://richardzhuang0412.github.io

AI & ML interests

LLM Routing, LLM + Games, Post-Training, Agents

Recent Activity

updated a dataset about 8 hours ago

DCAgent3/terminal_bench_2_fresh_gptlongtezos_step5100__Qwen3_32B_20260515_082058

published a dataset about 8 hours ago

DCAgent3/terminal_bench_2_fresh_gptlongtezos_step5100__Qwen3_32B_20260515_082058

updated a dataset about 9 hours ago

DCAgent3/terminal_bench_2_fresh_gptlongtezos_step5400__Qwen3_32B_20260515_082005

View all activity

Organizations

upvoted a paper about 1 month ago

On Data Engineering for Scaling LLM Terminal Capabilities

Paper • 2602.21193 • Published Feb 24 • 102

upvoted a paper 3 months ago

SkillOrchestra: Learning to Route Agents via Skill Transfer

Paper • 2602.19672 • Published Feb 23 • 58

upvoted 2 collections 5 months ago

OpenThinker-Agent

Collection

5 items • Updated Dec 6, 2025 • 10

Olmo 3 Post-training

Collection

All artifacts for post-training Olmo 3. Datasets follow the model that resulted from training on them. • 32 items • Updated Dec 23, 2025 • 54

upvoted a paper 5 months ago

DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle

Paper • 2512.04324 • Published Dec 3, 2025 • 159

upvoted a paper 8 months ago

DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search

Paper • 2509.25454 • Published Sep 29, 2025 • 148

upvoted an article 9 months ago

Article

SmolLM3: smol, multilingual, long-context reasoner

eliebak, cmpatino, anton-l, edbeeching, m-ric, nouamanetazi, akseljoonas, guipenedo, hynky, clefourrier, SaylorTwift, kashif, qgallouedec, hlarcher, glutamatt, Xenova, reach-vb, ngxson, craffel, lewtun, loubnabnl, lvwerra, thomwolf

•

Jul 8, 2025

• 775

upvoted 2 collections 10 months ago

Reasoning Datasets

Collection

50 items • Updated Jun 8, 2025 • 11

Reasoning Models

Collection

53 items • Updated Jun 8, 2025 • 1

upvoted an article about 1 year ago

Article

Reasoning Datasets Competition

bespokelabs

•

Apr 9, 2025

• 38

upvoted 2 papers over 1 year ago

PokerBench: Training Large Language Models to become Professional Poker Players

Paper • 2501.08328 • Published Jan 14, 2025 • 19

EmbedLLM: Learning Compact Representations of Large Language Models

Paper • 2410.02223 • Published Oct 3, 2024 • 3

Richard Zhuang PRO

AI & ML interests

Recent Activity

Organizations

RZ412's activity

SmolLM3: smol, multilingual, long-context reasoner

Reasoning Datasets Competition