Resources for "Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs"
Xin Lai
xinlai
AI & ML interests
Multimodal LLM, LLM Reasoning, Point Cloud Segmentation, Image Segmentation
Recent Activity
liked a model 16 days ago
tencent/Hy3 upvoted a paper 29 days ago
Training Open Models for Agentic Phone Use upvoted a paper about 1 month ago
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?Organizations
None yet