Submitted by Xiangyi Li 25 ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces BenchFlow 35 2
Submitted by Xiangyi Li 65 SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks BenchFlow 1.75k 4