HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 5 days ago • 238
view article Article Measuring benchmark optimization in speech recognition +5 tlebryk02, bezzam, aliceebaird, dayllon, jpc, jens-hume-ai, tzirakis • 16 days ago • 65
view article Article We changed one line and the benchmark score moved 0.21 AUROC FINAL-Bench • 14 days ago • 14
view article Article Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers +1 tomaarsen, NohTow, raphaelsty • 19 days ago • 106
SWE-bench Collection SWE-bench (Lite, Verified, Multimodal, Multilingual) all in one place! • 5 items • Updated Dec 14, 2025 • 16
view article Article State of Open Models: Summer 2026 Observations +1 AdinaY, multimodalart, irenesolaiman • 23 days ago • 186
view article Article LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge LiquidAI • 24 days ago • 50
Muse Glimmer Collection Muse Glimmer 30B: multimodal agentic model for local deployment. BF16 weights, GGUF k-quants, ExecuTorch builds, DFlash drafter. • 4 items • Updated 26 days ago • 107
view article Article Meta is back with Muse Glimmer: local, agentic, multimodal, and open source +2 pcuenq, merve, burtenshaw, ariG23498 • 27 days ago • 110
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks? Paper • 2605.11086 • Published May 11 • 3
view article Article Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident +2 hlarcher, XciD, raphael-gl, chris-rannou • Jul 27 • 487