Benchmark
AI doesn't find bugs unless you tell it what's wrong

About Benchmark
Given a real repository, an agent must discover & repair as many bugs as they can. Agents are not given any hint about the type of bug or its location.
In the maker’s words
New benchmark from researchers at Meta, Stanford, Harvard, UW, including the researchers who've worked on SWE-bench, ProgramBench etc. Most benchmarks just test if AI can fix a problem you've already pointed out. But obviously it would be much better to fix problems before you or any user runs into it. Like, isn't it crazy that we still have to wait for people to open tickets before a lot of obvious bugs get found? We wanted to test that capability at scale. Turns out that models still are terrible at it (best setup we tested still fixed < 5% of bugs) We have 100 repos of 22 languages and 4k bugs between. All the bugs are real-world bugs from github. We do a lot of filtering to ensure everything can be solved in this setting. =========================================== Sol 5.6 (xhigh) 4.7% $7,230 Luna 5.6 (xhigh) 2.5% $224 Terra 5.6 (xhigh) 1.5% $357 Luna 5.6 (high) 1.4% $28 Opus 5 (xhig…
Where people found it
- Hacker NewsShow HN: Benchmark: AI doesn't find bugs unless you tell it what's wrong5 points2 comments2 days ago
More sites like Benchmark
- Giving Opus 5.5 a simulated paint canvasstillwet.art
- AIHOT一个自己找热点、自己写日报的网站框架。把信源和精选标准换成你的,它就是你的行业热点站。
- Offrunmanage every coding agent from one workspace
- OpenDotsYour always-on AI coworkers that move between text, calls, and Slack.
- Pi podRun your pi coding agent in sandboxes on your own server
- Ledge.shRunnable Markdown Notes