JevBench, a reproducible benchmark for typed decision models
About JevBench, a reproducible benchmark for typed decision models
JevBench Capability Score ranks decision models within cost and latency caps. Compare intelligence, calibration, speed, and cost.
In the maker’s words
Hi HN! I built JevBench because Jev kicks ass, and the world deserves to know how the serious open source and fake lookalike projects really perform in comparison. Jev-class models return bounded choices and probabilities instead of text, and are disruptively faster and cheaper than LLMs, while being similarly intelligent on the text input they operate on. JevBench allows looking at accuracy, latency and price all at once, in a weighted way - you can even configure the weighting. A full run asks 534 English decisions. The v1.3 score combines chance-corrected Intelligence, Calibration, Speed and Cost. Leaderboard right now: #1 - Jev 74.4 #2 - SemIf 73.1 #3 - djev 73.0 #4 - Winnow-12B Q8 71.2 #5 reflex 4B 70.3. MIT harness, public items, frozen artifacts, scoring code and public per-task outcomes: https://github.com/fstandhartinger/jevbench Two no-signup demos: https://who-is-right.app.min…
Where people found it
- Hacker NewsShow HN: JevBench, a reproducible benchmark for typed decision models153 points39 comments11 days ago
- Hacker NewsJevBench: Benchmark for Jev-Class Models3 points0 comments12 days ago
More sites like JevBench, a reproducible benchmark for typed decision models
- Giving Opus 5.5 a simulated paint canvasstillwet.art
- AIHOT一个自己找热点、自己写日报的网站框架。把信源和精选标准换成你的,它就是你的行业热点站。
- Offrunmanage every coding agent from one workspace
- OpenDotsYour always-on AI coworkers that move between text, calls, and Slack.
- Pi podRun your pi coding agent in sandboxes on your own server
- Ledge.shRunnable Markdown Notes