DEV Community

#benchmarks

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
voiceloop: the fastest voice agent loop in the browser is now open source

voiceloop: the fastest voice agent loop in the browser is now open source

2
Comments 1
5 min read
Our task categorizer routed 46% of tasks. A $0.04/MTok decision model routes 97%.

Our task categorizer routed 46% of tasks. A $0.04/MTok decision model routes 97%.

Comments
3 min read
94.7% on LoCoMo — and why most of the gap between published memory numbers isn't the memory

94.7% on LoCoMo — and why most of the gap between published memory numbers isn't the memory

Comments
9 min read
Fast Decisions in Agent Workflows: Laya vs TypeSafe Jev

Fast Decisions in Agent Workflows: Laya vs TypeSafe Jev

Comments
5 min read
OCR that looked like it worked

OCR that looked like it worked

1
Comments 1
6 min read
DeepMind agents blew the whistle on cheating agents

DeepMind agents blew the whistle on cheating agents

5
Comments
4 min read
Giving a coding agent more time barely helps

Giving a coding agent more time barely helps

Comments 1
2 min read
The render said clean, the file said broken: why AI visual QA lies

The render said clean, the file said broken: why AI visual QA lies

Comments
2 min read
Claude Formalized Fermat in 11 Days. The Math Isn't New.

Claude Formalized Fermat in 11 Days. The Math Isn't New.

Comments
3 min read
13 of 14 Models Write Messier Code Than the Human Who Fixed the Same Bug

13 of 14 Models Write Messier Code Than the Human Who Fixed the Same Bug

Comments
6 min read
Efficiency Hallucination: Every Model Rewrote Code That Couldn't Get Faster

Efficiency Hallucination: Every Model Rewrote Code That Couldn't Get Faster

Comments 1
7 min read
Laya is a 421M open-weights answer to Jev

Laya is a 421M open-weights answer to Jev

6
Comments 2
4 min read
GPT-6 Astra scores 95% on one robot task, 10% on another

GPT-6 Astra scores 95% on one robot task, 10% on another

6
Comments
4 min read
Benchmaxing: Winning the Exam Is Not Doing Better Work

Benchmaxing: Winning the Exam Is Not Doing Better Work

Comments
8 min read
Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.

Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.

1
Comments 1
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.