DEV Community

#benchmarks

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
We hit 99.95% on the LoCoMo memory benchmark. Here's the catch, and why it still matters.

We hit 99.95% on the LoCoMo memory benchmark. Here's the catch, and why it still matters.

5
Comments 1
4 min read
Twelve LLMs Played Werewolf. The Real Wolf Was the Thinking Knob.

Twelve LLMs Played Werewolf. The Real Wolf Was the Thinking Knob.

1
Comments
8 min read
Meta's Muse Code Clears 59% on Deep Software Engineering

Meta's Muse Code Clears 59% on Deep Software Engineering

Comments
2 min read
Kimi K3 for Coding: Real-World Performance Tests and Benchmarks (2026)

Kimi K3 for Coding: Real-World Performance Tests and Benchmarks (2026)

Comments
9 min read
SynthDocBench: A New Benchmark for Long-Context Visual Document Understanding Reveals VLM Weaknesses

SynthDocBench: A New Benchmark for Long-Context Visual Document Understanding Reveals VLM Weaknesses

Comments
4 min read
UniClawBench: A New Benchmark for Proactive AI Agents in Real-World Scenarios

UniClawBench: A New Benchmark for Proactive AI Agents in Real-World Scenarios

Comments
3 min read
AI News Roundup: Grok 4.5 Hits Tesla, Perplexity's Orchestrator Beats Opus, and Meta Undercuts Pricing

AI News Roundup: Grok 4.5 Hits Tesla, Perplexity's Orchestrator Beats Opus, and Meta Undercuts Pricing

Comments
2 min read
Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.

Half the answer keys in text-to-SQL benchmarks are wrong. So I generated the database from the answer key.

Comments
7 min read
Agent Leaderboards Measure Score. We Added Price.

Agent Leaderboards Measure Score. We Added Price.

Comments
5 min read
How I tried to write an article about slow Chinese LLMs

How I tried to write an article about slow Chinese LLMs

17
Comments 16
10 min read
Turn the camera away, and the AI's world freezes

Turn the camera away, and the AI's world freezes

Comments
3 min read
Reliable, and still wrong

Reliable, and still wrong

Comments
3 min read
Put AI agents in charge of a Civilization game and they reach for the nukes

Put AI agents in charge of a Civilization game and they reach for the nukes

Comments
3 min read
Claude Fable 5 Scores 95% on SWE-bench, Then Hands Off to Opus 4.8

Claude Fable 5 Scores 95% on SWE-bench, Then Hands Off to Opus 4.8

Comments
3 min read
An LLM benchmark is only useful for as long as it's hard

An LLM benchmark is only useful for as long as it's hard

2
Comments
10 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.