DEV Community

#benchmarks

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
DeepSeek's Vision Flash Plays the Agent Game

DeepSeek's Vision Flash Plays the Agent Game

Comments
2 min read
Frontier Models Fail at Research-Level Reasoning

Frontier Models Fail at Research-Level Reasoning

Comments
2 min read
Frontier Models Hit a Wall on Research Thinking

Frontier Models Hit a Wall on Research Thinking

Comments
2 min read
Ranking Language Models by How Well They Spot Liars

Ranking Language Models by How Well They Spot Liars

Comments
9 min read
The Benchmarkpocalypse: Why AI Benchmarks Are Broken — and What Dan Luu Says We Should Do About It

The Benchmarkpocalypse: Why AI Benchmarks Are Broken — and What Dan Luu Says We Should Do About It

1
Comments
4 min read
Labs are ditching factual knowledge for reasoning speed

Labs are ditching factual knowledge for reasoning speed

Comments
2 min read
Why AI Benchmarks Mean Less Than You Think

Why AI Benchmarks Mean Less Than You Think

Comments
6 min read
An AI Capture-the-Flag Tournament: What the Scoreboard Counted

An AI Capture-the-Flag Tournament: What the Scoreboard Counted

Comments
6 min read
Twelve LLMs Played Werewolf. The Real Wolf Was the Thinking Knob.

Twelve LLMs Played Werewolf. The Real Wolf Was the Thinking Knob.

1
Comments
8 min read
Meta's Muse Code Clears 59% on Deep Software Engineering

Meta's Muse Code Clears 59% on Deep Software Engineering

Comments
2 min read
The Same GraphRAG Comparison Wins and Loses. It Depends Which Instrument Judged It.

The Same GraphRAG Comparison Wins and Loses. It Depends Which Instrument Judged It.

10
Comments 23
3 min read
Kimi K3 for Coding: Real-World Performance Tests and Benchmarks (2026)

Kimi K3 for Coding: Real-World Performance Tests and Benchmarks (2026)

Comments
9 min read
SynthDocBench: A New Benchmark for Long-Context Visual Document Understanding Reveals VLM Weaknesses

SynthDocBench: A New Benchmark for Long-Context Visual Document Understanding Reveals VLM Weaknesses

Comments
4 min read
UniClawBench: A New Benchmark for Proactive AI Agents in Real-World Scenarios

UniClawBench: A New Benchmark for Proactive AI Agents in Real-World Scenarios

Comments
3 min read
AI News Roundup: Grok 4.5 Hits Tesla, Perplexity's Orchestrator Beats Opus, and Meta Undercuts Pricing

AI News Roundup: Grok 4.5 Hits Tesla, Perplexity's Orchestrator Beats Opus, and Meta Undercuts Pricing

Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.