DEV Community

#benchmarking

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Our benchmark was leaking the answers to the model. The numbers looked fine the whole time.

Our benchmark was leaking the answers to the model. The numbers looked fine the whole time.

Comments
3 min read
Fairness Under the Microscope: Why HY3 Beats Nemotron 3 Ultra on lforla's Bias Stereotypes Audit

Fairness Under the Microscope: Why HY3 Beats Nemotron 3 Ultra on lforla's Bias Stereotypes Audit

1
Comments
3 min read
Three Gemma 4 Deployments on One T4G for Under $3: What the Runtime Changes, and What It Doesn't

Three Gemma 4 Deployments on One T4G for Under $3: What the Runtime Changes, and What It Doesn't

Comments
14 min read
I benchmarked 5 managed graph databases — and the "obvious" winner changed depending on what I measured

I benchmarked 5 managed graph databases — and the "obvious" winner changed depending on what I measured

Comments
5 min read
Three Gemma 4 Deployments on One T4G for Under $3: What the Runtime Changes, and What It Doesn't

Three Gemma 4 Deployments on One T4G for Under $3: What the Runtime Changes, and What It Doesn't

8
Comments 2
14 min read
The Model Reading My Benchmark Mattered More Than the Memory System Did

The Model Reading My Benchmark Mattered More Than the Memory System Did

Comments
7 min read
I benchmarked CognoDB against four other graph databases. The most interesting result had nothing to do with CognoDB.

I benchmarked CognoDB against four other graph databases. The most interesting result had nothing to do with CognoDB.

Comments 1
5 min read
I benchmarked 5 graph databases. The first four hours measured the Indian Ocean.

I benchmarked 5 graph databases. The first four hours measured the Indian Ocean.

Comments
6 min read
Last night we entered a memory benchmark against Tencent and Mem0. The score isn't back yet - and I'm publishing it either way.

Last night we entered a memory benchmark against Tencent and Mem0. The score isn't back yet - and I'm publishing it either way.

1
Comments
4 min read
Don't Trust a New Model's Benchmarks Until You Run Your Own 30-Minute Smoke Test

Don't Trust a New Model's Benchmarks Until You Run Your Own 30-Minute Smoke Test

Comments
5 min read
Before You Adopt MiniMax H3, Run a Twenty-Minute Model Audit

Before You Adopt MiniMax H3, Run a Twenty-Minute Model Audit

Comments
3 min read
DeepSeek Harness: What "Everything is a Plugin" Actually Means for Agent Frameworks

DeepSeek Harness: What "Everything is a Plugin" Actually Means for Agent Frameworks

Comments
2 min read
Free Model Endpoints Are an Evaluation Problem, Not a Hosting Strategy

Free Model Endpoints Are an Evaluation Problem, Not a Hosting Strategy

Comments
3 min read
New Benchmark for Evaluating Long-Horizon Agents in Online Environments

New Benchmark for Evaluating Long-Horizon Agents in Online Environments

Comments
4 min read
Native HTTP Engine for Node: Performance Benchmarks Against uWS, Bun, Fastify, and Hono

Native HTTP Engine for Node: Performance Benchmarks Against uWS, Bun, Fastify, and Hono

Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.