DEV Community

#benchmarking

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
The Benchmark That's Half Traps — and Why That's Brilliant

The Benchmark That's Half Traps — and Why That's Brilliant

8
Comments
9 min read
A 4 GB Laptop GPU vs a 6-Core CPU on Gemma 4, Re-Measured in ABBA Order: 4.1x

A 4 GB Laptop GPU vs a 6-Core CPU on Gemma 4, Re-Measured in ABBA Order: 4.1x

9
Comments 2
11 min read
Counting bugs is the hard part of comparing AI review tools

Counting bugs is the hard part of comparing AI review tools

Comments
3 min read
When to use which: Dragonfly vs Redis vs Valkey

When to use which: Dragonfly vs Redis vs Valkey

1
Comments
4 min read
Our AI agents' "verified success" claims: 10 out of 10 failed independent recompute — including ours

Our AI agents' "verified success" claims: 10 out of 10 failed independent recompute — including ours

1
Comments
3 min read
Operational simplicity: the cross-key tax, quantified

Operational simplicity: the cross-key tax, quantified

Comments 1
3 min read
Memory efficiency: bytes per key

Memory efficiency: bytes per key

Comments 1
2 min read
Apple M2 Compute Performance: Benchmarking AMX, GPU, and Neural Engine

Apple M2 Compute Performance: Benchmarking AMX, GPU, and Neural Engine

Comments
6 min read
Latency under load: the same story as throughput, from the other side

Latency under load: the same story as throughput, from the other side

Comments 3
3 min read
Qwen 3.8 4-bit Benchmark RTX 4090: 1-bit is a Trap

Qwen 3.8 4-bit Benchmark RTX 4090: 1-bit is a Trap

Comments
7 min read
The sleep loop is the tell: agents that pay per action optimize to do nothing

The sleep loop is the tell: agents that pay per action optimize to do nothing

Comments 1
2 min read
Why a normal client gets 4.6M ops/s out of a Redis cluster that can do 40M

Why a normal client gets 4.6M ops/s out of a Redis cluster that can do 40M

Comments 3
5 min read
Our benchmark was leaking the answers to the model. The numbers looked fine the whole time.

Our benchmark was leaking the answers to the model. The numbers looked fine the whole time.

Comments
3 min read
We Measured the 200x Claim, and Got It Wrong Twice First

Real-world benchmark data beats vendor claims

We Measured the 200x Claim, and Got It Wrong Twice First

11
Comments 6
11 min read
Fairness Under the Microscope: Why HY3 Beats Nemotron 3 Ultra on lforla's Bias Stereotypes Audit

Fairness Under the Microscope: Why HY3 Beats Nemotron 3 Ultra on lforla's Bias Stereotypes Audit

1
Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.