Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
benchmarking
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Our benchmark was leaking the answers to the model. The numbers looked fine the whole time.
AIOil Security Shield
AIOil Security Shield
AIOil Security Shield
Follow
Sep 2
Our benchmark was leaking the answers to the model. The numbers looked fine the whole time.
#
security
#
ai
#
benchmarking
#
opensource
Comments
Add Comment
3 min read
Fairness Under the Microscope: Why HY3 Beats Nemotron 3 Ultra on lforla's Bias Stereotypes Audit
RESK
RESK
RESK
Follow
Sep 1
Fairness Under the Microscope: Why HY3 Beats Nemotron 3 Ultra on lforla's Bias Stereotypes Audit
#
ai
#
llm
#
fairness
#
benchmarking
1
 reaction
Comments
Add Comment
3 min read
Three Gemma 4 Deployments on One T4G for Under $3: What the Runtime Changes, and What It Doesn't
xbill
xbill
xbill
Follow
for
AWS Community Builders
Sep 1
Three Gemma 4 Deployments on One T4G for Under $3: What the Runtime Changes, and What It Doesn't
#
aws
#
machinelearning
#
benchmarking
#
python
Comments
Add Comment
14 min read
I benchmarked 5 managed graph databases — and the "obvious" winner changed depending on what I measured
Ahmed Amer
Ahmed Amer
Ahmed Amer
Follow
Aug 27
I benchmarked 5 managed graph databases — and the "obvious" winner changed depending on what I measured
#
database
#
graphdatabase
#
benchmarking
Comments
Add Comment
5 min read
Three Gemma 4 Deployments on One T4G for Under $3: What the Runtime Changes, and What It Doesn't
xbill
xbill
xbill
Follow
for
Google Developer Experts
Sep 1
Three Gemma 4 Deployments on One T4G for Under $3: What the Runtime Changes, and What It Doesn't
#
aws
#
machinelearning
#
benchmarking
#
python
8
 reactions
Comments
2
 comments
14 min read
The Model Reading My Benchmark Mattered More Than the Memory System Did
Pranab Sarkar
Pranab Sarkar
Pranab Sarkar
Follow
Aug 25
The Model Reading My Benchmark Mattered More Than the Memory System Did
#
ai
#
llm
#
benchmarking
#
opensource
Comments
Add Comment
7 min read
I benchmarked CognoDB against four other graph databases. The most interesting result had nothing to do with CognoDB.
Sachin
Sachin
Sachin
Follow
Aug 21
I benchmarked CognoDB against four other graph databases. The most interesting result had nothing to do with CognoDB.
#
webdev
#
database
#
devops
#
benchmarking
Comments
1
 comment
5 min read
I benchmarked 5 graph databases. The first four hours measured the Indian Ocean.
Abhinav Bahuguna
Abhinav Bahuguna
Abhinav Bahuguna
Follow
Aug 21
I benchmarked 5 graph databases. The first four hours measured the Indian Ocean.
#
showdev
#
database
#
benchmarking
#
devops
Comments
Add Comment
6 min read
Last night we entered a memory benchmark against Tencent and Mem0. The score isn't back yet - and I'm publishing it either way.
Bryan Williams
Bryan Williams
Bryan Williams
Follow
Aug 23
Last night we entered a memory benchmark against Tencent and Mem0. The score isn't back yet - and I'm publishing it either way.
#
showdev
#
ai
#
buildinpublic
#
benchmarking
1
 reaction
Comments
Add Comment
4 min read
Don't Trust a New Model's Benchmarks Until You Run Your Own 30-Minute Smoke Test
Dakota Wu
Dakota Wu
Dakota Wu
Follow
Aug 17
Don't Trust a New Model's Benchmarks Until You Run Your Own 30-Minute Smoke Test
#
ai
#
programming
#
benchmarking
#
opensource
Comments
Add Comment
5 min read
Before You Adopt MiniMax H3, Run a Twenty-Minute Model Audit
Casey Li
Casey Li
Casey Li
Follow
Aug 14
Before You Adopt MiniMax H3, Run a Twenty-Minute Model Audit
#
ai
#
opensource
#
programming
#
benchmarking
Comments
Add Comment
3 min read
DeepSeek Harness: What "Everything is a Plugin" Actually Means for Agent Frameworks
Cole Halton
Cole Halton
Cole Halton
Follow
Aug 13
DeepSeek Harness: What "Everything is a Plugin" Actually Means for Agent Frameworks
#
deepseek
#
aiagents
#
opensource
#
benchmarking
Comments
Add Comment
2 min read
Free Model Endpoints Are an Evaluation Problem, Not a Hosting Strategy
Riley Xu
Riley Xu
Riley Xu
Follow
Aug 14
Free Model Endpoints Are an Evaluation Problem, Not a Hosting Strategy
#
ai
#
opensource
#
benchmarking
#
python
Comments
Add Comment
3 min read
New Benchmark for Evaluating Long-Horizon Agents in Online Environments
David DĂaz
David DĂaz
David DĂaz
Follow
Aug 12
New Benchmark for Evaluating Long-Horizon Agents in Online Environments
#
ai
#
benchmarking
#
agents
#
onlineservices
Comments
Add Comment
4 min read
Native HTTP Engine for Node: Performance Benchmarks Against uWS, Bun, Fastify, and Hono
Raiyan C
Raiyan C
Raiyan C
Follow
Aug 10
Native HTTP Engine for Node: Performance Benchmarks Against uWS, Bun, Fastify, and Hono
#
node
#
performance
#
http
#
benchmarking
Comments
Add Comment
3 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account