Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
benchmark
Follow
Hide
Posts
Left menu
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Why I’m Building a New AI Memory Benchmark (And Why the Existing Ones Fall Short)
woochan
woochan
woochan
Follow
Sep 7
Why I’m Building a New AI Memory Benchmark (And Why the Existing Ones Fall Short)
#
ai
#
benchmark
#
startup
#
opensource
1
reaction
Comments
Add Comment
2 min read
Elysia 2 vs NestJS 12: Runtime +64.6%, Framework +10.6%
Davron Yuldashev
Davron Yuldashev
Davron Yuldashev
Follow
Sep 7
Elysia 2 vs NestJS 12: Runtime +64.6%, Framework +10.6%
#
benchmark
#
bunjs
#
elysia
#
performance
Comments
Add Comment
14 min read
I Tested Q4_K_M vs MXFP4 on the Same Laptop — The Supposedly-Faster New Format Lost
Pitambar Mahato
Pitambar Mahato
Pitambar Mahato
Follow
Sep 5
I Tested Q4_K_M vs MXFP4 on the Same Laptop — The Supposedly-Faster New Format Lost
#
llm
#
benchmark
#
performance
#
ai
Comments
Add Comment
6 min read
ช่องว่าง 0.3% แต่ราคาต่าง 2 เท่า, อ่านตาราง Terminal-Bench 4.0 ให้เป็น
Nokka
Nokka
Nokka
Follow
Sep 5
ช่องว่าง 0.3% แต่ราคาต่าง 2 เท่า, อ่านตาราง Terminal-Bench 4.0 ให้เป็น
#
ai
#
benchmark
#
llm
#
programming
Comments
Add Comment
2 min read
เมื่อ Benchmark โกหกคุณ, SWE-Bench ProMax กับคะแนนจริงที่โมเดลเก่งสุดทำได้แค่ 41.2%
Nokka
Nokka
Nokka
Follow
Sep 5
เมื่อ Benchmark โกหกคุณ, SWE-Bench ProMax กับคะแนนจริงที่โมเดลเก่งสุดทำได้แค่ 41.2%
#
ai
#
benchmark
#
programming
#
machinelearning
Comments
Add Comment
2 min read
Benchmarking Real-Time Voice AI APIs: Cartesia vs Deepgram vs ElevenLabs (2026)
mrzitoun
mrzitoun
mrzitoun
Follow
Sep 3
Benchmarking Real-Time Voice AI APIs: Cartesia vs Deepgram vs ElevenLabs (2026)
#
ai
#
webdev
#
voice
#
benchmark
Comments
Add Comment
1 min read
I built a benchmark for AI companion apps because every "best AI girlfriend" list is affiliate spam
Maya Stone
Maya Stone
Maya Stone
Follow
Sep 3
I built a benchmark for AI companion apps because every "best AI girlfriend" list is affiliate spam
#
ai
#
llm
#
benchmark
Comments
1
comment
5 min read
A Benchmark Is Only as Honest as Its Harness
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 2
A Benchmark Is Only as Honest as Its Harness
#
ai
#
benchmark
#
llm
#
testing
Comments
Add Comment
4 min read
Extraction APIs invented values for 17% of the fields that aren't in the document. Ours included.
Velrim
Velrim
Velrim
Follow
Sep 2
Extraction APIs invented values for 17% of the fields that aren't in the document. Ours included.
#
ai
#
machinelearning
#
benchmark
#
llm
Comments
Add Comment
12 min read
A Small Transformer Trained in 1.5 Hours Beat Many LLMs on ARC
Anaz S. Aji
Anaz S. Aji
Anaz S. Aji
Follow
for
Codecora Dev
Sep 2
A Small Transformer Trained in 1.5 Hours Beat Many LLMs on ARC
#
machinelearning
#
vectorsearch
#
quantization
#
benchmark
Comments
Add Comment
5 min read
I Ran 3 Open-Weight LLMs Head-to-Head on a 24GB Mac — One Was 3x Faster
Pitambar Mahato
Pitambar Mahato
Pitambar Mahato
Follow
Sep 4
I Ran 3 Open-Weight LLMs Head-to-Head on a 24GB Mac — One Was 3x Faster
#
llm
#
benchmark
#
opensource
#
ai
2
reactions
Comments
4
comments
5 min read
Benchmark a Free AI Coding Tier on a Cold Server
Avery Wang
Avery Wang
Avery Wang
Follow
Aug 30
Benchmark a Free AI Coding Tier on a Cold Server
#
ai
#
benchmark
#
opensource
#
testing
Comments
Add Comment
4 min read
AxonASP vs. Native IIS ASP: Performance Benchmarks and Engine Architecture
Lucas Guimarães
Lucas Guimarães
Lucas Guimarães
Follow
Aug 29
AxonASP vs. Native IIS ASP: Performance Benchmarks and Engine Architecture
#
iis
#
asp
#
vbscript
#
benchmark
Comments
Add Comment
2 min read
We ran 160 agent tasks across two frameworks. The frameworks tied. Then we changed the model.
benchclawio
benchclawio
benchclawio
Follow
for
benchclaw
Aug 29
We ran 160 agent tasks across two frameworks. The frameworks tied. Then we changed the model.
#
python
#
ai
#
testing
#
benchmark
Comments
Add Comment
4 min read
A LongMemEval-S number you can reproduce
Przemek Marzec
Przemek Marzec
Przemek Marzec
Follow
for
Sovantica
Aug 27
A LongMemEval-S number you can reproduce
#
ai
#
agents
#
benchmark
#
python
Comments
4
comments
6 min read
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account