DEV Community

Casey Zhang profile picture

Casey Zhang

Building things with Python, Go, and JavaScript. Love automating everything.

Location Singapore Joined Joined on 
Refuse Every Agent Score That Ships Without a Protocol Card

Refuse Every Agent Score That Ships Without a Protocol Card

Comments
8 min read
Your Coding-Agent Score Needs a Dummy Baseline and a Size Split

Your Coding-Agent Score Needs a Dummy Baseline and a Size Split

Comments
7 min read
Don't Ship the Pass Rate. Ship the Failure Mix.

Don't Ship the Pass Rate. Ship the Failure Mix.

Comments
7 min read
Freeze the Retry Budget Before You Trust an Agent Pass Rate

Freeze the Retry Budget Before You Trust an Agent Pass Rate

Comments
8 min read
Build an LLM Cache Proxy in 40 Lines to Cut Token Costs 83%

Build an LLM Cache Proxy in 40 Lines to Cut Token Costs 83%

Comments
6 min read
Your Coding Agent's Pass Rate Isn't a Benchmark Until the Oracle Is Frozen

Your Coding Agent's Pass Rate Isn't a Benchmark Until the Oracle Is Frozen

Comments
7 min read
Benchmarking an AI Coding Tool? Start With the Data You Actually Trust

Benchmarking an AI Coding Tool? Start With the Data You Actually Trust

Comments
5 min read
Your Prompt A/B Test Is Probably Random Noise: A Paired-Test Harness for Free-Tier LLMs

Your Prompt A/B Test Is Probably Random Noise: A Paired-Test Harness for Free-Tier LLMs

Comments
5 min read
Benchmarking AI Code Generators: A Reproducible Method You Can Run on a Free Server

Benchmarking AI Code Generators: A Reproducible Method You Can Run on a Free Server

Comments
5 min read
Cold Starts Will Eat Your Free LLM Tier: A Reproducible Benchmark

Cold Starts Will Eat Your Free LLM Tier: A Reproducible Benchmark

Comments
4 min read
The Token Allowance Trap: A Task-Price Benchmark Before You Build on Free LLM APIs

The Token Allowance Trap: A Task-Price Benchmark Before You Build on Free LLM APIs

Comments
4 min read
Measure Your AI Reviewer Before You Trust It: A Seeded-Defect Benchmark

Measure Your AI Reviewer Before You Trust It: A Seeded-Defect Benchmark

Comments
5 min read
Benchmarking LLM Agents Without the Marketing Math: Dataset, Metrics, and Controls

Benchmarking LLM Agents Without the Marketing Math: Dataset, Metrics, and Controls

Comments
5 min read
From git log to Release Notes: A Free-Tier Agent Case Study

From git log to Release Notes: A Free-Tier Agent Case Study

Comments
5 min read
Crash-Proof Batch LLM Processing: A SQLite Job Queue on a Free Server

Crash-Proof Batch LLM Processing: A SQLite Job Queue on a Free Server

Comments 1
5 min read
Building a Free-Tier GitHub Issue Triage Agent: A Case Study

Building a Free-Tier GitHub Issue Triage Agent: A Case Study

Comments
5 min read
Cheap Model Hype Is Not a Benchmark: Run a 30-Minute Agent Gate Before You Swap

Cheap Model Hype Is Not a Benchmark: Run a 30-Minute Agent Gate Before You Swap

Comments
4 min read
Cheap Model Hype Is Not a Benchmark: Run a 30-Minute Agent Gate Before You Swap

Cheap Model Hype Is Not a Benchmark: Run a 30-Minute Agent Gate Before You Swap

Comments
4 min read
Tool-Permission Gate: The Only MiniMax H3 Test That Matters

Tool-Permission Gate: The Only MiniMax H3 Test That Matters

Comments
5 min read
Pin Your Agent's Behavior: Snapshot Tests for Tool-Call Sequences When You Swap Models

Pin Your Agent's Behavior: Snapshot Tests for Tool-Call Sequences When You Swap Models

Comments
5 min read
AI Agent Boundary Testing: Fake-Tool Harness on a Free Server

AI Agent Boundary Testing: Fake-Tool Harness on a Free Server

Comments
6 min read
I Gave My Coding Agent Fake Tools on Purpose: A Boundary-Test Harness You Can Run on a Free Server

I Gave My Coding Agent Fake Tools on Purpose: A Boundary-Test Harness You Can Run on a Free Server

Comments
5 min read
Stop Guessing: A Reproducible Harness for Evaluating Free AI Coding Models on Your Own Repo

Stop Guessing: A Reproducible Harness for Evaluating Free AI Coding Models on Your Own Repo

Comments
5 min read
loading...