Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
evals
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
When Building an AI Agent, the Journey Matters as Much as the Destination
Raj Kundalia
Raj Kundalia
Raj Kundalia
Follow
Aug 26
When Building an AI Agent, the Journey Matters as Much as the Destination
#
ai
#
evals
Comments
Add Comment
12 min read
What to actually measure when your agent "works"
Walker Miller
Walker Miller
Walker Miller
Follow
Aug 22
What to actually measure when your agent "works"
#
evals
#
testing
#
metrics
#
llmasjudge
Comments
Add Comment
9 min read
Evals are the unit tests of prompts
INTFRAME
INTFRAME
INTFRAME
Follow
Aug 15
Evals are the unit tests of prompts
#
ai
#
evals
Comments
Add Comment
2 min read
One pass of my eval bills $9.14 on the API and $0 through the CLI
Dylan Merigaud
Dylan Merigaud
Dylan Merigaud
Follow
Aug 12
One pass of my eval bills $9.14 on the API and $0 through the CLI
#
ai
#
evals
#
claude
#
tooling
Comments
Add Comment
2 min read
AI per developer: cosa accelera davvero (e cosa ti fa perdere tempo)
frontendfacile.it
frontendfacile.it
frontendfacile.it
Follow
Aug 7
AI per developer: cosa accelera davvero (e cosa ti fa perdere tempo)
#
workflowaicoding
#
agentillm
#
evals
#
reliabilitytesting
Comments
Add Comment
4 min read
I have been Vibecoding Evals (works better than I thought)
juan pablo hernández
juan pablo hernández
juan pablo hernández
Follow
Aug 3
I have been Vibecoding Evals (works better than I thought)
#
ai
#
llm
#
python
#
evals
Comments
Add Comment
3 min read
Rewriting prose until the tests pass: everything passed, but the check that mattered never ran once
matsumotory
matsumotory
matsumotory
Follow
Aug 4
Rewriting prose until the tests pass: everything passed, but the check that mattered never ran once
#
testing
#
aiwriting
#
evals
#
promptengineering
Comments
Add Comment
6 min read
Rewriting research prose until the tests pass
matsumotory
matsumotory
matsumotory
Follow
Aug 4
Rewriting research prose until the tests pass
#
testing
#
research
#
aiwriting
#
evals
Comments
Add Comment
5 min read
The 12-Prompt Eval I Run Before I Trust Any Model Upgrade
Agnel Nieves
Agnel Nieves
Agnel Nieves
Follow
for
Promptway
Jul 29
The 12-Prompt Eval I Run Before I Trust Any Model Upgrade
#
prompting
#
evals
#
modelmigration
#
claude
Comments
Add Comment
3 min read
How to Build AI Evals for Tool-Calling Agents
Dhanush Reddy
Dhanush Reddy
Dhanush Reddy
Follow
Aug 8
How to Build AI Evals for Tool-Calling Agents
#
ai
#
evals
#
aievals
#
agents
1
 reaction
Comments
2
 comments
17 min read
Predicting agent failure before you ship it
Walker Miller
Walker Miller
Walker Miller
Follow
Aug 14
Predicting agent failure before you ship it
#
failuremodes
#
testing
#
evals
#
reliability
Comments
Add Comment
6 min read
Loop drift: how agents convince themselves they're making progress
Walker Miller
Walker Miller
Walker Miller
Follow
Aug 13
Loop drift: how agents convince themselves they're making progress
#
failuremodes
#
evals
#
postmortem
#
loops
Comments
Add Comment
7 min read
Your LLM-as-judge is lying to you
Walker Miller
Walker Miller
Walker Miller
Follow
Aug 12
Your LLM-as-judge is lying to you
#
evals
#
llmasjudge
#
testing
#
bias
Comments
Add Comment
8 min read
Do not choose an AI model from a leaderboard alone
Edward Li
Edward Li
Edward Li
Follow
Jul 8
Do not choose an AI model from a leaderboard alone
#
ai
#
api
#
llm
#
evals
Comments
Add Comment
3 min read
Evaluating your evals: how to know the LLM judge is right
Walker Miller
Walker Miller
Walker Miller
Follow
Aug 10
Evaluating your evals: how to know the LLM judge is right
#
evals
#
llmasjudge
#
testing
#
metrics
Comments
Add Comment
5 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account