Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
← All Trends
Custom LLM Eval Decks Over Hype
59 posts in this trend in the last 7 days
•
Active about 2 hours ago
Pick Your LLM With a Scoreboard, Not a Hunch: A Two-File Eval That Costs Nothing to Re-Run
Riley Wu
Riley Wu
Riley Wu
Follow
Aug 10
Pick Your LLM With a Scoreboard, Not a Hunch: A Two-File Eval That Costs Nothing to Re-Run
#
llm
#
testing
#
python
#
ai
Comments
Add Comment
6 min read
Stop Asking Coding Models to Write Code. Test Whether They Can Review a Patch
Finley Zhou
Finley Zhou
Finley Zhou
Follow
Aug 13
Stop Asking Coding Models to Write Code. Test Whether They Can Review a Patch
#
ai
#
python
#
testing
#
codequality
Comments
Add Comment
5 min read
Stop Arguing About Which Model Is Best. Build a Two-Tier Habit Instead.
Avery Wang
Avery Wang
Avery Wang
Follow
Aug 13
Stop Arguing About Which Model Is Best. Build a Two-Tier Habit Instead.
#
ai
#
productivity
#
programming
#
tooling
Comments
Add Comment
5 min read
Weekly 'Game-Changer' Models Burned Me Twice. Now They Earn Production Access Through Gates.
Finley Zhou
Finley Zhou
Finley Zhou
Follow
Aug 13
Weekly 'Game-Changer' Models Burned Me Twice. Now They Earn Production Access Through Gates.
#
ai
#
productivity
#
llm
#
programming
Comments
1
comment
6 min read
Make AI Patches Prove Themselves Before They Touch Main
Taylor Wang
Taylor Wang
Taylor Wang
Follow
Aug 12
Make AI Patches Prove Themselves Before They Touch Main
#
ai
#
programming
#
testing
#
productivity
Comments
Add Comment
7 min read
« First
‹ Prev
1
2
3
4
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account