Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
llmasjudge
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Your LLM-as-judge is lying to you
Walker Miller
Walker Miller
Walker Miller
Follow
Aug 12
Your LLM-as-judge is lying to you
#
evals
#
llmasjudge
#
testing
#
bias
Comments
Add Comment
8 min read
Evaluating your evals: how to know the LLM judge is right
Walker Miller
Walker Miller
Walker Miller
Follow
Aug 10
Evaluating your evals: how to know the LLM judge is right
#
evals
#
llmasjudge
#
testing
#
metrics
Comments
Add Comment
5 min read
Reliable, and still wrong
Breach Protocol
Breach Protocol
Breach Protocol
Follow
Jul 1
Reliable, and still wrong
#
evaluation
#
llmasjudge
#
benchmarks
Comments
Add Comment
3 min read
Beyond Scores: A Critical Review of Benchmark Reports for Evaluating Large Language Models
Ismail zamareh
Ismail zamareh
Ismail zamareh
Follow
May 17
Beyond Scores: A Critical Review of Benchmark Reports for Evaluating Large Language Models
#
llmevaluation
#
benchmarkcontamination
#
reproducibility
#
llmasjudge
Comments
Add Comment
7 min read
Build a Production RAG System on AWS Bedrock from Scratch
Joyson Fernandes
Joyson Fernandes
Joyson Fernandes
Follow
May 31
Build a Production RAG System on AWS Bedrock from Scratch
#
llmevaluation
#
llmasjudge
#
apigateway
#
bedrock
1
 reaction
Comments
Add Comment
29 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account