A developer uses Claude as an LLM judge to score outputs from their RAG pipeline on a 1-5 quality scale. They notice the judge consistently assigns scores around 3.5 regardless of actual output quality — both excellent and poor outputs receive similar scores. What is the most likely cause and how should it be fixed?
This question requires Pro
Unlock all 5 questions in this certification
Written by certified professionals · Aligned to official exam objectives
View all Pro featuresMore Eval Testing and Debugging Questions
5 questions
Full Claude Certified Developer – Foundations Practice Test
All topics covered
All Claude Certified Developer – Foundations Questions
Browse by topic
Related Questions
A developer updates a system prompt to improve clarity and receives reports that Claude now produces...
A developer deploys a new system prompt and receives user feedback that the bot now hallucinated fac...
A developer builds an eval suite to test a RAG pipeline that answers questions from a 500-page inter...
A developer builds a RAG system and creates an eval suite with 100 test questions. They use Claude O...
Educational Content — CertQnA practice questions are written against official exam objectives, covering the same domains tested on the real exam. All content is original and independent — not actual exam questions, not affiliated with any certification vendor. Learn more about our content policy