A developer builds a RAG system and creates an eval suite with 100 test questions. They use Claude Opus 4.8 as an LLM judge to score answers on a 1-10 quality scale. After running the eval, they observe that 95% of scores are between 6 and 8, with very few scores at the extremes (1-3 or 9-10), even though some answers are clearly excellent and others are clearly poor. What is the most likely cause and how should the eval be improved?
This question requires Pro
Unlock all 5 questions in this certification
Written by certified professionals · Aligned to official exam objectives
View all Pro featuresMore Eval Testing and Debugging Questions
5 questions
Full Claude Certified Developer – Foundations Practice Test
All topics covered
All Claude Certified Developer – Foundations Questions
Browse by topic
Related Questions
A developer updates a system prompt to improve clarity and receives reports that Claude now produces...
A developer deploys a new system prompt and receives user feedback that the bot now hallucinated fac...
A developer uses Claude as an LLM judge to score outputs from their RAG pipeline on a 1-5 quality sc...
A developer builds an eval suite to test a RAG pipeline that answers questions from a 500-page inter...
Educational Content — CertQnA practice questions are written against official exam objectives, covering the same domains tested on the real exam. All content is original and independent — not actual exam questions, not affiliated with any certification vendor. Learn more about our content policy