Skip to content
CCDV-F
Eval Testing and Debugging
hard
Question 5 of 5

A developer builds a RAG system and creates an eval suite with 100 test questions. They use Claude Opus 4.8 as an LLM judge to score answers on a 1-10 quality scale. After running the eval, they observe that 95% of scores are between 6 and 8, with very few scores at the extremes (1-3 or 9-10), even though some answers are clearly excellent and others are clearly poor. What is the most likely cause and how should the eval be improved?

A
B
C
D

This question requires Pro

Unlock all 5 questions in this certification

Written by certified professionals · Aligned to official exam objectives

View all Pro features

Educational Content — CertQnA practice questions are written against official exam objectives, covering the same domains tested on the real exam. All content is original and independent — not actual exam questions, not affiliated with any certification vendor. Learn more about our content policy