Skip to content
CCDV-F
Eval Testing and Debugging
hard
Question 3 of 5

A developer uses Claude as an LLM judge to score outputs from their RAG pipeline on a 1-5 quality scale. They notice the judge consistently assigns scores around 3.5 regardless of actual output quality — both excellent and poor outputs receive similar scores. What is the most likely cause and how should it be fixed?

A
B
C
D

This question requires Pro

Unlock all 5 questions in this certification

Written by certified professionals · Aligned to official exam objectives

View all Pro features

Educational Content — CertQnA practice questions are written against official exam objectives, covering the same domains tested on the real exam. All content is original and independent — not actual exam questions, not affiliated with any certification vendor. Learn more about our content policy