Skip to content
CCDV-F
Eval Testing and Debugging
hard
Question 4 of 5

A developer builds an eval suite to test a RAG pipeline that answers questions from a 500-page internal knowledge base. The eval suite includes 100 test questions covering all topics. After prompt changes, aggregate accuracy across all 100 questions remains 85% on both old and new prompts. However, user complaints spike about incorrect answers in a specific domain. What is the issue with the current eval approach, and what is the fix?

A
B
C
D

This question requires Pro

Unlock all 5 questions in this certification

Written by certified professionals · Aligned to official exam objectives

View all Pro features

Educational Content — CertQnA practice questions are written against official exam objectives, covering the same domains tested on the real exam. All content is original and independent — not actual exam questions, not affiliated with any certification vendor. Learn more about our content policy