A developer builds an eval suite to test a RAG pipeline that answers questions from a 500-page internal knowledge base. The eval suite includes 100 test questions covering all topics. After prompt changes, aggregate accuracy across all 100 questions remains 85% on both old and new prompts. However, user complaints spike about incorrect answers in a specific domain. What is the issue with the current eval approach, and what is the fix?
This question requires Pro
Unlock all 5 questions in this certification
Written by certified professionals · Aligned to official exam objectives
View all Pro featuresMore Eval Testing and Debugging Questions
5 questions
Full Claude Certified Developer – Foundations Practice Test
All topics covered
All Claude Certified Developer – Foundations Questions
Browse by topic
Related Questions
A developer updates a system prompt to improve clarity and receives reports that Claude now produces...
A developer deploys a new system prompt and receives user feedback that the bot now hallucinated fac...
A developer uses Claude as an LLM judge to score outputs from their RAG pipeline on a 1-5 quality sc...
A developer builds a RAG system and creates an eval suite with 100 test questions. They use Claude O...
Educational Content — CertQnA practice questions are written against official exam objectives, covering the same domains tested on the real exam. All content is original and independent — not actual exam questions, not affiliated with any certification vendor. Learn more about our content policy