CCAR-P - Evaluation, Testing & Optimization (16% of the exam) - Section 4.4

Diagnose system issues (prompt failure, hallucinations, model mismatch).

Locating the cause of a failure: an ambiguous prompt, missing or wrong retrieved context, a hallucination, or a model mismatched to the task. The guide's sample item points at retrieval when answers go wrong after a document refresh with the model unchanged. Candidates should use what changed to narrow the cause.

prompt failureretrieval failurehallucinationmodel mismatchwhat changed

Practice question for this objective

Free sampleEvaluation, Testing & Optimizationhard

A wealth management firm's retrieval-based assistant answers adviser questions about fund fee schedules. Last month the ingestion job was changed to split documents into chunks half the previous size with no overlap, and since then the share of adviser-flagged wrong answers has risen from 2 to 9 per cent. The model, system prompt and latency are unchanged, and most wrong answers quote a real fee figure from the right fund but attach it to the wrong share class. Per-request token spend must not rise. What should the architect recommend?

  • AMove the assistant to a more capable model tier so that it can reason across fragmented chunks and infer which share class each fee belongs to
  • BRaise the number of retrieved chunks per question from five to twenty so that the share-class heading chunk is more likely to be returned alongside
  • CAdd a system prompt rule telling the model to confirm the share class named in a chunk before it quotes any fee figure from that chunk to an adviser
  • DReplay the flagged questions against the old and new indexes, and fix the splitter so each fee table row stays in a chunk with its share-class heading Correct
When accuracy drops after an ingestion change with the model unchanged, isolate the retrieval layer and fix how documents are chunked before changing the model or prompt. Using what changed narrows the search: the model, prompt and latency are constant, and only the chunking moved. Halving chunk size with no overlap separates table rows from the headings that give them meaning, so retrieval returns figures stripped of their share class and the model attaches them to the wrong one. Replaying the failing questions against both indexes proves this, and a splitter that keeps rows with their headings removes the cause without adding retrieved tokens.

Why A is wrong: This is tempting because a more capable tier is better at multi-step inference. It is wrong because the model did not change while the chunking did, and when the share-class heading is missing from the retrieved text no tier can reliably recover it; a larger tier also raises per-request spend, which the stem rules out.

Why B is wrong: This is tempting because it works at the retrieval layer and may sometimes pull the heading back in. It is wrong because it roughly quadruples the retrieved context per request, breaking the spend constraint, and even when the heading arrives the model must still guess which orphaned figure belongs under it.

Why C is wrong: This is tempting because it is cheap and aims directly at the symptom. It is wrong because the new chunks no longer contain the share-class heading, so there is nothing in the chunk for the model to confirm; an instruction cannot restore context that the ingestion change removed.

Why D is correct: The only change was the chunking, and the failure pattern (right figure, wrong share class) is what splitting a table away from its heading produces. Replaying the flagged questions against both indexes confirms the cause at the retrieval layer, and a structure-aware splitter fixes it without adding tokens to each request.

See more CCAR-P practice questions, answers explained.

Exam traps in Evaluation, Testing & Optimization

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAR-P bank for this domain.

  • Retrieval is returning outdated rule passages for the eligibility step, so the model reasons correctly over the wrong rules.

    Why it is wrong: Tempting because retrieval is the usual first suspect for wrong RAG answers. It is wrong because the request logs confirm the correct rule passages are in context for every audited case, which rules retrieval out.

  • The library holds superseded documents, so the assistant is quoting procedure numbers from obsolete versions that it retrieved.

    Why it is wrong: Tempting because superseded documents are a common source of wrong answers. It is wrong because the failing requests received no passages at all, and the cited numbers do not exist anywhere in the library, so nothing was quoted from a retrieved document.

  • The generation model's behaviour has drifted, so it now misreads older policy passages that it handled correctly before the switch.

    Why it is wrong: Tempting because blaming the model is a common reflex when answer quality drops. It is wrong because the generation model is unchanged, and model drift would not split the failures neatly by whether a page was edited after the index switch.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.