CCAR-P - Integration (19% of the exam) - Section 3.6

Apply retrieval strategies matched to data shape and query pattern.

Choosing how to retrieve given the data and the questions asked: semantic search, keyword search, hybrid search with reranking, or structured queries against tabular data. Candidates should recognise when vector search is the wrong tool, for example for exact identifiers or aggregations.

semantic searchkeyword searchhybrid retrieval and rerankingstructured queries

Practice question for this objective

Free sampleIntegrationhard

A national DIY retailer replaced the keyword index behind its in-store product assistant with a vector index built from the same catalogue, embedding each product as one short record as before. Accuracy on evaluation questions such as which drill suits masonry rose, but accuracy on questions naming a part number, such as a replacement filter for model HX-4410, fell from 94 percent to 61 percent. Retrieval traces for the failing questions show the wrong records, for example HX-4401 and HX-4140, already ranked first before generation begins, and the index was rebuilt from the current catalogue the night before the evaluation. What is the most likely cause of the part-number failures?

  • ADense embeddings capture meaning rather than exact character sequences, so codes differing by a character or two sit close together and lose exact matching Correct
  • BThe model is altering part numbers while drafting its answer, because its sampling settings let it substitute similar-looking product codes
  • CThe vector index is serving vectors from an older catalogue build, so newly listed part numbers are missing and the nearest older products are returned
  • DSemantic retrieval needs a larger top-k than keyword search did, so the correct product is retrieved but cut off before it reaches the model's context
Dense vector retrieval loses the exact-match signal that exact identifiers such as part numbers need, so those queries call for lexical or hybrid retrieval. Embedding models map text to points that reflect meaning, and alphanumeric codes that differ by one or two characters land very close together, so nearest-neighbour search cannot reliably separate HX-4410 from HX-4401. Keyword retrieval scores exact token matches, which is why part-number accuracy was high before the swap while descriptive queries, where meaning matters, improved after it. The usual remedy is hybrid retrieval that keeps a lexical signal for identifiers.

Why A is correct: Correct. The only change was swapping lexical for dense retrieval, and dense vectors place near-identical identifiers close together without rewarding an exact token match, which is precisely what the keyword index supplied for part-number queries.

Why B is wrong: Tempting because the visible symptom is a wrong code in the answer, and generation can corrupt identifiers. It is wrong because the traces show the wrong records already ranked first before generation begins, so the fault is upstream of the model.

Why C is wrong: Tempting because a stale index is a classic cause of confident wrong retrieval after a change. It is wrong because the stem states the index was rebuilt from the current catalogue the night before the evaluation, so freshness is ruled out.

Why D is wrong: Tempting because raising k is a common first lever when the right item is missing from context. It is wrong because the traces show the wrong products ranked first; a larger k would add more candidates without fixing a ranking signal that cannot tell near-identical codes apart.

See more CCAR-P practice questions, answers explained.

Exam traps in Integration

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAR-P bank for this domain.

  • Fine-tune the embedding model on pairs of error codes and their pages, so that rare identifiers land close to the right documents

    Why it is wrong: Tempting because domain-tuned embeddings do improve retrieval of specialist vocabulary. It is wrong here because new codes arrive weekly and the team has stated it cannot retrain on that cadence, so every new code would be poorly placed until the next tuning run.

  • Raise the retrieval depth from 40 to 400 rows and add a cross-encoder reranker so that more of the qualifying disputes reach the model

    Why it is wrong: This is tempting because the visible failure looks like missing rows, and deeper retrieval plus reranking does raise recall on text corpora. It is wrong because a ranked top-k list is still a similarity sample with no guarantee that it contains every row meeting the amount, date and region conditions, and the model is still doing the arithmetic, so the totals cannot be relied on to reconcile with month-end reporting.

  • The embedding model represents ward names poorly, so permits from the asked-about ward score below permits from neighbouring wards with similar names

    Why it is wrong: Tempting because weak handling of proper nouns is a known embedding weakness. It is wrong because the counts are capped at about 20 in total and small wards are counted correctly, which points to a ceiling on how many records are seen, not to ward confusion.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.