CCAR-P - Claude Models, Prompting & Context Engineering (13% of the exam) - Section 2.1

Select appropriate Claude models based on trade-offs.

Matching a model tier to a workload by weighing capability against latency and cost: the most capable tier for hard reasoning, faster and cheaper tiers for high-volume or latency-sensitive work, and routing between tiers within one system. Items test the trade-off, not the current model names.

capability versus latency versus costmodel tieringrouting between modelstask fit

Practice question for this objective

Free sampleClaude Models, Prompting & Context Engineeringmedium

A national records office answers freedom of information requests within a statutory deadline. For each request, a coordinator agent splits the request into search tasks, subagents each screen a batch of about 300 documents for relevance, and the coordinator then decides which statutory exemptions apply and drafts the reasoning. A trial on 150 closed requests found the fastest tier agreed with archivists on relevance screening 96 per cent of the time but chose the correct exemptions only 61 per cent of the time, while the most capable tier reached 93 per cent on exemptions. The office requires 90 per cent agreement on exemptions, and running every step on the most capable tier exceeds its cost ceiling per request. Which design best meets the requirement?

  • AFastest tier for every step, with an instruction asking the coordinator to reason with more care.
  • BMost capable tier for the screening subagents, fastest tier for the coordinator to keep drafts quick.
  • CMid tier for every step, so that cost and exemption accuracy both fall between the two results.
  • DMost capable tier for the coordinator's exemption decisions, fastest tier for screening subagents. Correct
Within one multi-agent system, route high-volume simple subtasks to a fast tier and reserve the most capable tier for the reasoning step that needs it. Model selection does not have to be a single choice per system. When a workflow mixes many simple, high-volume calls with a few calls that need complex judgement, assigning tiers per step lets each step meet its own measured requirement. Because document screening dominates the number of calls, moving it to the fastest tier is what makes room in the cost ceiling for the most capable tier on the exemption decisions.

Why A is wrong: This is tempting because it keeps cost at its lowest and prompt changes are cheap to try. It is wrong because a 61 per cent result on applying statutory exemptions is a capability gap on complex reasoning, not a missing instruction, and nothing in the trial suggests a wording change would close a 29-point shortfall against the requirement.

Why B is wrong: It can seem sensible to put the strongest model on the step that touches every document. It is wrong because it inverts the measured fit: screening already meets the bar on the fastest tier, while the exemption decisions that failed at 61 per cent stay on that tier, and paying the higher rate across the high-volume step raises cost rather than lowering it.

Why C is wrong: A uniform middle tier is a common compromise when two figures pull in opposite directions. It is wrong because the mid tier's exemption accuracy was never measured, so there is no evidence it clears the 90 per cent requirement, and it still pays more than necessary for screening work the fastest tier already handles.

Why D is correct: This is correct because it matches each tier to the work it was measured on: high-volume relevance screening, where the fastest tier already agrees with archivists 96 per cent of the time, and the low-volume legal reasoning on exemptions, where only the most capable tier clears the 90 per cent requirement. Most document-level calls move to the cheaper tier, which is what brings the request back under the cost ceiling.

See more CCAR-P practice questions, answers explained.

Exam traps in Claude Models, Prompting & Context Engineering

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAR-P bank for this domain.

  • The curriculum retrieval omits the worked methods for algebra word problems, so the fast tier answers them unguided.

    Why it is wrong: Missing source material is a frequent cause of wrong answers in a retrieval-backed assistant, so this is tempting. The audit found that the passages attached to the failing requests contain the correct method, which rules retrieval out.

  • The most capable tier without extended thinking, since a stronger tier needs no thinking budget.

    Why it is wrong: This is tempting because a stronger tier is often assumed to make extra reasoning unnecessary. It is wrong because the team's own evaluation shows that configuration at 87 per cent, below the 90 per cent floor, and the most capable tier exceeds the fixed annual budget in either configuration.

  • The retrieval index stopped returning statistics passages after the switch, so worked solutions lost their source material.

    Why it is wrong: Retrieval failures do produce confident wrong answers, which makes this tempting. It is wrong because the stem states the retrieval index is unchanged and the logistics questions that use the same index still score 96 per cent, so retrieval is not what moved.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.