CCAR-P - Claude Models, Prompting & Context Engineering (13% of the exam) - Section 2.4

Optimize context windows and manage token usage.

Deciding what belongs in the context window: trimming irrelevant material, summarising long histories, retrieving only what a step needs, and keeping critical instructions where the model will attend to them. Candidates should treat context as a budget with both a cost and a quality impact.

context budgethistory summarisationselective retrievaltoken usage

Practice question for this objective

Free sampleClaude Models, Prompting & Context Engineeringmedium

A state licensing agency runs an assistant that answers caseworker questions by placing up to 25 retrieved regulation excerpts into each request. The system prompt sets two binding rules: cite the excerpt behind every statement, and decline to give a determination on any individual application. The prompt is currently ordered as rules first, then the caseworker's question, then the excerpts. On 200 held-out questions, rule adherence is 97 per cent when fewer than five excerpts are retrieved and 81 per cent when more than twenty are, while retrieval precision is unchanged. Some answers genuinely need every excerpt, and the agency must stay within its current per-request budget. Which change should the architect recommend?

  • APut the excerpts first in the user turn, then restate both binding rules and the question at the end Correct
  • BMove the assistant to the most capable model tier so it can hold the rules across longer prompts
  • CRestate the two rules in capital letters at the very start of the system prompt to raise their weight
  • DCap retrieval at five excerpts per request so every prompt stays inside the high-adherence range
With long retrieved material in context, put the documents first and critical instructions and the question last, where the model attends to them. Adherence falls only as the number of excerpts grows while retrieval quality stays constant, which points at prompt structure rather than at retrieval or model capability. Placing long-form material at the top of the user turn and restating the binding rules with the query at the end keeps the instructions next to the point where the model begins generating, which improves adherence in long contexts. The change keeps every needed excerpt and adds negligible cost, whereas a larger model raises spend and capping retrieval removes context some answers require.

Why A is correct: This is correct because with long material in context, models attend more reliably to instructions and the query placed after the documents, close to where generation starts. It keeps every excerpt, adds only the few tokens of the restated rules, so it fits the budget while addressing the length-linked drop in adherence.

Why B is wrong: This is tempting because a more capable model often follows instructions more reliably in long prompts. It is wrong because it raises per-request cost against a fixed budget and treats a prompt structure problem as a capability problem; the evidence points at where the rules sit relative to a long block of excerpts.

Why C is wrong: This is tempting because emphasis is a common first reaction to ignored instructions. It is wrong because it leaves the rules at the far end of the prompt from where the model generates its answer, with the long excerpt block still sitting between them, so the positional cause of the drop is untouched.

Why D is wrong: This is tempting because the data shows adherence is high with few excerpts. It is wrong because the stem states some answers need every excerpt, so capping retrieval trades away answer completeness, which is truncating required context to protect a metric.

See more CCAR-P practice questions, answers explained.

Exam traps in Claude Models, Prompting & Context Engineering

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAR-P bank for this domain.

  • Widen the kept window to the most recent 40 turns so early details stay in context for most chats

    Why it is wrong: This is tempting because it directly keeps more of the early conversation and needs almost no engineering. It is wrong because it raises the tokens sent on every turn against a fixed cost ceiling, and chats that run past 40 turns still lose the order number and address, so it moves the failure rather than removing it.

  • The model is inventing diagnoses at random, so a lower sampling temperature is needed to restore summary accuracy

    Why it is wrong: Fabricated clinical content is a serious concern, so hallucination is a tempting diagnosis. It is wrong because the flagged problems are real entries from the patient's own history, not invented ones, and the errors are systematic rather than random; temperature was also unchanged.

  • Summarise the accumulated transcript before each stage so later stages receive a shorter version of it

    Why it is wrong: This is tempting because it shrinks what each later stage reads without redesigning the handoffs. It is wrong because a free-text summary can paraphrase or drop exact allergen and nutrition values, which the claims check must work from, and it still carries material the later stages never use.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.