CCAR-P - Evaluation, Testing & Optimization (16% of the exam) - Section 4.5

Optimize token usage, latency, and cost-performance trade-offs.

Reducing cost and latency without breaking quality: caching, smaller models for simple steps, batching non-urgent work, trimming context and limiting output length. Candidates should pick the optimisation that addresses the stated constraint and verify quality holds afterwards.

prompt cachingmodel routingbatch processingcontext trimming

Practice question for this objective

Free sampleEvaluation, Testing & Optimizationmedium

A city council's planning enquiry service sends Claude's reply to residents as a single text message once generation has finished, so streaming partial output gives the resident nothing earlier. Replies average about 450 words and take around nine seconds to generate, against a stated budget of four seconds from question to message. A review of 300 replies, which the team keeps as a labelled eval set, found that the first two or three sentences of each reply already held the complete answer residents needed. Which change best fits the stated constraints?

  • AInstruct a short answer of two or three sentences, set a ceiling that fits it, and re-run the eval set Correct
  • BStream the reply token by token so residents see the start of the answer within the budget
  • CMove the service to the Message Batches API so that long replies no longer hold open a connection
  • DLower the output token ceiling sharply and leave the prompt as it is so replies are cut off early
Control output length through the prompt, with a ceiling as a backstop, when generation time breaks a latency budget, and verify quality afterwards. Total response time is dominated by how many tokens the model generates, so cutting a 450-word reply to two or three sentences removes most of the nine seconds. Instructing the shorter format lets the model write a complete short answer instead of being truncated, and the eval re-run provides evidence that the optimisation did not break correctness.

Why A is correct: Correct. Generation time scales with the number of output tokens, so asking for the short answer residents actually need brings the reply inside the budget, the ceiling acts as a backstop, and re-running the labelled set confirms the shorter replies remain correct.

Why B is wrong: Tempting because streaming is the standard fix for perceived latency in chat interfaces. It is wrong here because the reply is delivered as one text message after generation finishes, so streaming does not change when the resident receives anything.

Why C is wrong: Tempting because batch processing is a well-known way to handle long generations more cheaply. It is wrong because batch requests are processed asynchronously with no promise of a fast turnaround, which moves the service further from a four-second budget.

Why D is wrong: Tempting because a lower ceiling does stop generation sooner and is a one-line change. It is wrong because the model still plans a 450-word reply, so the ceiling cuts it mid-sentence; a hard stop is a compensating control where shaping the reply itself is the structural fix.

See more CCAR-P practice questions, answers explained.

Exam traps in Evaluation, Testing & Optimization

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAR-P bank for this domain.

  • Switch the weekly run to the smallest model tier, since the descriptions are short and formulaic.

    Why it is wrong: This is tempting because a smaller tier is cheaper per token and product copy looks like simple work. It is wrong as stated because the switch is made on an assumption rather than a measurement: the merchandising team requires reviewed-set scores to hold, and nothing here shows the smallest tier keeps brand voice and attribute accuracy on that set, whereas batching and caching cut cost without changing the model at all.

  • The new lines make the model write longer, more qualified replies for each region, and generation time and cost grow with that output

    Why it is wrong: This is tempting because longer output does raise both latency and cost, and an added instruction can change reply length. It is wrong because the stem states that output tokens per request match the baseline, and the cost increase is on the input side.

  • Replace earlier turns with a running summary after each reply so each request carries fewer tokens

    Why it is wrong: Tempting because summarising history is a common way to shrink long conversations. It is wrong because the audit rule requires the caller's exact earlier wording to stay available, and a summary paraphrases it away.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.