CCAR-P - Evaluation, Testing & Optimization (16% of the exam) - Section 4.1

Define evaluation metrics (accuracy, latency, cost, safety, security).

Deciding what to measure before changing anything: task accuracy, latency, cost per request, and safety and security behaviour. Candidates should tie each metric to a requirement and recognise when a single headline metric hides a failure in another dimension.

accuracy metricslatency and cost metricssafety and security metricsmetrics tied to requirements

Practice question for this objective

Free sampleEvaluation, Testing & Optimizationmedium

A mortgage lender uses Claude to extract and check figures from application packs, and finance has set a ceiling on cost per completed application, where completed means the pack passed every check without manual rework. To cut spend, the team moved the extraction step to a smaller model tier. Cost per request fell by 38 per cent, but the automatic retry rate rose from 4 to 19 per cent and more packs now go to an underwriter for manual correction. Which metric should decide whether the team keeps the change?

  • ACost per request, since a 38 per cent fall on each call carries through to the monthly bill
  • BTotal model and rework cost per completed application, including retries and underwriter time Correct
  • CAverage input and output tokens per request, compared before and after the model change
  • DExtraction accuracy on the first attempt, against the previous tier on a labelled set of packs
Measure cost in the unit the requirement uses, such as cost per completed task including retries and human rework, not cost per model request. A cheaper model can lower the price of each call while raising the number of calls and the human effort needed to finish the work. Because finance set the ceiling per completed application, the deciding metric must sum every model request, every retry and the underwriter's correction time for each application that finishes. Cost per request hides exactly the retries and rework that the change introduced.

Why A is wrong: This is tempting because the fall is real and directly measured. It is wrong because retries multiply the number of requests per application and underwriter rework is not counted at all, while the ceiling finance set is per completed application.

Why B is correct: This matches the unit finance set the ceiling on. It folds in the extra requests from retries and the cost of underwriter correction, so it shows whether the cheaper tier lowers or raises what each finished application actually costs.

Why C is wrong: Token counts are tempting because they drive the model bill. It is wrong because they are a per-request proxy that ignores the retry rate and the cost of manual correction, so they cannot show whether the ceiling per completed application is met.

Why D is wrong: Accuracy is tempting because the retries and rework stem from extraction errors. It is wrong because an accuracy figure alone does not answer a cost ceiling; the decision turns on what the errors cost per completed application, which accuracy does not express.

See more CCAR-P practice questions, answers explained.

Exam traps in Evaluation, Testing & Optimization

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAR-P bank for this domain.

  • A larger benign question set of 2,000 items, so the accuracy figure has a tighter confidence interval

    Why it is wrong: This is tempting because more items give a statistically firmer number. It is wrong because a bigger benign set still contains no injection attempts, so the security requirement remains entirely unmeasured however precise the accuracy figure becomes.

  • The longer prompt has pushed the static prefix out of the cache, so every request now pays full input processing before any output is generated.

    Why it is wrong: A prompt edit that breaks caching is a familiar cause of rising cost, so this is tempting. It is wrong because the stem states the cache hit rate and input tokens are unchanged, which rules out a change on the input side.

  • Mean model response time from the orchestration logs, alerting if it rises above 1.0 seconds

    Why it is wrong: This is tempting because the figure is already collected and looks comfortably inside budget. It is wrong because a mean hides the slow tail the 99 per cent requirement is about, and the orchestration logs exclude the lookup calls and network time that the gateway experiences.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.