CCAR-P - Evaluation, Testing & Optimization (16% of the exam) - Section 4.6

Monitor system performance using logging and observability tools.

Running ongoing monitoring in production: logging requests, responses and tool calls, tracking quality and cost over time, and alerting on drift. Candidates should connect monitoring back to the evaluation metrics so a regression is caught from production data.

production loggingdrift detectionalertingcost tracking

Practice question for this objective

Free sampleEvaluation, Testing & Optimizationmedium

A national pensions agency uses Claude to draft replies to written enquiries, and a caseworker reviews and edits every draft before it is sent; the review tool already stores both the draft and the sent version. Release quality is judged on a 400-case regression set, re-run before each release, which has scored 92 percent on every release this year. The service lead sets two requirements: a fall in live draft quality must be detected within a week, and any failure pattern found in live traffic must be caught by the pre-release regression set from then on. No extra caseworker or reviewer hours are available for monitoring. Which two changes best meet these requirements? Select TWO.

  • AHave the model rate its own confidence in each draft, and alert when the weekly mean confidence score falls below baseline
  • BTrack the share of drafts that caseworkers heavily rewrite, by edit distance and enquiry type, and alert on a rise over baseline Correct
  • CHave a senior caseworker re-review a random sample of 200 sent replies each week and score each one against the rubric
  • DAdd heavily rewritten live drafts to the regression set each week, using the caseworker's sent reply as the reference answer Correct
  • ERe-run the 400-case regression set nightly against the production system, and alert when its score drops below 92 percent
Production monitoring should reuse signals the workflow already produces, such as reviewer edits, and feed live failures back into the regression set. Caseworker edits turn every draft into a labelled example: the distance between draft and sent reply measures how far the output fell short, so tracking the heavy-rewrite rate by enquiry type detects a live regression quickly at no extra cost. Detection alone does not stop a recurrence, so the rewritten cases, with the approved reply as reference, are added to the regression set, which then tests for that pattern before every release. Monitoring and evaluation become one loop.

Why A is wrong: This is tempting because it costs no reviewer time and produces a number per draft. It is wrong because a model's self-reported confidence is not a calibrated accuracy signal: a drafting regression can leave stated confidence unchanged, so the alert can stay silent while caseworkers rewrite more drafts.

Why B is correct: Correct. Every draft is already paired with the version a caseworker approved, so the size of each edit is a free, expert label on live output. A rising rewrite rate in any enquiry type surfaces a quality fall within days, with no added review hours.

Why C is wrong: This is tempting because human scoring on live replies is a trusted quality measure. It is wrong here because the stem rules out any extra caseworker or reviewer hours, and the existing edits already carry the same expert judgment at no added cost.

Why D is correct: Correct. Feeding live failures into the regression set, with the approved reply as the reference, is what makes the pre-release gate catch that failure pattern on every later release, which is the second stated requirement.

Why E is wrong: This is tempting because it automates a check the team already trusts. It is wrong because a fixed set re-run against an unchanged system returns the same score each night; it cannot see failures in enquiry types the set does not hold, and it adds nothing new to the set.

See more CCAR-P practice questions, answers explained.

Exam traps in Evaluation, Testing & Optimization

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAR-P bank for this domain.

  • Log the complete prompt for each request, including the nurse's typed question word for word, in the platform

    Why it is wrong: This is tempting because the full prompt gives the richest debugging record. It is wrong because nurses sometimes type patient names, so storing the question verbatim puts identifiers into a platform the privacy office has ruled out, which breaks a stated requirement and the principle of data minimisation.

  • A shift in incoming notes toward case types the assistant handles poorly, which lowered the real quality of its suggestions

    Why it is wrong: Input mix drift is the most common cause of a live quality drop and is a reasonable first guess. It is wrong here because the specialty mix is unchanged, and the canary set's outputs are stored and identical, so no change in incoming notes can explain why they now score lower.

  • Expand the offline eval set with new synthetic questions each night and alert when the nightly rerun falls below the launch score of 91 percent

    Why it is wrong: This is tempting because a larger eval set feels like better coverage and it reuses the existing nightly job. It is wrong because synthetic questions are invented by the team rather than drawn from what students actually ask, so the set still cannot see the shift in the live question mix that caused the complaints.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.