CCAR-P - Evaluation, Testing & Optimization (16% of the exam) - Section 4.3

Conduct A/B testing and iterative improvements.

Comparing a change against a baseline on the same evaluation set or live traffic split, changing one variable at a time, and keeping a change only when it improves the target metric without regressing others.

baseline comparisonone variable at a timeregression checksiteration

Practice question for this objective

Free sampleEvaluation, Testing & Optimizationmedium

A university runs a Claude-based course assistant for 22,000 students. Three teams change it on separate weekly schedules: one edits the system prompt, one rebuilds the course-material index, and one adjusts which model tier serves each request type. An automated grader scores a 2 percent sample of live conversations for faithfulness, and that score fell from 0.91 to 0.83 last month, while the offline evaluation set scored the same before and after every release. Nobody could say which change caused the drop. The provost's office requires that any future drop be traced to a specific change within one working day, and none of the teams may pause its release schedule. Which two changes best meet the requirement? Select TWO.

  • ABundle the three teams' changes into one combined monthly release, so the date of any drop identifies the cause
  • BRe-run the offline evaluation set after every release and compare its score with the previous release's result
  • CStamp every logged request with the prompt version, the index build ID and the model tier that served it Correct
  • DChart the live grader score split by each of those version fields, with each team's releases marked on the timeline Correct
  • ERaise the live grader's sample from 2 to 20 percent of conversations so that a drop becomes visible sooner
Tagging every production request with the versions that served it lets monitored quality be split by change, so a live regression can be attributed. When several components change on independent schedules, an aggregate quality trend cannot say which change moved it. Recording the prompt version, index build and model tier on each logged request, then breaking the live grader score down by those fields with releases marked, shows the version at which the score fell. This attributes the drop from production data alone, without freezing releases and without relying on an offline set that did not register the regression.

Why A is wrong: This is tempting because fewer release events seem easier to line up against a trend. It is wrong because it pauses the teams' weekly schedules, which the stem forbids, and bundling three changes into one release makes it harder, not easier, to say which of them caused a drop.

Why B is wrong: This is tempting because per-release evaluation is good practice. It is wrong here because the stem shows the offline set did not move while live quality fell, so it cannot see this kind of regression and cannot attribute a live drop to a change.

Why C is correct: Correct. Recording the exact configuration behind each logged request means every graded live conversation can be tied to the prompt version, index build and tier that produced it, which is the raw material for attribution.

Why D is correct: Correct. Splitting the monitored score by prompt version, index build and model tier, with releases annotated, shows which field's change lines up with the fall, so the cause can be named within a day while all three teams keep shipping.

Why E is wrong: This is tempting because a larger sample does detect a drop faster and with less noise. It is wrong because the requirement is attribution, not detection: a bigger sample with no version data still shows only that quality fell, not which change caused it.

See more CCAR-P practice questions, answers explained.

Exam traps in Evaluation, Testing & Optimization

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAR-P bank for this domain.

  • Promote the variant, since correctness is the target metric and the gain is significant at this sample size

    Why it is wrong: This is tempting because a five point gain on roughly 9,000 conversations per arm is very likely real. It is wrong because the variant regresses latency past a stated product requirement, and a gain on one metric does not license breaking another.

  • Serve the candidate on 5 percent of applications as a canary, with automatic rollback if underwriter overrides rise

    Why it is wrong: Tempting because a small canary with a rollback trigger is a sound way to limit risk in many releases. It is wrong because underwriters would see drafts from an unapproved configuration, which the compliance policy forbids regardless of traffic share.

  • Release the candidate since it improves on the 200 new cases, and reopen a corrected category if a complaint recurs

    Why it is wrong: This is tempting because the candidate has a measured gain and the release window is close. It is wrong because the new cases say nothing about the corrected categories, and waiting for complaints means discovering a regression in front of claimants, which breaks the agency's public commitment.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.