CCAR-P - Evaluation, Testing & Optimization (16% of the exam) - Section 4.2

Design evaluation datasets and test frameworks using mixed methodologies.

Building evaluation sets that represent real traffic and its edge cases, and combining code-based checks, model-graded evaluation and human review. Candidates should know where each method is reliable and why a model grader needs its own validation against human judgment.

representative evaluation setsedge casescode-based gradingmodel-graded evaluationhuman review

Practice question for this objective

Free sampleEvaluation, Testing & Optimizationhard

A hospital group is building an evaluation set for a Claude assistant that summarises discharge notes. The first version holds 150 notes the team wrote by hand and scores 97 percent, yet clinicians report errors on live notes with dictation artefacts, mixed medication lists and unusual abbreviations. Live traffic is 70 percent general medicine, with paediatric and oncology notes making up most of the rest. A business associate agreement is in place, and the privacy office permits only de-identified notes for testing. Which approach best produces a representative evaluation set?

  • AAsk Claude to generate 2,000 synthetic discharge notes across every specialty, then have clinicians spot-check a tenth of them for realism
  • BSample de-identified live notes in proportion to the specialty mix, add an oversampled slice of known hard cases, and report that slice apart Correct
  • CDraw a simple random sample of 500 de-identified live notes and score them together, so the set mirrors traffic with no manual selection
  • DCopy a month of identifiable live notes into the evaluation store unchanged, because de-identification strips the artefacts behind errors
A representative evaluation set samples real traffic in proportion and adds a separately reported slice of deliberately chosen edge cases. Hand-written cases reflect what the authors imagined, not what the system receives, which is why a 97 percent score coexisted with live errors. Proportional sampling of de-identified live notes reproduces the real input distribution, and an oversampled hard-case slice gives enough examples of rare patterns to measure them. Reporting that slice separately stops strong performance on common notes from masking failures on the cases that matter most.

Why A is wrong: This is tempting because it avoids patient data and scales quickly. It is wrong because model-generated notes tend to be cleaner and more regular than dictated ones, so they repeat the original set's flaw of missing the real-world artefacts that cause the errors.

Why B is correct: This is correct because sampling real de-identified traffic captures the artefacts the hand-written notes lacked and mirrors the specialty mix, while a separately reported edge-case slice keeps rare, high-risk patterns visible instead of diluted in the average.

Why C is wrong: This is tempting because random sampling from live traffic is representative on average. It is wrong because rare but important patterns appear only a handful of times, and one blended score lets a failure on them disappear inside the average.

Why D is wrong: This is tempting because it maximises realism. It is wrong because the privacy office permits only de-identified notes for testing, and under the minimum necessary standard identifiers are not needed to evaluate summarisation quality, whatever agreement covers the platform.

See more CCAR-P practice questions, answers explained.

Exam traps in Evaluation, Testing & Optimization

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAR-P bank for this domain.

  • The prompt revision was tuned on the 300 evaluation cases, so it is overfitted to the graded set and will not generalise

    Why it is wrong: Overfitting a prompt to a fixed evaluation set is a real and common failure, so it is tempting. It is wrong here because the clinicians reviewed summaries drawn from that same set: an overfitted revision would have looked better to them on those cases, yet they found no improvement.

  • Run the grader on a more capable model tier than the drafting model, so its rubric marks are more trustworthy than the drafts it scores

    Why it is wrong: This is tempting because a stronger model is often a better judge. It is wrong because capability is not evidence of agreement with tutors; the board asked for a demonstrated link to tutor judgment, and a stronger grader can still apply the rubric differently from the tutors.

  • Grade every field with a model grader that compares each output to the verified record, so one consistent method covers the whole suite

    Why it is wrong: This is tempting because a single method is simpler to maintain. It is wrong because model grading adds cost, latency and run-to-run variation to fields that have one correct value, where an exact comparison is both cheaper and fully reliable.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.