CCAR-P domain - 16% of the exam

Evaluation, Testing & Optimization

Evaluation, Testing & Optimization is 16% of the Claude Certified Architect - Professional (CCAR-P) exam. These are the objectives it covers, each with practice questions, with every answer explained.

Objectives in this domain

Sample question from this domain

Free sampleEvaluation, Testing & Optimizationmedium

A hospital group is piloting Claude to draft discharge summaries from inpatient records, and the clinical safety officer has set one requirement in writing: no summary may leave out a documented drug allergy or a medication change. The team's evaluation plan scores each draft against a clinician-written reference summary using an overall similarity score, and the pilot average is 0.87. About one record in six carries an allergy or a medication change. What should the architect recommend the team measure before the pilot widens?

  • AThe overall similarity score, with its pass threshold raised from 0.87 to 0.92 across the pilot
  • BThe share of summaries the model itself rates as complete, from a self-check added to each draft
  • CRecall of documented allergies and medication changes per summary, on a labelled set, reported apart Correct
  • DClinician satisfaction with each draft, as a five-point rating from the ward doctors who sign it
Derive the metric from the stated requirement, measuring recall on safety-critical items separately so an aggregate quality score cannot hide omissions. The requirement is defined by a failure on specific items, so the metric has to count those items. An overall similarity score averages over all text and over records that contain no allergy or medication change, which lets a 0.87 mean coexist with missed allergies. Item-level recall on a labelled set, with a threshold tied to the safety officer's requirement and reported separately, exposes exactly the failure the requirement forbids.

Why A is wrong: Raising the bar is tempting because it looks like a stricter safety gate. It is wrong because the similarity score averages over the whole narrative and over the five in six records with no allergy or medication change, so a dropped allergy line barely moves it and a higher threshold still does not target the stated requirement.

Why B is wrong: A self-check is tempting because it is automatic and cheap to run on every draft. It is wrong because a model's own judgement of completeness is not a ground-truth accuracy signal; the same model that omitted an allergy can rate its draft complete.

Why C is correct: The requirement names a specific failure, the omission of an allergy or medication change, so the metric must count those items: recall against records labelled with them, reported separately from any aggregate so a good average cannot hide a miss.

Why D is wrong: Clinician ratings are tempting because the raters are the domain experts. It is wrong because a satisfaction score reflects overall impression and readability; a busy signer does not check every line against the record, so omissions go unmeasured.

Other domains in this exam

See also the CCAR-P cert hub, the study guide, and the cheat sheet.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.