CCAR-P - Governance, Safety & Risk Management (14% of the exam) - Section 5.3

Apply human-in-the-loop validation strategies.

Placing human review where it matters: before irreversible or high-impact actions, on low-confidence or high-risk outputs, and on a sample for ongoing quality assurance. Candidates should calibrate review to risk rather than reviewing everything or nothing.

approval before irreversible actionsrisk-based reviewsampling for quality assuranceescalation paths

Practice question for this objective

Free sampleGovernance, Safety & Risk Managementmedium

A hospital trust uses a Claude-based assistant to draft discharge letters from the inpatient record, about 1,800 letters a day. Clinicians currently sign off every letter before it goes to the patient's GP, but they have capacity for roughly 450 reviews a day and the backlog now delays letters by up to four days. An audit shows that errors with clinical consequence sit almost entirely in the 20 percent of letters that change a medication or dose, while administrative letters had a correction rate under 1 percent. The clinical safety officer requires that a medication change is not sent without clinician sign-off. What should the architect recommend?

  • ARequire clinician sign-off on every letter that changes a medication or dose, release the rest automatically, and review a random weekly sample of released letters. Correct
  • BAsk the assistant to rate its confidence in each letter, route any letter rated below a set threshold to a clinician, and release the remainder automatically.
  • CKeep clinician sign-off on every letter and add a drafting pass in which the assistant checks its own letter against the record before it joins the queue.
  • DRelease every letter automatically with a footer telling the GP it was drafted by an assistant, and investigate any error that a receiving practice reports.
Calibrate human review to the risk of the content, mandating sign-off on high-impact outputs and sampling the rest, rather than reviewing everything or relying on self-reported confidence. Review capacity is the binding constraint, so it must be spent where an error has consequence. The audit shows the clinically significant errors concentrate in medication and dose changes, a property of the letter's content that can be detected deterministically rather than inferred from the model's own opinion of itself. Routing that 20 percent to mandatory sign-off satisfies the safety requirement within capacity, and a random sample of the auto-released letters provides the quality assurance evidence that the low-risk tier stays low risk.

Why A is correct: This places mandatory review on the content the audit identified as high risk, about 360 letters a day, which fits inside the 450-review capacity and meets the safety officer's requirement. The spare capacity funds a random sample of auto-released letters, which keeps the low error rate of administrative letters under ongoing measurement.

Why B is wrong: This is tempting because it cuts the review volume and sounds like risk-based routing. It is wrong because a model's self-reported confidence is not a calibrated accuracy signal, so a confidently wrong dose change could be released without the sign-off the safety officer requires.

Why C is wrong: A self-check pass may improve draft quality, which makes it attractive. It is wrong because it leaves 1,800 letters a day competing for 450 reviews, so the four-day backlog, the actual stated problem, is untouched.

Why D is wrong: This clears the backlog at once and adds transparency, so it can look pragmatic. It is wrong because medication changes would reach GPs without clinician sign-off, breaching the stated requirement, and errors would only be found if a practice happened to report them.

See more CCAR-P practice questions, answers explained.

Exam traps in Governance, Safety & Risk Management

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAR-P bank for this domain.

  • Move the assistant to a more capable model tier so that it tags emergencies more accurately, and keep the urgent messages in the general clinical inbox.

    Why it is wrong: Better detection sounds like a safety improvement, which makes this tempting. It is wrong because the tagging was already correct in 22 of 23 cases; the failure is that correctly tagged messages had no route to a clinician overnight, and a different model leaves that unchanged.

  • Keep approval on every tool call and move the agent to a more capable model tier so that each batch completes sooner and leaves more time before the cut-off.

    Why it is wrong: A more capable model can reduce rework, which makes this tempting. It is wrong because the delay comes from a human approving 30 to 40 steps, most of them reversible lookups, and a different model does nothing to remove that waiting time.

  • Reviewers are grading leniently and passing wrong answers, so their grades need calibrating against an expert-written answer key

    Why it is wrong: Lenient grading would hide errors, which makes this tempting. The stem rules it out: grades are already calibrated monthly against the product team's answer key and match it on 95 per cent of items.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.