CCAR-P - Governance, Safety & Risk Management (14% of the exam) - Section 5.1

Implement guardrails and safety controls.

Layering safety controls around a Claude system: input screening, output filtering, restricted tool permissions and controls enforced outside the model. Candidates should recognise that a control enforced in code is a guarantee, while an instruction to the model is not.

input screeningoutput filteringpermission restrictionsenforcement outside the model

Practice question for this objective

Free sampleGovernance, Safety & Risk Managementmedium

A state benefits agency runs a triage agent that reads incoming claims and records an initial assessment. Policy says only a human caseworker may close a claim, so the team removed the close-claim tool from the triage agent's tool list. Request logs from every replica confirm that the close-claim definition has been absent since the change and that no call to it has been made. A week later an audit finds 14 claims closed with the triage agent's service account recorded as the actor. The agent still holds an update-claim tool that it uses to record assessment notes and the status of a claim. What is the most likely cause?

  • AThe update-claim tool can still set a claim's status to closed, so the capability survived the tool removal Correct
  • BPrompt caching kept serving the earlier tool list, so cached requests still offered the close-claim tool
  • CThe model inferred the closing step from its instructions and carried it out without needing any tool
  • DThe tool removal was deployed to only some replicas, so a share of requests still carried the old tool
Removing a named tool restricts a capability only if no remaining tool can reach the same state change, so restrictions belong in handlers and backend permissions. A permission restriction is about which state changes the agent can cause, not which tool names it can see. The close-claim tool is provably gone, yet the agent's account still closed claims, and its remaining update-claim tool writes the status of a claim. A general-purpose write tool reaches the same outcome, so the control must be enforced outside the model, in the handler or the case system's own permissions.

Why A is correct: The audit names the agent's account, and the only write path it still holds is a tool that sets a claim's status. Removing a named tool removes a capability only if no remaining tool reaches the same state change, so the restriction must be enforced in the update handler or the case system, for example by rejecting a closed status from this account.

Why B is wrong: Caching is easy to suspect when a configuration change seems not to take effect. It is wrong because prompt caching reuses processing of an identical prefix and does not change what a request contains, and the logs confirm the close-claim definition was absent from every request.

Why C is wrong: This is tempting if the model is pictured as acting directly on the case system. It is wrong because a model cannot change external state on its own; every state change happens through a tool call that application code executes, so a write path must exist.

Why D is wrong: A partial rollout is a common reason for a change appearing not to work, which makes this plausible. It is wrong because the stem states that logs from every replica show the definition absent and no call to it, which rules out a stale replica.

See more CCAR-P practice questions, answers explained.

Exam traps in Governance, Safety & Risk Management

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAR-P bank for this domain.

  • Citizens' enquiries contain embedded instructions telling the model to omit the statement, and the model obeys them.

    Why it is wrong: Prompt injection through user content is a real risk and worth considering. It is wrong here because the generation logs show the model produced the statement in 597 of 600 drafts, so the omission happens after generation, not inside it.

  • Set temperature to zero for every classification call so that rerunning the same product returns the same code during an audit

    Why it is wrong: Tempting because lower temperature does make outputs more consistent. It is wrong because it does not guarantee identical output across runs, and any later prompt or model change would shift codes again, so a regenerated code cannot be relied on as the recorded decision.

  • Call the vendor's public API through an egress proxy in the authorised environment that inspects and logs every request.

    Why it is wrong: A proxy adds visibility and feels like control, but the request still leaves the boundary and is processed outside it, so the stated requirement on processing is broken regardless of how well the traffic is logged.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.