CCAR-P - Stakeholder Communication & Lifecycle Management (14% of the exam) - Section 6.3

Manage stakeholder feedback loops and expectation alignment (including SLAs).

Setting realistic expectations about what an AI system will and will not do, agreeing SLAs that it can meet, and building feedback loops so stakeholder input reaches the system. Candidates should handle the gap between an expectation and what the system can deliver.

expectation settingSLAsfeedback loopsprobabilistic output

Practice question for this objective

Free sampleStakeholder Communication & Lifecycle Managementmedium

A government benefits agency uses Claude to draft replies to citizen letters, and a caseworker approves each draft before it is sent. The agency's legal counsel must be able to explain any reply months later if a citizen appeals. Counsel asks the architect to confirm in writing that the system will give the same reply every time it receives the same letter. In testing with the sampling temperature set to zero, the team still saw occasional wording differences between runs on identical input. What should the architect tell counsel?

  • AThat identical replies cannot be guaranteed, so each reply's input, retrieved policy text, draft and caseworker approval are recorded for appeal Correct
  • BThat setting the temperature to zero makes generation deterministic, so the agency can commit to identical replies for identical letters
  • CThat each reply is cached against its letter's text, so a repeated letter returns the stored reply and the appeal can rely on that reply
  • DThat an evaluation found 97 percent wording similarity across repeated runs, so the remaining variation is too small to affect any appeal
Explain a non-deterministic system to legal stakeholders honestly and meet their explainability need through per-decision records, not promises of identical output. Generation can vary between runs even with identical input and a temperature of zero, so a written promise of identical replies would be false. What counsel actually needs is the ability to explain a specific reply on appeal, and that comes from recording each reply's input, the policy passages retrieved, the draft produced and the caseworker's approval at the time. The guarantee moves from reproducing the generation to reproducing the evidence trail, which the architecture can deliver.

Why A is correct: Correct. It states the system's non-deterministic behaviour honestly and then gives counsel the control that actually meets the appeal requirement: a per-reply record of what went in, what policy was retrieved, what was drafted and who approved it, so any reply can be explained later regardless of whether a rerun would match.

Why B is wrong: This is tempting because a temperature of zero is widely described as making output deterministic. It is wrong because model output is not guaranteed to be identical across runs even at temperature zero, and the team has already observed differences, so putting this commitment in writing would give counsel an assurance the system cannot honour.

Why C is wrong: This is tempting because a cache does force identical output for byte-identical input. It is wrong because citizen letters are almost never byte-identical, so the cache rarely applies, and it does nothing to let counsel explain how a particular reply was produced; the appeal requirement is about traceability of each decision, not repetition.

Why D is wrong: This is tempting because a measured similarity figure sounds like rigorous evidence for a risk audience. It is wrong because an aggregate similarity score says nothing about whether a specific contested reply can be explained, and it invites counsel to treat residual variation as negligible when a single differing reply could be the one under appeal.

See more CCAR-P practice questions, answers explained.

Exam traps in Stakeholder Communication & Lifecycle Management

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAR-P bank for this domain.

  • The refusal instructions in the prompt are too strict, so the assistant now declines personal calculations it is permitted to attempt.

    Why it is wrong: Over-strict refusal wording is a real cause of unhelpful declines, so this is tempting. It is wrong because the prompt has not changed, so the same refusals happened during the year the complaint rate sat at 3 per cent, and declining personal calculations is the designed behaviour rather than an error.

  • An architecture overview explaining each component, how the index is built and how the tool integration works

    Why it is wrong: Tempting because on-call engineers do need to understand the system. It is wrong because an overview explains how the service works, not what to do when the index refresh fails, so an engineer would still be diagnosing from first principles against a 30 minute target.

  • Users rate the tone and speed of replies rather than their accuracy, so polite but wrong answers still score highly

    Why it is wrong: Rating bias towards tone is a known weakness of satisfaction scores, so this is a plausible suspicion. It does not explain the evidence, because the callers describe giving up partway through, and users who leave by closing the window are never shown the rating at all, so how raters weight tone cannot account for failures that never reach the rating.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.