A government benefits agency uses Claude to draft replies to citizen letters, and a caseworker approves each draft before it is sent. The agency's legal counsel must be able to explain any reply months later if a citizen appeals. Counsel asks the architect to confirm in writing that the system will give the same reply every time it receives the same letter. In testing with the sampling temperature set to zero, the team still saw occasional wording differences between runs on identical input. What should the architect tell counsel?
- AThat identical replies cannot be guaranteed, so each reply's input, retrieved policy text, draft and caseworker approval are recorded for appeal Correct
- BThat setting the temperature to zero makes generation deterministic, so the agency can commit to identical replies for identical letters
- CThat each reply is cached against its letter's text, so a repeated letter returns the stored reply and the appeal can rely on that reply
- DThat an evaluation found 97 percent wording similarity across repeated runs, so the remaining variation is too small to affect any appeal
Why A is correct: Correct. It states the system's non-deterministic behaviour honestly and then gives counsel the control that actually meets the appeal requirement: a per-reply record of what went in, what policy was retrieved, what was drafted and who approved it, so any reply can be explained later regardless of whether a rerun would match.
Why B is wrong: This is tempting because a temperature of zero is widely described as making output deterministic. It is wrong because model output is not guaranteed to be identical across runs even at temperature zero, and the team has already observed differences, so putting this commitment in writing would give counsel an assurance the system cannot honour.
Why C is wrong: This is tempting because a cache does force identical output for byte-identical input. It is wrong because citizen letters are almost never byte-identical, so the cache rarely applies, and it does nothing to let counsel explain how a particular reply was produced; the appeal requirement is about traceability of each decision, not repetition.
Why D is wrong: This is tempting because a measured similarity figure sounds like rigorous evidence for a risk audience. It is wrong because an aggregate similarity score says nothing about whether a specific contested reply can be explained, and it invites counsel to treat residual variation as negligible when a single differing reply could be the one under appeal.