A national pensions agency uses Claude to draft replies to written enquiries, and a caseworker reviews and edits every draft before it is sent; the review tool already stores both the draft and the sent version. Release quality is judged on a 400-case regression set, re-run before each release, which has scored 92 percent on every release this year. The service lead sets two requirements: a fall in live draft quality must be detected within a week, and any failure pattern found in live traffic must be caught by the pre-release regression set from then on. No extra caseworker or reviewer hours are available for monitoring. Which two changes best meet these requirements? Select TWO.
- AHave the model rate its own confidence in each draft, and alert when the weekly mean confidence score falls below baseline
- BTrack the share of drafts that caseworkers heavily rewrite, by edit distance and enquiry type, and alert on a rise over baseline Correct
- CHave a senior caseworker re-review a random sample of 200 sent replies each week and score each one against the rubric
- DAdd heavily rewritten live drafts to the regression set each week, using the caseworker's sent reply as the reference answer Correct
- ERe-run the 400-case regression set nightly against the production system, and alert when its score drops below 92 percent
Why A is wrong: This is tempting because it costs no reviewer time and produces a number per draft. It is wrong because a model's self-reported confidence is not a calibrated accuracy signal: a drafting regression can leave stated confidence unchanged, so the alert can stay silent while caseworkers rewrite more drafts.
Why B is correct: Correct. Every draft is already paired with the version a caseworker approved, so the size of each edit is a free, expert label on live output. A rising rewrite rate in any enquiry type surfaces a quality fall within days, with no added review hours.
Why C is wrong: This is tempting because human scoring on live replies is a trusted quality measure. It is wrong here because the stem rules out any extra caseworker or reviewer hours, and the existing edits already carry the same expert judgment at no added cost.
Why D is correct: Correct. Feeding live failures into the regression set, with the approved reply as the reference, is what makes the pre-release gate catch that failure pattern on every later release, which is the second stated requirement.
Why E is wrong: This is tempting because it automates a check the team already trusts. It is wrong because a fixed set re-run against an unchanged system returns the same score each night; it cannot see failures in enquiry types the set does not hold, and it adds nothing new to the set.