A national benefits agency's caseworker summarisation assistant has run in production for four months, built and operated by an external vendor that runs a weekly evaluation against 500 labelled case files. The contract ends next month and the service passes to the agency's in-house platform team, which has no machine learning specialists but already runs the agency's service dashboards. Agency policy requires the agency itself to detect any drop in summary quality within one week. What should the architect require before the vendor exits?
- AHave the vendor deliver a final evaluation report, a written runbook and the prompt history, so the in-house team understands how the summariser was built and tuned
- BKeep the vendor on a reduced support contract to investigate and fix any summary quality problems that caseworkers raise through the agency's service desk
- CTransfer the evaluation suite and labelled set, name an in-house owner for it, and schedule it to run weekly with an alert when scores fall below the current baseline Correct
- DAdd latency, error-rate and token-usage alerts to the in-house team's existing dashboards, so degradation in the summariser is visible on the screens they watch
Why A is wrong: This is tempting because documentation is a normal part of any handoff and helps the new team understand the system. It is wrong because a final report is a snapshot: once the vendor leaves, nobody runs the weekly evaluation, so a quality drop would go undetected and the one-week detection policy fails.
Why B is wrong: This is tempting because it retains the expertise that built the system. It is wrong because it is reactive: quality problems are found only when caseworkers notice and report them, which breaks the requirement that the agency itself detect a drop within a week.
Why C is correct: This is correct because it keeps the quality measurement running after the handoff and gives it a named owner, and an automated scheduled run with a baseline alert does not require machine learning specialists to operate, so the one-week detection requirement continues to be met.
Why D is wrong: This is tempting because it fits the in-house team's existing skills and tooling. It is wrong because those are operational health signals; a summariser can stay fast, error-free and on budget while its summaries become less accurate, so this does not measure quality at all.