A hospital trust's clinical coding assistant proposes diagnosis codes from discharge notes, and human coders accept or amend each proposal. Its monitoring alerts on API error rate, p95 latency and tool call failure rate, and all three stayed green for a quarter in which a monthly manual audit found the coders' amendment rate had risen from 12 percent to 31 percent. Traces later showed that for six weeks the terminology lookup tool had returned successful responses drawn from a code set the trust had retired, after a configuration change pointed the tool at an archived database. Which gap in the monitoring design most directly explains why six weeks passed before anyone noticed?
- AThe tool call failure threshold was set too high, so the lookup tool's errors stayed below the alert line.
- BTraces were sampled at too low a rate, so the requests that returned retired codes were not being stored.
- CThe assistant's self-reported confidence was not logged, so its low-confidence proposals were not flagged.
- DEvery alert watched whether calls completed, and none tracked the coders' amendment rate as a quality signal. Correct
Why A is wrong: This is tempting because the fault sat in a tool and a loose threshold is a common reason alerts stay quiet. It is wrong because the lookup tool returned successful responses throughout, so there were no tool errors for any threshold to catch.
Why B is wrong: This is tempting because trace sampling does limit what can be investigated after the fact. It is wrong because the traces did capture the retired codes, and stored traces raise no alert on their own; the missing piece was a metric that measured output quality.
Why C is wrong: This is tempting because a confidence score looks like a cheap quality signal. It is wrong because a model given valid-looking codes by its own tool has no reason to report low confidence, and self-reported confidence is not a reliable accuracy measure; the human amendment rate is an observed outcome.
Why D is correct: Correct. The fault produced well-formed, successful results that were clinically wrong, so error, latency and failure metrics could not move. A quality signal already present in the workflow, the rate at which coders amend proposals, rose sharply and would have surfaced the degradation within days if it had been tracked and alerted on.