A hospital group is building an evaluation set for a Claude assistant that summarises discharge notes. The first version holds 150 notes the team wrote by hand and scores 97 percent, yet clinicians report errors on live notes with dictation artefacts, mixed medication lists and unusual abbreviations. Live traffic is 70 percent general medicine, with paediatric and oncology notes making up most of the rest. A business associate agreement is in place, and the privacy office permits only de-identified notes for testing. Which approach best produces a representative evaluation set?
- AAsk Claude to generate 2,000 synthetic discharge notes across every specialty, then have clinicians spot-check a tenth of them for realism
- BSample de-identified live notes in proportion to the specialty mix, add an oversampled slice of known hard cases, and report that slice apart Correct
- CDraw a simple random sample of 500 de-identified live notes and score them together, so the set mirrors traffic with no manual selection
- DCopy a month of identifiable live notes into the evaluation store unchanged, because de-identification strips the artefacts behind errors
Why A is wrong: This is tempting because it avoids patient data and scales quickly. It is wrong because model-generated notes tend to be cleaner and more regular than dictated ones, so they repeat the original set's flaw of missing the real-world artefacts that cause the errors.
Why B is correct: This is correct because sampling real de-identified traffic captures the artefacts the hand-written notes lacked and mirrors the specialty mix, while a separately reported edge-case slice keeps rare, high-risk patterns visible instead of diluted in the average.
Why C is wrong: This is tempting because random sampling from live traffic is representative on average. It is wrong because rare but important patterns appear only a handful of times, and one blended score lets a failure on them disappear inside the average.
Why D is wrong: This is tempting because it maximises realism. It is wrong because the privacy office permits only de-identified notes for testing, and under the minimum necessary standard identifiers are not needed to evaluate summarisation quality, whatever agreement covers the platform.