A university runs a Claude-based course assistant for 22,000 students. Three teams change it on separate weekly schedules: one edits the system prompt, one rebuilds the course-material index, and one adjusts which model tier serves each request type. An automated grader scores a 2 percent sample of live conversations for faithfulness, and that score fell from 0.91 to 0.83 last month, while the offline evaluation set scored the same before and after every release. Nobody could say which change caused the drop. The provost's office requires that any future drop be traced to a specific change within one working day, and none of the teams may pause its release schedule. Which two changes best meet the requirement? Select TWO.
- ABundle the three teams' changes into one combined monthly release, so the date of any drop identifies the cause
- BRe-run the offline evaluation set after every release and compare its score with the previous release's result
- CStamp every logged request with the prompt version, the index build ID and the model tier that served it Correct
- DChart the live grader score split by each of those version fields, with each team's releases marked on the timeline Correct
- ERaise the live grader's sample from 2 to 20 percent of conversations so that a drop becomes visible sooner
Why A is wrong: This is tempting because fewer release events seem easier to line up against a trend. It is wrong because it pauses the teams' weekly schedules, which the stem forbids, and bundling three changes into one release makes it harder, not easier, to say which of them caused a drop.
Why B is wrong: This is tempting because per-release evaluation is good practice. It is wrong here because the stem shows the offline set did not move while live quality fell, so it cannot see this kind of regression and cannot attribute a live drop to a change.
Why C is correct: Correct. Recording the exact configuration behind each logged request means every graded live conversation can be tied to the prompt version, index build and tier that produced it, which is the raw material for attribution.
Why D is correct: Correct. Splitting the monitored score by prompt version, index build and model tier, with releases annotated, shows which field's change lines up with the fall, so the cause can be named within a day while all three teams keep shipping.
Why E is wrong: This is tempting because a larger sample does detect a drop faster and with less noise. It is wrong because the requirement is attribution, not detection: a bigger sample with no version data still shows only that quality fell, not which change caused it.