A government benefits agency assembles each decision-letter prompt from separately owned modules: a tone module, an eligibility module and a statutory-notices module carrying the mandatory appeal-rights paragraph. Each module has its own test suite, and every suite passes. Three weeks ago the communications team added a plain-language module, appended after the others, that limits letters to 200 words. Since then reviewers find the appeal-rights paragraph missing from 12 percent of letters. The model and the original three modules are unchanged. What should the team investigate first?
- AWhether a larger model tier would follow each module's instructions more reliably now the assembled prompt is longer
- BThe fully assembled prompt for affected letters, checking whether the new word limit conflicts with the mandatory paragraph Correct
- CThe statutory-notices module's own test suite, since its passing tests must lack a case for the appeal paragraph
- DThe model provider's release notes, since a silent change to the model's weights would explain a sudden regression
Why A is wrong: Tempting because a more capable model can handle more instructions at once. It is wrong because the regression began with a prompt change while the model stayed the same, so a model swap addresses the wrong layer and would leave the conflicting instructions in place.
Why B is correct: Correct. The only change was a new module appended last, and a 200-word ceiling can compete directly with a mandatory paragraph from another module. Reading the rendered prompt as the model receives it is the quickest way to confirm that cross-module conflict.
Why C is wrong: Tempting because missing test coverage is a common cause of silent regressions. It is wrong because that module and its tests are unchanged and passed before and after; a test that exercises one module in isolation cannot reveal a conflict that only exists once modules are combined.
Why D is wrong: Tempting because unexplained regressions are often blamed on the model. It is wrong because the stem states the model is unchanged, and the onset coincides exactly with the new module, which points to the prompt rather than the weights.