A homeware retailer's shopping assistant has a system prompt that restricts it to the retailer's own catalogue, orders and delivery questions. For six months a weekly review of 400 sampled conversations found fewer than 1 per cent of replies off scope. Three weeks ago the marketing team appended a block of seasonal campaign rules to the same system prompt, one of which tells the assistant to help shoppers compare gift ideas from anywhere they have seen them. The scope rule, the model and the retrieval setup are unchanged, and off-scope replies have since risen to 7 per cent, almost all of them recommending products sold by other shops. What is the most likely cause?
- AThe scope rule now sits far from the end of a longer prompt, so the model has stopped reading it
- BRetrieval now returns competitor product pages, which the model summarises as its suggestions
- CThe new campaign rule conflicts with the scope rule, and the model resolves the clash unevenly Correct
- DShoppers have learned to jailbreak the assistant, and the campaign season brought more of them
Why A is wrong: Instruction position can matter, which makes this tempting, but a positional effect would not explain why the failures are concentrated on recommending other shops' products, which is exactly what the new campaign rule invites.
Why B is wrong: Retrieval is a sensible first suspect for wrong content, but the stem states the retrieval setup is unchanged, and the timing matches the prompt edit rather than any index or data change.
Why C is correct: The scope rule did not change, but a newly added instruction now licenses talking about products from anywhere, so the prompt contains two competing directives and the model sometimes follows the newer, more specific one.
Why D is wrong: Adversarial users are a real risk, but nothing in the evidence suggests hostile input, and the failures line up with the wording of a newly added instruction rather than with unusual user behaviour.