A state licensing agency runs an assistant that answers caseworker questions by placing up to 25 retrieved regulation excerpts into each request. The system prompt sets two binding rules: cite the excerpt behind every statement, and decline to give a determination on any individual application. The prompt is currently ordered as rules first, then the caseworker's question, then the excerpts. On 200 held-out questions, rule adherence is 97 per cent when fewer than five excerpts are retrieved and 81 per cent when more than twenty are, while retrieval precision is unchanged. Some answers genuinely need every excerpt, and the agency must stay within its current per-request budget. Which change should the architect recommend?
- APut the excerpts first in the user turn, then restate both binding rules and the question at the end Correct
- BMove the assistant to the most capable model tier so it can hold the rules across longer prompts
- CRestate the two rules in capital letters at the very start of the system prompt to raise their weight
- DCap retrieval at five excerpts per request so every prompt stays inside the high-adherence range
Why A is correct: This is correct because with long material in context, models attend more reliably to instructions and the query placed after the documents, close to where generation starts. It keeps every excerpt, adds only the few tokens of the restated rules, so it fits the budget while addressing the length-linked drop in adherence.
Why B is wrong: This is tempting because a more capable model often follows instructions more reliably in long prompts. It is wrong because it raises per-request cost against a fixed budget and treats a prompt structure problem as a capability problem; the evidence points at where the rules sit relative to a long block of excerpts.
Why C is wrong: This is tempting because emphasis is a common first reaction to ignored instructions. It is wrong because it leaves the rules at the far end of the prompt from where the model generates its answer, with the long excerpt block still sitting between them, so the positional cause of the drop is untouched.
Why D is wrong: This is tempting because the data shows adherence is high with few excerpts. It is wrong because the stem states some answers need every excerpt, so capping retrieval trades away answer completeness, which is truncating required context to protect a metric.