A wealth management firm's retrieval-based assistant answers adviser questions about fund fee schedules. Last month the ingestion job was changed to split documents into chunks half the previous size with no overlap, and since then the share of adviser-flagged wrong answers has risen from 2 to 9 per cent. The model, system prompt and latency are unchanged, and most wrong answers quote a real fee figure from the right fund but attach it to the wrong share class. Per-request token spend must not rise. What should the architect recommend?
- AMove the assistant to a more capable model tier so that it can reason across fragmented chunks and infer which share class each fee belongs to
- BRaise the number of retrieved chunks per question from five to twenty so that the share-class heading chunk is more likely to be returned alongside
- CAdd a system prompt rule telling the model to confirm the share class named in a chunk before it quotes any fee figure from that chunk to an adviser
- DReplay the flagged questions against the old and new indexes, and fix the splitter so each fee table row stays in a chunk with its share-class heading Correct
Why A is wrong: This is tempting because a more capable tier is better at multi-step inference. It is wrong because the model did not change while the chunking did, and when the share-class heading is missing from the retrieved text no tier can reliably recover it; a larger tier also raises per-request spend, which the stem rules out.
Why B is wrong: This is tempting because it works at the retrieval layer and may sometimes pull the heading back in. It is wrong because it roughly quadruples the retrieved context per request, breaking the spend constraint, and even when the heading arrives the model must still guess which orphaned figure belongs under it.
Why C is wrong: This is tempting because it is cheap and aims directly at the symptom. It is wrong because the new chunks no longer contain the share-class heading, so there is nothing in the chunk for the model to confirm; an instruction cannot restore context that the ingestion change removed.
Why D is correct: The only change was the chunking, and the failure pattern (right figure, wrong share class) is what splitting a table away from its heading produces. Replaying the flagged questions against both indexes confirms the cause at the retrieval layer, and a structure-aware splitter fixes it without adding tokens to each request.