A health insurer's clinical coding assistant is a single model call augmented with retrieval over coding guidelines and a tool that looks up code descriptions. Its error rate on audited claims is 8 percent, and an engineer proposes rebuilding it as an agent that can plan, search repeatedly and check its own work. Review of the 240 miscoded claims finds that in 205 of them the retrieved guideline passages came from the previous year's codebook edition, which the index still holds alongside the current one, and that the model applied those passages faithfully. What does this evidence indicate?
- AThe single call has no chance to check its own work, so an agent that searches again would catch the superseded passages before answering.
- BThe model is applying outdated guidance absorbed in training, so fine-tuning on the current codebook edition is needed to override that knowledge.
- CThe model's reasoning is too weak for coding rules, so moving the same call to a larger model tier would reduce the share of misapplied passages.
- DThe errors originate in retrieval returning the superseded edition, so restricting retrieval to the current edition fixes them without changing the pattern. Correct
Why A is wrong: This is tempting because self-checking agents can catch some errors. It is wrong because a repeat search would query the same index holding both editions and could return the same superseded passages, so the agent adds cost and variability without removing the source of the errors.
Why B is wrong: This is tempting because outdated model knowledge is a real failure mode. It is wrong because the review shows the outdated guidance came from retrieved passages, not from the model's training, so fine-tuning would leave the index serving the wrong edition.
Why C is wrong: This is tempting because a larger tier is a common response to accuracy problems. It is wrong because the passages were applied faithfully, so the reasoning is sound and the input is wrong; a larger model given the same superseded passages would produce the same codes.
Why D is correct: Correct. Most errors trace to the retrieval augmentation serving outdated passages that the model then applied correctly, so filtering the index to the current edition addresses the cause while the single augmented call stays as simple and predictable as before.