A software company's internal IT helpdesk agent answers staff questions about laptops, access requests and VPN faults. Its weekly score on a fixed 300-question evaluation set fell from 91 percent to 77 percent between two runs. The change log for that week shows the model version, system prompt and IT knowledge base unchanged, and one platform change: the agent's six scoped tools were replaced by a shared enterprise bundle of 58 tools also used by the finance, HR and facilities agents. In 62 of the 69 newly failing cases, the agent's first tool call went to a system outside IT. What should the architect investigate first?
- ARerun the evaluation on a larger model tier, since a stronger model should pick the right tool even from a long tool list.
- BCompare the tools the agent now holds with the six its role uses, since the bundle swap is the only change in the failing window. Correct
- CCheck the IT knowledge base index for stale entries, since retrieval faults are a frequent cause of a sudden accuracy drop.
- DEnlarge the evaluation set before acting, since a 14-point drop on 300 questions could be sampling noise rather than a real change.
Why A is wrong: This is tempting because larger models are often better at tool selection. It is wrong as a first step because the model did not change, so a model swap would mask rather than explain a regression caused by the tool set, and it adds cost without removing the 52 tools the role does not need.
Why B is correct: Correct. The 'what changed' evidence points to one change, and the failure pattern of first calls going to non-IT systems is the signature of wrong-tool selection in a bloated tool set. Confirming that the extra 52 tools are what the agent is reaching for leads directly to scoping it back to its six.
Why C is wrong: This is tempting because retrieval and indexing are a common source of confident-but-wrong answers. It is wrong here because the knowledge base was unchanged and most failures began with a call to a system outside IT, before any IT retrieval took place.
Why D is wrong: This is tempting because small evaluation sets can produce misleading swings. It is wrong because on a fixed set of 300 questions a drop from 91 to 77 percent is far larger than run-to-run noise, and the 69 new failures share one clear pattern.