A cloud software company is building an internal platform assistant for 2,000 engineers that must act across 180 tools spread over 14 internal services. Last quarter's tickets show that 40 percent of tasks touch three or more services, and that most tools are needed by fewer than one task in a hundred, though every tool is needed by some task. With all 180 definitions loaded on every request, the prototype picks the wrong tool on 11 percent of turns and exceeds the cost ceiling per task set by finance. The platform lead requires that rarely used tools stay reachable within the same conversation as the rest of the task. What should the architect recommend?
- ALoad only the 30 most used tools and retire the other 150, sending long-tail requests to a ticket queue for the owning service team.
- BSplit the assistant into 14 service-specific agents behind a router that sends each request to the single agent owning that service.
- CExpose a search over the tool catalogue that returns matching definitions on demand, and test on a task set that needed tools are found. Correct
- DKeep every definition loaded and move to a more capable model tier, so tool selection improves across the full catalogue of 180 tools.
Why A is wrong: Pruning to frequently used tools is the right move when tools are genuinely unneeded, and it cuts both cost and confusion. Here every tool is needed by some task, so retiring 150 of them breaks the requirement that rarely used tools stay reachable within the conversation.
Why B is wrong: Specialised agents each hold a small tool set, which improves selection and lowers cost, so this is a credible design. It fails the stated workload: 40 percent of tasks span three or more services, and routing each request to a single service agent leaves those tasks without the tools they need in one conversation.
Why C is correct: Discovering tools progressively keeps only a small search capability in the starting context, which reduces cost and the number of competing definitions behind wrong-tool choices, while every tool remains reachable in the same conversation. Measuring discovery recall on representative tasks addresses the main risk of the pattern, that a needed tool is never found.
Why D is wrong: A stronger model can choose more accurately among many tools, which makes this attractive. It addresses the wrong layer and moves cost the wrong way: all 180 definitions are still sent on every request, so the per-task ceiling is breached by more, and the distraction from irrelevant definitions remains.