A retail bank's head of complaints asks for a customer-facing chatbot. Discovery shows that 18 per cent of complaints breach the regulator's five-business-day acknowledgement deadline, and case timing shows most of the delay sits in staff manually sorting each complaint into one of 12 regulatory categories before it can be routed. The team has eight weeks and must show the sponsor a measurable improvement in that breach rate. What should the architect recommend?
- ABuild the customer-facing chatbot as requested, measured by how many complaints customers submit through it rather than by email each week
- BFine-tune a model on five years of past complaints so it learns the 12 categories, and begin routing only once that training has finished
- CDeploy an agent that investigates each complaint, decides the outcome and sends the final response letter to the customer without staff review
- DUse Claude to classify each new complaint into the 12 categories and route it, measured by breach rate and by accuracy on a labelled sample Correct
Why A is wrong: It is tempting because it delivers exactly what the sponsor asked for. It is wrong because the measured delay sits in internal categorisation, not in how complaints arrive, so a new intake channel leaves the breach rate untouched and its success metric says nothing about the original problem.
Why B is wrong: It is tempting because historical labelled complaints look like ideal training data. It is wrong because fine-tuning comes before any prompt-based classification has been tried, adds data preparation and evaluation work that threatens the eight-week deadline, and may not beat a well-prompted model on a 12-category task.
Why C is wrong: It is tempting because it promises to clear the whole backlog at once. It is wrong because it is far larger than the stated problem, removes human judgement from a high-impact regulated decision, and the acknowledgement deadline only requires faster categorisation and routing.
Why D is correct: This targets the step the timing data identified as the bottleneck, uses a language model for a text classification task it suits, and ties success to the breach rate the sponsor cares about plus an accuracy check that catches misrouting.