A mortgage lender uses Claude to extract and check figures from application packs, and finance has set a ceiling on cost per completed application, where completed means the pack passed every check without manual rework. To cut spend, the team moved the extraction step to a smaller model tier. Cost per request fell by 38 per cent, but the automatic retry rate rose from 4 to 19 per cent and more packs now go to an underwriter for manual correction. Which metric should decide whether the team keeps the change?
- ACost per request, since a 38 per cent fall on each call carries through to the monthly bill
- BTotal model and rework cost per completed application, including retries and underwriter time Correct
- CAverage input and output tokens per request, compared before and after the model change
- DExtraction accuracy on the first attempt, against the previous tier on a labelled set of packs
Why A is wrong: This is tempting because the fall is real and directly measured. It is wrong because retries multiply the number of requests per application and underwriter rework is not counted at all, while the ceiling finance set is per completed application.
Why B is correct: This matches the unit finance set the ceiling on. It folds in the extra requests from retries and the cost of underwriter correction, so it shows whether the cheaper tier lowers or raises what each finished application actually costs.
Why C is wrong: Token counts are tempting because they drive the model bill. It is wrong because they are a per-request proxy that ignores the retry rate and the cost of manual correction, so they cannot show whether the ceiling per completed application is met.
Why D is wrong: Accuracy is tempting because the retries and rework stem from extraction errors. It is wrong because an accuracy figure alone does not answer a cost ceiling; the decision turns on what the errors cost per completed application, which accuracy does not express.