A city council's planning portal receives about 9,000 public comments a month on planning applications. Two requirements apply: when a comment is submitted, the portal must tell the resident within 5 seconds whether it raises a material planning consideration or needs more detail, and each comment must have a structured summary in the planning officer's register within 10 working days. On a 600-comment evaluation, a smaller model passed the completeness check on 97 percent of comments against a 95 percent bar, answering well inside 5 seconds, but produced acceptable register summaries on only 74 percent against a 90 percent bar, while a larger model met both bars. The current design runs both tasks in realtime on the larger model and exceeds the council's fixed monthly budget. What should the architect recommend?
- ARun both tasks in realtime on the smaller model, which keeps the 5 second response and brings spend well inside the budget.
- BRun both tasks through the Message Batches API on the larger model, as batch processing suits a register updated over 10 days.
- CRun the completeness check in realtime on the smaller model and send register summaries through the Message Batches API on the larger model. Correct
- DKeep both tasks in realtime on the larger model and shorten each register summary to a single sentence to cut output spend.
Why A is wrong: It is tempting because one cheaper model simplifies the design and fixes the budget. It is wrong because the smaller model's register summaries scored 74 percent against a 90 percent bar, so it trades away a stated accuracy requirement.
Why B is wrong: Batching keeps the stronger model and lowers cost, which fits the summary task well. It is wrong because batch results are asynchronous, so the resident would not receive the completeness check within the required 5 seconds.
Why C is correct: Each task gets the cheapest configuration that meets its own requirement: the smaller model clears the 95 percent check bar at 97 percent inside the 5 second window, and asynchronous batch processing on the larger model meets the 90 percent bar well within 10 working days at lower cost than realtime.
Why D is wrong: Cutting output tokens does reduce spend, so this looks like a low-effort budget fix. It is wrong because it truncates a structured summary the officers require, trading away the content of the register when a scheduling change saves money without that loss.