A software company's support assistant drafts email replies through a seven-call prompt chain: detect the language, extract the product and version, retrieve knowledge-base articles, write an outline, expand the outline into a reply, adjust the tone, and check that every claim cites a retrieved article. Code validates the extracted product and version against the release catalogue, and legal requires the citation check to pass before any reply is sent. The outline, expansion and tone calls each rewrite the previous call's prose with no check between them, and together they take 11 of the chain's 14 seconds at p95 against a 7-second target. In trials, a single drafting call matched their combined rubric score in about 3 seconds, while moving every call to a smaller model tier lowered the rubric score from 4.4 to 3.7. What should the architect recommend?
- ARun all seven calls in parallel on the incoming email and combine their outputs in code before the reply is sent
- BDrop the citation check call and rely on the drafting prompts' instruction to cite only retrieved articles
- CMerge the outline, expansion and tone calls into one drafting call, keeping extraction and the citation check separate Correct
- DMove every call in the chain to the smaller model tier and add a second tone pass to recover the lost rubric score
Why A is wrong: Parallelisation is tempting because it attacks wall-clock time directly, and it is the right pattern for independent subtasks. These steps are not independent: retrieval needs the extracted product, expansion needs the outline and the citation check needs the finished reply, so they cannot run side by side.
Why B is wrong: Removing a call does cut latency, and an instruction to cite sources often improves grounding. It trades away the one step legal requires to pass before sending, replacing a verifiable gate with a prompt instruction that nothing checks.
Why C is correct: A chain boundary earns its latency when the step's output can be checked on its own. The three prose-rewriting calls have no check between them, so merging them loses no verification point, matches their quality in the trial and brings p95 to about 6 seconds, while the validated extraction and the required citation gate stay as distinct, checkable steps.
Why D is wrong: A smaller tier is usually faster, which makes this a natural first idea for a latency target. The trial already showed the quality loss, and adding another rewriting call adds latency and another unchecked step rather than removing the steps that carry no verification value.