A software company runs two Claude-based review paths. An editor plug-in suggests fixes as developers type and must return its first token within 1.5 seconds. A nightly job reviews every merged pull request that touches concurrency code, and engineers read its findings the next morning. On 200 held-out concurrency pull requests, enabling extended thinking on the nightly reviewer raised the share of seeded defects found from 58 to 81 per cent and added about 40 seconds per review, and the engineering director has approved the higher cost for that job. Where should extended thinking be enabled?
- AOn both paths, so that developers and the nightly job receive reasoning of the same depth.
- BOn the editor plug-in only, because interactive suggestions reach developers the soonest.
- COn the nightly concurrency review only, leaving the editor plug-in to answer without it. Correct
- DOn neither path, moving the nightly job to the fastest tier to offset the cost of the reviews.
Why A is wrong: Consistency across tools is an appealing principle, and the nightly gain suggests the plug-in might improve too. It is wrong because extended thinking adds generation time before the answer, which the 1.5 second first-token budget on the editor path cannot absorb, and the plug-in's quick fix suggestions were never shown to need deeper reasoning.
Why B is wrong: It is tempting to put the best reasoning where users see it first. It is wrong on both counts: the plug-in has the tight latency budget that extended thinking would break, and the measured 23-point gain in defects found belongs to the nightly concurrency review, which this option leaves without it.
Why C is correct: This is correct because extended thinking earns its latency where a task needs multi-step reasoning and the consumer is not waiting: the nightly review has no one reading it until morning, a measured gain from 58 to 81 per cent, and approved cost. The editor plug-in keeps its 1.5 second budget by answering directly.
Why D is wrong: Cost reduction is a reasonable instinct for a job that runs on every merged pull request. It is wrong because the director has already approved the higher cost, and dropping both extended thinking and model capability on hard concurrency reasoning gives up the measured improvement in defect detection that the job exists to deliver.