An operations lead at a logistics company used one prompt asking Claude to summarise three depot incident reports, identify root causes, propose fixes and draft an email to depot managers. The summaries were accurate and detailed, but the root causes and fixes were generic, and the email simply repeated them. What is the most likely cause of the weak result?
- AThe incident reports were too long for Claude to read, so it only took in the opening of each report.
- BThe model chosen is not capable enough for analysis, and a more powerful model would fix all four parts.
- CClaude lacks specialist knowledge of logistics, so it cannot identify root causes in depot incident reports.
- DFour distinct tasks were packed into one request, so the analysis steps were not done or reviewed in their own right. Correct
Why A is wrong: Length limits are a real concern with large files, but the summaries were accurate and detailed. That shows Claude read the reports, so the weakness lies elsewhere.
Why B is wrong: A stronger model can help with analysis, so this is tempting. The pattern of good summaries but generic analysis points to how the request was structured, and a model switch does not add review between steps.
Why C is wrong: Domain knowledge matters, but the root causes would come from the reports themselves, which Claude summarised well. Concluding it cannot do the task skips the more likely fix of decomposing it.
Why D is correct: Bundling four tasks into one prompt spreads effort thinly and gives no chance to check the root causes before fixes and the email build on them. Splitting the work and reviewing each step targets the weak parts directly.