A marketing manager at a software firm is redesigning the onboarding email sequence for new customers with Claude. After each draft she asks, 'Score this sequence out of 10 and improve it.' Over five rounds the score rose from 6 to 9. When she shared the final version, the customer success team said it ignores the three questions new customers ask most often. What best explains this result?
- AClaude scored each draft against its own general idea of a good email, so the team's knowledge of customer questions never entered the loop. Correct
- BFive rounds was too few, and further rounds of self-scoring would have brought the sequence in line with customer needs.
- CThe conversation grew too long, so Claude lost the details of the target audience given in the very first prompt.
- DA faster model was used, and a more capable model would have produced self-scores that reflected customer needs.
Why A is correct: The rising score only measured the drafts against Claude's general sense of a good onboarding email. Nobody fed in what customers actually ask, so the iterations improved polish while the real gap stayed untouched.
Why B is wrong: More rounds feels like more refinement, so this is tempting. But each extra round repeats the same closed loop: Claude keeps optimising against its own criteria, so the missing customer questions would still never enter the work.
Why C is wrong: Long conversations can blur earlier detail, which makes this a reasonable guess. Five rounds is not a long conversation, though, and the stem never says the customers' main questions were given to Claude in the first place.
Why D is wrong: Reaching for a stronger model is a common reflex. No model can score against customer questions it was never told about, and a self-assigned score is not evidence of accuracy at any tier.