CCAO-F - Output Evaluation and Validation (21% of the exam) - Section 2.2

Identify hallucinations, inconsistencies, and biases in responses.

Recognising fabricated specifics such as citation numbers, statistics or quotes, internal contradictions, and one-sided framing. The guide's sample item turns on a confident summary citing a specific subsection that must be checked before it is shared.

hallucinated specificsfabricated citationsinternal inconsistencybias in framing

Practice question for this objective

Free sampleOutput Evaluation and Validationmedium

An HR officer at a supermarket chain asks Claude to summarise 300 free-text staff survey responses about a proposed new shift pattern. The summary says staff 'overwhelmingly oppose' the change and supports this with four strongly worded quotes. The board will decide whether to go ahead based on the summary. Which check would best show whether the summary gives a fair picture of staff views?

  • AConfirm that each of the four quotes appears word for word in the survey responses before the summary goes to the board.
  • BAsk Claude whether its summary is balanced, and to rewrite any parts it judges one-sided before it goes to the board.
  • CAsk Claude to soften the wording to neutral phrasing, such as 'many staff have concerns', before sending it.
  • DCount how many responses oppose, support or are neutral, and compare that split with the summary's 'overwhelmingly oppose'. Correct
Selection bias in a summary is tested by comparing its quantifiers and chosen examples with the full data, not by confirming quotes are genuine. A summary can be one-sided without fabricating anything: it can pick the most vivid examples and describe them with sweeping words like 'overwhelmingly'. Verifying that quotes exist rules out fabrication but not unrepresentative selection. Measuring the actual split of views in the source data is what shows whether the summary's framing matches reality.

Why A is wrong: Checking quotes against the source is good practice and catches fabricated wording, so this feels thorough. But genuine quotes can still be the most extreme few out of 300, so this check cannot show whether 'overwhelmingly oppose' is true.

Why B is wrong: Asking for a self-review is quick and sounds responsible. It is not independent evidence, though: Claude is judging its own selection without anyone counting the responses the summary is supposed to represent.

Why C is wrong: Toning down strong language seems like a fix for bias. But rewording without checking the data only swaps one unverified claim for another, and 'many staff have concerns' could still misrepresent a workforce that mostly supports the change.

Why D is correct: The claim at risk is about proportion, so the check must measure proportion. Tallying the responses tests whether the summary's strongest word and its chosen quotes reflect the whole survey or a vocal minority.

See more CCAO-F practice questions, answers explained.

Exam traps in Output Evaluation and Validation

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAO-F bank for this domain.

  • Claude's summaries become less reliable when reused for a different audience, so the summary should be generated again for customers.

    Why it is wrong: It can seem that a new audience calls for a fresh draft. Reuse does not make an already checked text less accurate, and regenerating it would produce new, unchecked wording while leaving the real gap, approval for public release, unaddressed.

  • Claude has a built-in preference for environmentally friendly options and leans towards them whatever the prompt asks for.

    Why it is wrong: It is easy to assume the model has its own agenda. Here the prompt itself explains the slant: Claude was asked to make the case for electric vans, so there is no need to look for a hidden preference.

  • The association removed the survey from its website after Claude read it, so the figure is still sound to use.

    Why it is wrong: Tempting because published material does get taken down. It is wrong because the association itself has confirmed it never ran the survey, so there was nothing to remove and the figure has no basis.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.