CCAR-P - Integration (19% of the exam) - Section 3.3

Evaluate accuracy-latency trade-offs and justify configuration decisions.

Configuring a system for a stated latency target without giving up the accuracy the task needs: model choice, reasoning depth, retrieval breadth, streaming and batching. Candidates should justify a configuration against the stated requirement rather than maximise either dimension alone.

latency targetsaccuracy requirementsstreamingbatch versus realtime

Practice question for this objective

Free sampleIntegrationmedium

A city council's planning portal receives about 9,000 public comments a month on planning applications. Two requirements apply: when a comment is submitted, the portal must tell the resident within 5 seconds whether it raises a material planning consideration or needs more detail, and each comment must have a structured summary in the planning officer's register within 10 working days. On a 600-comment evaluation, a smaller model passed the completeness check on 97 percent of comments against a 95 percent bar, answering well inside 5 seconds, but produced acceptable register summaries on only 74 percent against a 90 percent bar, while a larger model met both bars. The current design runs both tasks in realtime on the larger model and exceeds the council's fixed monthly budget. What should the architect recommend?

  • ARun both tasks in realtime on the smaller model, which keeps the 5 second response and brings spend well inside the budget.
  • BRun both tasks through the Message Batches API on the larger model, as batch processing suits a register updated over 10 days.
  • CRun the completeness check in realtime on the smaller model and send register summaries through the Message Batches API on the larger model. Correct
  • DKeep both tasks in realtime on the larger model and shorten each register summary to a single sentence to cut output spend.
Split a workload by each task's own latency and accuracy requirement, using realtime only where someone waits and batch processing where nobody does. The two tasks have different binding constraints. The completeness check has a person waiting and a task the smaller model handles well, so a fast realtime call fits. The register summary needs the larger model's accuracy but has a 10 working day deadline, so asynchronous batch processing can run it at lower cost than realtime. Matching each task to its own requirement meets both bars and relieves the budget.

Why A is wrong: It is tempting because one cheaper model simplifies the design and fixes the budget. It is wrong because the smaller model's register summaries scored 74 percent against a 90 percent bar, so it trades away a stated accuracy requirement.

Why B is wrong: Batching keeps the stronger model and lowers cost, which fits the summary task well. It is wrong because batch results are asynchronous, so the resident would not receive the completeness check within the required 5 seconds.

Why C is correct: Each task gets the cheapest configuration that meets its own requirement: the smaller model clears the 95 percent check bar at 97 percent inside the 5 second window, and asynchronous batch processing on the larger model meets the 90 percent bar well within 10 working days at lower cost than realtime.

Why D is wrong: Cutting output tokens does reduce spend, so this looks like a low-effort budget fix. It is wrong because it truncates a structured summary the officers require, trading away the content of the register when a scheduling change saves money without that loss.

See more CCAR-P practice questions, answers explained.

Exam traps in Integration

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAR-P bank for this domain.

  • The larger model with a large thinking budget, streamed, since it scores highest and streaming hides the wait.

    Why it is wrong: This is tempting because it maximises correctness and streaming is a familiar latency fix. It is wrong because the ceiling applies to the complete reply, and streaming does not shorten the 14 seconds it takes for the full solution to arrive.

  • Adopt the smaller model, because a 1.2 second label gets urgent messages to the duty nurse about ten seconds sooner.

    Why it is wrong: Faster paging sounds like a clinical benefit, so this is tempting. It is wrong because 91.3 percent recall misses the 98 percent floor, meaning more urgent messages are labelled routine, while the extra speed improves on a 30 second target that is already met.

  • The 3-point overall fall is within run-to-run noise on a 500-incident set, so the configuration can ship as it stands without any further breakdown.

    Why it is wrong: Tempting because a 3-point aggregate movement can look like noise, and shipping the latency win is attractive. It is wrong because the breakdown shows no change at all on 410 items and a loss of 15 of 90 on one category, a structured pattern that noise does not produce.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.