CCAR-P - Integration (19% of the exam) - Section 3.4

Analyze observability challenges and select monitoring strategies at scale.

What to capture from a production LLM system so failures can be diagnosed: request and response traces, tool calls, token usage, latency and quality signals. Candidates should choose monitoring that surfaces silent quality degradation, not only errors and outages.

tracingtool call loggingquality signalssilent degradation

Practice question for this objective

Free sampleIntegrationhard

A hospital trust's clinical coding assistant proposes diagnosis codes from discharge notes, and human coders accept or amend each proposal. Its monitoring alerts on API error rate, p95 latency and tool call failure rate, and all three stayed green for a quarter in which a monthly manual audit found the coders' amendment rate had risen from 12 percent to 31 percent. Traces later showed that for six weeks the terminology lookup tool had returned successful responses drawn from a code set the trust had retired, after a configuration change pointed the tool at an archived database. Which gap in the monitoring design most directly explains why six weeks passed before anyone noticed?

  • AThe tool call failure threshold was set too high, so the lookup tool's errors stayed below the alert line.
  • BTraces were sampled at too low a rate, so the requests that returned retired codes were not being stored.
  • CThe assistant's self-reported confidence was not logged, so its low-confidence proposals were not flagged.
  • DEvery alert watched whether calls completed, and none tracked the coders' amendment rate as a quality signal. Correct
Monitoring that tracks only errors, latency and failures cannot detect silent quality degradation; production systems need an alerting quality signal tied to real outcomes. Operational metrics answer whether a request completed, not whether its output was right. A tool that returns successful but wrong data passes every operational check, so the only place the fault shows is in an outcome measure such as how often reviewers override the output. Tracking that rate with its own alert turns a silent degradation into a visible one.

Why A is wrong: This is tempting because the fault sat in a tool and a loose threshold is a common reason alerts stay quiet. It is wrong because the lookup tool returned successful responses throughout, so there were no tool errors for any threshold to catch.

Why B is wrong: This is tempting because trace sampling does limit what can be investigated after the fact. It is wrong because the traces did capture the retired codes, and stored traces raise no alert on their own; the missing piece was a metric that measured output quality.

Why C is wrong: This is tempting because a confidence score looks like a cheap quality signal. It is wrong because a model given valid-looking codes by its own tool has no reason to report low confidence, and self-reported confidence is not a reliable accuracy measure; the human amendment rate is an observed outcome.

Why D is correct: Correct. The fault produced well-formed, successful results that were clinically wrong, so error, latency and failure metrics could not move. A quality signal already present in the workflow, the rate at which coders amend proposals, rose sharply and would have surfaced the degradation within days if it had been tracked and alerted on.

See more CCAR-P practice questions, answers explained.

Exam traps in Integration

Answers that look right on this material and are not. Each one is a distractor from a different question in the CCAR-P bank for this domain.

  • This term's answers differ from earlier cohorts, so the rubric fits the new responses less well than it did.

    Why it is wrong: This is tempting because cohort shift is a real cause of drift in marking systems. It is wrong because median answer length is unchanged and a change in what students write cannot cut input tokens per request by nearly two thirds in a single week.

  • Tool results from the new servers are appended to the conversation, so each turn carries their payloads into the next request.

    Why it is wrong: This is tempting because large tool results are a common driver of growing input tokens in agentic loops. It is wrong because the logs show no call to any tool on the new servers, so they have produced no results to append.

  • Tighten the alert on the weekly unedited-sign rate so that it fires on any fall of one point across all practices

    Why it is wrong: This is tempting because a more sensitive alert seems to close the gap. It is wrong because a 30-point fall in a type that is 4 percent of volume moves the overall figure by little more than a point, inside its normal two-point band, so a one-point alert would fire constantly on ordinary noise without showing which type had failed, and it would fire only after the release had reached every practice.

Examworthy is not affiliated with or endorsed by Anthropic. Original, blueprint-aligned practice material only.