Anthropic free practice

Free CCAR-P practice questions

21 real CCAR-P sample questions, each with an explanation of why every option is right or wrong. No account, no card. This is the reasoning the CCAR-P tests: knowing why the tempting answer is wrong, not just spotting the right one.

The real CCAR-P is 63 questions in 120 minutes, pass mark 720 / 1000. For a domain-by-domain breakdown and a study plan, read the CCAR-P study guide. The full bank has 318 questions.

Integration (19% of the exam)

Free sampleIntegrationmedium

A hospital network is piloting a discharge-summary agent that reads a patient's notes, medications and results and drafts a summary for the attending clinician to edit. The agent was built on the hospital's shared clinical tool gateway and inherited every tool on it, including place_medication_order and cancel_appointment, which the summary role never calls. Clinicians have said they will abandon the pilot if it adds interruptions to their workflow, and go-live is in three weeks. The security lead asks for a recommendation on the two write tools. What should the architect recommend?

  • ARemove both write tools from the agent's configuration so its tool set covers only the read operations the summary role uses Correct
  • BKeep both tools but require the clinician to approve each order or cancellation in a confirmation dialog before it runs
  • CKeep both tools and log every invocation to the security monitoring platform, with an alert on any call from the agent
  • DAdd a system-prompt instruction forbidding medication orders and cancellations, and verify it with a red-team test suite
When an agent holds a capability its role never uses, removing that capability beats confirming, logging or instructing against its use. Least privilege is a structural control: a tool that is not in the agent's configuration cannot be invoked by any prompt, error or injected instruction. Confirmations, logs and prompt rules all leave the capability reachable and either add friction or act after the harm. Because the summary role is read-only, removing the write tools costs nothing in function while removing the risk entirely.

Why A is correct: Correct. The summary role never places orders or cancels appointments, so removing those tools eliminates the risk outright rather than detecting or gating it. It adds no clinician interruptions, shrinks the attack surface a manipulated prompt could reach, and is a configuration change that fits the three-week timeline.

Why B is wrong: A confirmation step feels safe because a human sees every write before it lands. It is wrong because it is a compensating control on a capability the role does not need, it adds exactly the workflow interruptions clinicians said would end the pilot, and approval fatigue tends to turn confirmations into reflexive clicks.

Why C is wrong: Logging with alerting is tempting because it adds no clinician friction and gives the security team visibility. It is wrong because it is a detective control: an erroneous medication order or cancelled appointment has already happened by the time the alert fires, and the capability itself remains.

Why D is wrong: An instruction plus red-team testing looks rigorous and costs nothing at runtime. It is wrong because a prompt instruction is a soft control that injected content or an unusual input can override, and passing a test suite does not prove the model will refuse every case when the tool is still available to call.

Free sampleIntegrationmedium

A retail bank's mortgage assistant answers signed-in customers' questions about their own repayment schedule. Its integration calls the core lending API with a service credential that holds read and write scopes on every loan account, because that credential was already provisioned for the batch servicing system. An internal audit finding requires that a compromised or manipulated assistant can neither change any loan nor read another customer's loan, and the bank wants to keep the assistant's current response time. What should the architect recommend?

  • AKeep the shared credential and route every write request the assistant makes through a human approval queue in the servicing team
  • BIssue a dedicated credential with read-only scope on schedule data, and attach the customer's loan scope server-side from the session Correct
  • CKeep the shared credential, log every lending API call with the customer's identifier, and review the log for anomalies daily
  • DRotate the shared credential monthly and instruct the model in its system prompt to pass only the signed-in customer's loan number
Give an assistant a dedicated credential scoped to its role and attach the customer scope server-side, rather than reusing a broad service credential. Reusing a credential provisioned for another system silently grants the assistant every scope that system needed. Least privilege here has two parts: shrink the credential to the read operations the role performs, and take the customer's account scope from the authenticated session instead of from anything the model produces. Both are enforced before a call reaches the lending API, which is what the audit finding demands.

Why A is wrong: An approval queue is tempting because it stops unreviewed loan changes. It is wrong because the assistant's role needs no writes at all, so this gates a capability that should not exist, and it does nothing about the audit's second requirement: the shared credential can still read any other customer's loan.

Why B is correct: Correct. A read-only credential limited to schedule data removes the write capability, and binding the loan scope from the authenticated session in the integration layer prevents cross-customer reads regardless of what the model asks for. Both are preventive controls and add no review step, so response time is unchanged.

Why C is wrong: Logging is tempting because it adds no latency and creates an audit trail the finding seems to ask for. It is wrong because the finding requires that the assistant cannot change or read other loans; a daily log review only discovers a write or a cross-customer read after it has happened.

Why D is wrong: Rotation and a scoping instruction both sound like hygiene. They are wrong because rotation does not narrow what the credential can do, and letting the model supply the loan number means a manipulated conversation can request another customer's loan; the scope must come from the session, not from model output.

Free sampleIntegrationmedium

A health research institute runs a cohort analysis agent whose data governance approval covers de-identified records only. A research lead asks for a tool that looks up identified patient records, so that the agent can resolve the roughly two queries a month where a study coordinator needs to confirm a participant's identity. The governance board will not extend the agent's approval to identified data, and coordinators already have a supervised re-identification process run by the data custodian. What should the architect recommend?

  • AAdd the lookup tool behind an approval workflow so the data custodian signs off on each identified lookup the agent makes
  • BAdd the lookup tool and write every identified lookup to an audit log that the governance board reviews at each meeting
  • CLeave the lookup tool out of the agent and route the rare identity checks through the custodian's existing supervised process Correct
  • DAdd the lookup tool with a system-prompt rule limiting its use to queries that a study coordinator has explicitly flagged
Do not add a capability outside an agent's approved role for a rare need; serve that need through an existing, appropriately governed path. Capability bloat often arrives one convenient tool at a time. A tool added for an occasional edge case widens what the agent can reach on every run, and here it would carry the agent past its data governance boundary. Approval workflows, logs and prompt rules only manage a capability that should not be present; keeping it out and using the custodian's supervised process preserves the boundary and the principle of data minimisation.

Why A is wrong: Custodian sign-off is tempting because it keeps a responsible person in the loop for each identified lookup. It is wrong because the agent's approval does not cover identified data at all, so granting the capability breaches the governance boundary regardless of who signs off, and it rebuilds a process that already exists.

Why B is wrong: An audit log reviewed by the board is tempting because it gives governance visibility of each use. It is wrong because it is a detective control: identified data has already flowed into an agent not approved to handle it before any review happens, which is precisely what the board refused to permit.

Why C is correct: Correct. The agent's role is defined by an approval for de-identified data, and the board has refused to widen it. Keeping the tool out holds the agent inside that boundary, and two queries a month are easily served by the supervised process coordinators already use, so no function is lost.

Why D is wrong: A usage rule is tempting because it seems to confine the tool to the rare legitimate case. It is wrong because a prompt instruction is not an enforceable boundary, the model can still call the tool on any query, and the agent would hold identified-data access its governance approval does not cover.

Solution Design & Architecture (17% of the exam)

Free sampleSolution Design & Architecturemedium

A retail bank's head of complaints asks for a customer-facing chatbot. Discovery shows that 18 per cent of complaints breach the regulator's five-business-day acknowledgement deadline, and case timing shows most of the delay sits in staff manually sorting each complaint into one of 12 regulatory categories before it can be routed. The team has eight weeks and must show the sponsor a measurable improvement in that breach rate. What should the architect recommend?

  • ABuild the customer-facing chatbot as requested, measured by how many complaints customers submit through it rather than by email each week
  • BFine-tune a model on five years of past complaints so it learns the 12 categories, and begin routing only once that training has finished
  • CDeploy an agent that investigates each complaint, decides the outcome and sends the final response letter to the customer without staff review
  • DUse Claude to classify each new complaint into the 12 categories and route it, measured by breach rate and by accuracy on a labelled sample Correct
Separate the outcome a sponsor needs from the mechanism they proposed, and aim the solution at the measured bottleneck with a matching success metric. The business problem is the acknowledgement breach rate, and the evidence places the delay in manual categorisation. Framing the solution as classification plus routing attacks that cause directly, fits an LLM's strength with unstructured text, and lets the team prove success with the breach rate and a labelled accuracy sample within the deadline. A chatbot answers the proposed mechanism rather than the problem.

Why A is wrong: It is tempting because it delivers exactly what the sponsor asked for. It is wrong because the measured delay sits in internal categorisation, not in how complaints arrive, so a new intake channel leaves the breach rate untouched and its success metric says nothing about the original problem.

Why B is wrong: It is tempting because historical labelled complaints look like ideal training data. It is wrong because fine-tuning comes before any prompt-based classification has been tried, adds data preparation and evaluation work that threatens the eight-week deadline, and may not beat a well-prompted model on a 12-category task.

Why C is wrong: It is tempting because it promises to clear the whole backlog at once. It is wrong because it is far larger than the stated problem, removes human judgement from a high-impact regulated decision, and the acknowledgement deadline only requires faster categorisation and routing.

Why D is correct: This targets the step the timing data identified as the bottleneck, uses a language model for a text classification task it suits, and ties success to the breach rate the sponsor cares about plus an accuracy check that catches misrouting.

Free sampleSolution Design & Architecturemedium

A hospital pharmacy team wants to flag paediatric prescriptions whose dose falls outside the weight-based range in its formulary table. The clinical safety officer requires that the same prescription and weight produce the same result every time, and that each flag can be traced to a specific table row during audit. A product owner proposes asking Claude to compute each safe range. What should the architect recommend?

  • ACompute each range in deterministic code against the formulary table, and limit Claude to drafting the plain-language alert text Correct
  • BHave Claude compute each range at a temperature of zero, since removing sampling randomness makes the calculation repeatable for audit
  • CHave Claude compute each range and route every flagged prescription to a pharmacist, so a human checks each calculation before release
  • DHave Claude compute each range with a confidence score, and send only those scored below a set threshold to a pharmacist for review
Assign rule-based calculations with reproducibility and audit requirements to deterministic code, and reserve generative components for tasks that tolerate variation. The safety officer's requirements are identical outputs for identical inputs and traceability to a table row. Deterministic code satisfies both by design, whereas any generative calculation, however constrained or reviewed, only approximates them. Splitting the system into a deterministic dose check and a generative alert-wording step uses each component where it fits.

Why A is correct: A table lookup and arithmetic in code is reproducible by construction and every flag points to the row that produced it, meeting both stated requirements. Claude is kept for the generative part, wording the alert, where variation does no harm.

Why B is wrong: It is tempting because lowering temperature does make outputs more consistent. It is wrong because it does not guarantee identical results across calls, and a generated calculation still cannot be traced to a specific formulary row in the way the safety officer requires.

Why C is wrong: It is tempting because human review feels like the safe choice in a clinical setting. It is wrong because it is a compensating control on a component that should not be generative at all, it misses unflagged prescriptions where the model erred, and it still fails the reproducibility requirement.

Why D is wrong: It is tempting because it appears to focus human effort where the risk is. It is wrong because a model's self-reported confidence is not a reliable accuracy signal, so confidently wrong ranges pass unchecked, and the approach still fails reproducibility and traceability.

Free sampleSolution Design & Architecturemedium

A software company's support director wants Claude deployed in its help centre and describes success as customers feeling better supported. Ticket data shows that 41 per cent of escalations to the tier 2 engineering queue are configuration questions already answered in the product documentation, and each escalation costs about 35 minutes of engineer time. The finance director will fund a second phase only on evidence tied to that problem. Which success criterion should the architect set before build begins?

  • AThe average satisfaction score from a survey shown after each assistant conversation, collected from launch and reported to finance monthly
  • BThe change in the tier 2 escalation rate for configuration questions against a pre-launch baseline, with accuracy checked on held-out tickets Correct
  • CThe number of conversations the assistant handles each week, on the basis that every handled conversation is one fewer ticket for the team
  • DThe assistant's own confidence rating on each answer, averaged weekly, with a target that at least 90 per cent of answers rate as high
Define success as a measurable change in the original business problem against a baseline, paired with a correctness check. A success criterion has to be traceable to the problem that justified the project. Here that is configuration-question escalations and the engineer time they consume. Capturing the escalation rate before launch makes any later change attributable, and checking answer accuracy on held-out tickets prevents the metric being met by an assistant that deflects users with wrong answers.

Why A is wrong: It is tempting because it matches the director's description of success. It is wrong because it has no pre-launch baseline, does not measure escalations or engineer time, and survey responses are a self-selected sample that finance cannot connect to the stated cost.

Why B is correct: This measures the exact problem the ticket data exposed, compares it with a baseline captured before launch so the change can be attributed, and pairs it with an accuracy check so deflection is not achieved by giving wrong answers.

Why C is wrong: It is tempting because volume is easy to count and rises quickly. It is wrong because a handled conversation is not a resolved one, so the count can grow while escalations stay flat, and it says nothing about correctness.

Why D is wrong: It is tempting because it looks like a built-in quality metric. It is wrong because a model's self-reported confidence is not a reliable measure of accuracy and has no link to the escalation cost finance wants evidence on.

Evaluation, Testing & Optimization (16% of the exam)

Free sampleEvaluation, Testing & Optimizationmedium

A hospital group is piloting Claude to draft discharge summaries from inpatient records, and the clinical safety officer has set one requirement in writing: no summary may leave out a documented drug allergy or a medication change. The team's evaluation plan scores each draft against a clinician-written reference summary using an overall similarity score, and the pilot average is 0.87. About one record in six carries an allergy or a medication change. What should the architect recommend the team measure before the pilot widens?

  • AThe overall similarity score, with its pass threshold raised from 0.87 to 0.92 across the pilot
  • BThe share of summaries the model itself rates as complete, from a self-check added to each draft
  • CRecall of documented allergies and medication changes per summary, on a labelled set, reported apart Correct
  • DClinician satisfaction with each draft, as a five-point rating from the ward doctors who sign it
Derive the metric from the stated requirement, measuring recall on safety-critical items separately so an aggregate quality score cannot hide omissions. The requirement is defined by a failure on specific items, so the metric has to count those items. An overall similarity score averages over all text and over records that contain no allergy or medication change, which lets a 0.87 mean coexist with missed allergies. Item-level recall on a labelled set, with a threshold tied to the safety officer's requirement and reported separately, exposes exactly the failure the requirement forbids.

Why A is wrong: Raising the bar is tempting because it looks like a stricter safety gate. It is wrong because the similarity score averages over the whole narrative and over the five in six records with no allergy or medication change, so a dropped allergy line barely moves it and a higher threshold still does not target the stated requirement.

Why B is wrong: A self-check is tempting because it is automatic and cheap to run on every draft. It is wrong because a model's own judgement of completeness is not a ground-truth accuracy signal; the same model that omitted an allergy can rate its draft complete.

Why C is correct: The requirement names a specific failure, the omission of an allergy or medication change, so the metric must count those items: recall against records labelled with them, reported separately from any aggregate so a good average cannot hide a miss.

Why D is wrong: Clinician ratings are tempting because the raters are the domain experts. It is wrong because a satisfaction score reflects overall impression and readability; a busy signer does not check every line against the record, so omissions go unmeasured.

Free sampleEvaluation, Testing & Optimizationmedium

A card issuer adds a Claude step that writes a short reason code for each transaction its fraud model holds. The checkout integration requires that 99 per cent of held transactions receive a reason within 1.5 seconds, measured from the moment the payment gateway sends the request. The step calls two internal lookup services before the model, and the team's dashboard reports a mean model response time of 0.6 seconds, taken from the orchestration service's own logs. What should the architect recommend as the latency metric for this requirement?

  • AMean model response time from the orchestration logs, alerting if it rises above 1.0 seconds
  • BOutput tokens per second from the model, tracked daily against the figure seen in the pilot
  • Cp99 model response time on a 50-request offline benchmark run in a quiet test environment
  • Dp99 end-to-end time measured from the gateway's send to its receipt, under production load Correct
Define a latency metric with the same percentile, start point and end point as the requirement, rather than a convenient mean from one component. A latency requirement is only testable when the metric matches its percentile and its measurement boundary. Here the business promises 99 per cent within 1.5 seconds from the gateway's point of view, so a component mean from orchestration logs cannot confirm or refute it: the mean hides the tail and the boundary omits the lookups and network. Measuring p99 end to end under production load makes the metric and the requirement the same quantity.

Why A is wrong: This is tempting because the figure is already collected and looks comfortably inside budget. It is wrong because a mean hides the slow tail the 99 per cent requirement is about, and the orchestration logs exclude the lookup calls and network time that the gateway experiences.

Why B is wrong: Throughput is tempting because it is a common way to describe model speed. It is wrong because it measures generation rate, not how long the gateway waits; it ignores queueing, the two lookups and time before the first token.

Why C is wrong: This is tempting because it uses the right percentile. It is wrong because 50 requests cannot estimate a 99th percentile with any confidence, a quiet environment removes production contention, and timing only the model call leaves out the lookups.

Why D is correct: The requirement states a percentile and a boundary, so the metric must use the same percentile and the same boundary: from the gateway's send to its receipt, under real load, including the lookups and network time that a model-only timer misses.

Free sampleEvaluation, Testing & Optimizationmedium

A software company's customer-facing assistant answers questions from files that each customer uploads to their own workspace, and it can call a tool that opens support tickets on the customer's behalf. The launch security requirement is that text inside an uploaded file must not cause the assistant to open or change a ticket, or to quote content from outside that workspace. The current evaluation suite holds 400 benign questions and reports 93 per cent answer accuracy. What should the architect add to the evaluation before launch?

  • AAn adversarial set of uploaded files with injected instructions, scored on unauthorised actions or leaks Correct
  • BA larger benign question set of 2,000 items, so the accuracy figure has a tighter confidence interval
  • CA weekly count of how often the assistant refuses a request in production, used as the security metric
  • DA system-prompt rule forbidding instructions from files, scored by asking the assistant if it complied
Give each security requirement its own adversarial test set and violation-rate metric, because accuracy on benign inputs says nothing about behaviour under attack. An evaluation can only measure behaviour its inputs provoke. A suite of benign questions never presents an injected instruction, so its 93 per cent accuracy is silent on the launch requirement. Building files that carry injected instructions and scoring the observed tool calls and disclosures gives a violation rate that maps directly onto the requirement, and keeping it separate from accuracy stops a strong accuracy figure from masking a security failure.

Why A is correct: The requirement concerns behaviour under attack, which a benign suite never exercises. A dedicated set of files carrying injected instructions, scored on the rate of unauthorised tool calls or cross-workspace disclosures against a target tied to the requirement, measures the security dimension directly and separately from accuracy.

Why B is wrong: This is tempting because more items give a statistically firmer number. It is wrong because a bigger benign set still contains no injection attempts, so the security requirement remains entirely unmeasured however precise the accuracy figure becomes.

Why C is wrong: Refusals are tempting because they look like evidence of caution. It is wrong because a refusal count has no ground truth: it can rise through over-refusal while successful injections still happen, and measuring in production after launch comes too late for a launch requirement.

Why D is wrong: A prompt rule is tempting because it is quick to add. It is wrong because it is a weak preventive instruction, not a measurement, and the assistant's own report of compliance is not evidence; only its actual tool calls and outputs under attack can show whether the requirement holds.

Governance, Safety & Risk Management (14% of the exam)

Free sampleGovernance, Safety & Risk Managementmedium

A retail bank gives its 600 relationship managers an assistant that looks up client accounts through a tool. The model supplies the account number as a tool argument, and the tool handler calls the core banking service with a shared service account that can read every client. The core banking service already holds each manager's book and enforces it when called with that manager's own credential. A conduct review found 14 lookups in one month of clients outside the requesting manager's own book, each after the manager typed a name that matched a different client. The regulator requires that a manager be prevented, not merely detected, from viewing a client outside their book, and managers must keep looking clients up by name. Which TWO changes meet the requirement? Select TWO.

  • AAdd the manager's list of client account numbers to the system prompt and instruct the model to refuse lookups of any other account.
  • BHave the tool handler check each requested account against the signed-in manager's book, taken from the session, and reject any mismatch. Correct
  • CLog every lookup with the manager's identity and send compliance a daily report of the accounts accessed outside the manager's book.
  • DMake the assistant ask the manager to confirm the client's full name and date of birth before it calls the lookup tool for any account.
  • EReplace the shared service account with a delegated credential for the signed-in manager, so the banking service applies their access rights. Correct
Enforce data-access scope from the authenticated session and delegated identity, never from model-supplied arguments or a shared all-access credential. The model chooses the account number, so any control that relies on the model choosing correctly can fail. Attaching the manager's identity from the session and checking the requested account against their book in the tool handler, and calling the banking service with the manager's delegated credential instead of an all-access service account, both put the boundary in code and systems the model cannot argue with. Lookup by name still works; only the out-of-book result is refused.

Why A is wrong: Tempting because it gives the model the information it needs to stay inside the book. It is wrong because the restriction still depends on the model obeying an instruction, and a name collision or a crafted message can lead it to request another account; the regulator asked for prevention, which an instruction cannot guarantee.

Why B is correct: Correct. The authorisation decision moves out of the model and into code that attaches the manager's identity server-side from the session, so an account outside the book is refused whatever argument the model supplies, while lookup by name keeps working.

Why C is wrong: Tempting because an audit trail is useful and conduct teams expect one. It is wrong because logging and reporting detect a breach after the client data has been shown, and the regulator explicitly requires prevention rather than detection.

Why D is wrong: Tempting because a confirmation step would catch some mistaken name matches. It is wrong because the manager making the request is the person confirming, the step is still carried out by the model, and nothing stops a confirmed lookup of a client outside the book.

Why E is correct: Correct. A shared credential that can read every client gives the tool far more reach than any one manager holds; calling the banking service as the manager means its own entitlement checks refuse out-of-book accounts, enforcing the boundary outside the model.

Free sampleGovernance, Safety & Risk Managementmedium

A national tax agency's online assistant answers questions about a citizen's own filings through a lookup tool whose input includes a taxpayer identifier that the model fills in. In a red-team exercise, testers signed in as one citizen and claimed to be acting for a relative, and in 7 of 40 attempts the assistant retrieved the relative's filing. The agency requires that a signed-in citizen can retrieve their own records alone, and the assistant must keep answering filing questions without staff involvement. What should the architect recommend?

  • AStrengthen the system prompt to say the assistant must look up the signed-in citizen's identifier and refuse others
  • BScreen incoming messages with a classifier that blocks requests mentioning a relative, an agent or another person
  • CHave the tool handler take the taxpayer identifier from the authenticated session and ignore any model-supplied value Correct
  • DLog each lookup with both identifiers and have the fraud team review mismatched pairs in a weekly audit report
Attach tenant or account scope to a tool call server-side from the authenticated session, never from an identifier the model supplies. The defect is that the model chooses whose record is fetched, so anything that persuades the model also widens access. Moving the identifier to the handler and binding it to the authenticated session removes that choice from the model entirely, which turns a probabilistic behaviour into an enforced boundary. Prompt wording and input classifiers lower the attack rate but leave the same capability in place, and audit logs only find the breach afterwards.

Why A is wrong: This is tempting because the failures came from the model being talked into a different identifier, and a firmer instruction would reduce that. It is wrong because the model still controls the identifier, so a sufficiently persuasive request can still succeed; an instruction is not an access control.

Why B is wrong: Input screening is a genuine layer and would catch the red team's phrasing. It is wrong because attackers can rephrase around the classifier, it blocks legitimate questions that mention family, and the identifier is still taken from model output.

Why C is correct: This is correct because the record scope is attached server-side from the session, so no wording in the conversation can change which taxpayer is looked up. The assistant keeps answering filing questions automatically, which meets the no-staff constraint.

Why D is wrong: Audit logging is valuable and an assessor would expect it. It is wrong as the answer because it detects a disclosure days after it has happened rather than preventing it, and it leaves the model in control of whose record is fetched.

Free sampleGovernance, Safety & Risk Managementmedium

A SaaS provider's customer support assistant answers questions by retrieving passages from 3,000 internal runbooks written by 40 engineering teams. An audit of 20,000 replies found 6 that quoted an internal hostname and 2 that quoted a live API key copied from a runbook. Security requires that no credential or internal hostname reaches a customer. The runbooks cannot be rewritten this quarter, and the assistant must keep answering from them because it resolves 38 per cent of tickets without a human agent. Which TWO changes should the architect recommend? Select TWO.

  • AInstruct the assistant in its system prompt to treat runbook credentials and hostnames as confidential and to leave them out of replies.
  • BRemove the runbooks from the retrieval index and answer from the public help centre articles until each team has cleaned its documents.
  • CRedact credentials and internal hostnames from runbook passages in the retrieval pipeline, before any passage enters the model's context. Correct
  • DScan every drafted reply in code for credential patterns and internal hostnames, and block or redact any match before it is sent. Correct
  • EAsk the assistant to end each reply with a statement confirming it holds no secrets, and release the replies that carry that statement.
Keep secrets out of the model's context at retrieval and filter outputs in code, rather than trusting instructions or the model's own assurance. Two independent controls in code cover the leak path from both ends. Redacting secrets in the retrieval pipeline removes them before the model can see them, without the manual runbook rewrite the quarter does not allow. A deterministic scan of each reply then catches any residue before it reaches the customer. Neither depends on the model following an instruction, and both keep the runbook-backed answers that drive the self-service resolution rate.

Why A is wrong: Tempting because it is quick to ship and will reduce the rate of leaks. It is wrong because the secrets remain in the model's context and the instruction is a request, not a guarantee; security asked that no credential reaches a customer, which needs a control enforced in code.

Why B is wrong: Tempting because it eliminates the source of the leak outright. It is wrong because it trades away the stated requirement that the assistant keep answering from the runbooks, putting the 38 per cent self-service resolution rate at risk for the whole quarter.

Why C is correct: Correct. Automated redaction at retrieval applies data minimisation without editing the source runbooks: a secret that never enters the context cannot be quoted, and the assistant keeps answering from the same corpus.

Why D is correct: Correct. A deterministic output filter is the last line of defence for anything the retrieval redaction misses, such as a hostname written in an unusual form, and it is enforced outside the model so it holds regardless of what the model generates.

Why E is wrong: Tempting because it looks like a gate on every reply. It is wrong because it trusts the model's self-report; the same model that quoted the key can append the confirmation, so the check adds no independent assurance.

Stakeholder Communication & Lifecycle Management (14% of the exam)

Free sampleStakeholder Communication & Lifecycle Managementmedium

A retail bank wants Claude to draft replies to customer complaints. Discovery has established that the team handles about 4,000 complaints a month, that the complaints policy library holds 2,300 documents which the compliance team revises weekly, and that the regulator requires each reply to cite the policy clause, and its version, that was in force on the date of the complaint. The sponsor proposes fine-tuning a model on two years of approved replies. What should the architect recommend?

  • AFine-tune a model on the two years of approved replies so policy positions and house style are carried in the model weights.
  • BLoad the whole policy library into the context window on each request so that retrieval cannot miss a relevant policy clause.
  • CPrompt the model with a summary of current policy positions, and have a reviewer add the clause citations before each reply is sent.
  • DRetrieve clauses from a versioned policy index filtered by complaint date, and have each draft cite the retrieved clause and version. Correct
When discovery reveals frequently revised source material and a version-specific citation duty, the architecture should retrieve from a versioned index rather than fine-tune. Discovery answers about data freshness and traceability decide the architecture. Fine-tuning bakes knowledge into weights that go stale on every revision and cannot attribute an answer to a source. Retrieval keeps knowledge outside the model, so a weekly revision is live once indexed, a date filter selects the version in force on the complaint date, and the retrieved text is what the reply cites to satisfy the regulator.

Why A is wrong: This is tempting because the sponsor proposed it and two years of approved replies look like ideal training data for tone and policy. It is wrong because a weekly revision cycle would leave the weights stale within days, and a fine-tuned model cannot point to which clause version it relied on, so the citation requirement cannot be met.

Why B is wrong: This is tempting because it removes retrieval misses as a failure mode. It is wrong because 2,300 documents on each of 4,000 requests a month is costly and slow, the model still has to pick the version in force on the complaint date from a mass of competing versions, and long-context stuffing is the over-engineered answer where targeted retrieval fits.

Why C is wrong: This is tempting because it keeps a human in the loop on a regulated output and avoids building an index. It is wrong because a summary of current positions cannot reflect the version in force on an older complaint date, and moving the citation duty to a manual step is a compensating control that adds cost to every reply where a structural fix exists.

Why D is correct: Correct. The discovered facts (weekly revisions, a large library, and a date-specific citation duty) all point to retrieval over a versioned index: updates reach answers as soon as they are indexed, the date filter selects the version in force, and the retrieved clause gives the draft something concrete to cite.

Free sampleStakeholder Communication & Lifecycle Managementmedium

An Irish hospital group wants Claude to turn discharge notes into summary letters for family doctors. Discovery has captured the stakeholders, the current process (a junior doctor spends about 25 minutes per letter) and a success criterion of a reviewed letter within an hour of discharge. The architect must now choose between calling a provider-hosted model endpoint and deploying through the group's existing cloud tenancy, and the information governance team has not yet been consulted. Which question must discovery answer before that choice is made?

  • AWhether patient-identifiable notes may be processed outside the group's own boundary, and under which data processing agreement. Correct
  • BHow many discharge letters each ward produces on a typical day, so the team can size throughput for the hosted endpoint.
  • CWhich layout and tone family doctors prefer in a discharge letter, so the prompt template can match the group's house format.
  • DWhether clinicians would accept a confidence score shown on each letter, so low-scoring letters can be routed for closer review.
Data residency and processing-agreement constraints for sensitive data must be gathered in discovery before choosing a deployment path, because they decide which paths are permitted. The two candidate designs differ mainly in where patient data travels and who processes it. Throughput, letter format and review routing can all be handled on either path, so they do not decide between them. Whether identifiable health data may leave the group's authorised boundary, and whether the required processing agreement is in place before any data is sent, determines which path is allowed at all.

Why A is correct: Correct. Where patient-identifiable data may be processed, and under what agreement, is a compliance constraint that rules deployment paths in or out. Until information governance answers it, choosing between a provider-hosted endpoint and the group's own tenancy is designing before a binding requirement is known.

Why B is wrong: This is tempting because volume is a standard discovery item and affects capacity planning. It is wrong because either deployment path can be sized for a hospital's daily discharge volume, so the answer does not separate the two options; it assumes the hosted path is already permitted.

Why C is wrong: This is tempting because the doctors receiving the letters are key stakeholders and their acceptance drives adoption. It is wrong because letter format is settled in the prompt and evaluation work on either path; it has no bearing on where patient data may be processed.

Why D is wrong: This is tempting because it looks like a sensible way to target clinical review. It is wrong because it does not bear on the deployment decision, and a model's self-reported confidence is not a reliable accuracy signal to route clinical review on.

Free sampleStakeholder Communication & Lifecycle Managementmedium

A managed hosting company wants a Claude-based agent to triage infrastructure alerts and run remediation runbooks. Discovery found that on-call engineers currently act on every alert, that the target is to cut mean time to acknowledge from 18 minutes to 5, and that runbooks range from restarting a service, which is reversible, to failing over a database or deleting storage volumes, which are not. The head of reliability states that a single wrong irreversible action would be a reportable customer incident. What should the architect recommend?

  • ARequire an on-call engineer to approve each runbook the agent proposes, including service restarts, before it is executed.
  • BLet the agent run reversible runbooks automatically, and require on-call approval before it runs any irreversible runbook. Correct
  • CLet the agent run all runbooks automatically, and record each action in an audit log that reliability reviews each morning.
  • DLet the agent act alone when its self-reported confidence is above 90 per cent, and send the remaining alerts to on-call.
Discovering which outputs trigger irreversible actions, and the tolerance for error on each, lets the architect place human approval only where it is required. Who acts on the output and what happens when it is wrong are discovery answers with direct architectural consequences. Here they split the action space into reversible and irreversible classes with different error tolerances. Gating approval by action class satisfies both the latency target and the zero-tolerance requirement, whereas uniform approval breaks latency, after-the-fact logging breaks the tolerance, and a confidence threshold relies on an uncalibrated signal.

Why A is wrong: This is tempting because it is the most cautious design and no action ever happens without a person. It is wrong because waiting for approval on routine reversible restarts keeps engineers in the loop on every alert, which defeats the 5 minute target the discovery work established.

Why B is correct: Correct. Discovery classified the actions by reversibility and stated the tolerance for error on each class. Automating the reversible runbooks meets the 5 minute acknowledgement target, while a hard approval gate on irreversible runbooks places human review exactly where a single mistake is unacceptable.

Why C is wrong: This is tempting because it gives the fastest response and an audit trail for accountability. It is wrong because a morning review detects an irreversible failover or deletion hours after the damage is done; logging is a compensating control where the stated tolerance requires prevention.

Why D is wrong: This is tempting because a threshold appears to tune automation against risk. It is wrong because a model's self-reported confidence is not a calibrated accuracy signal, and the gate ignores reversibility, so a confident but wrong volume deletion would still run unreviewed.

Claude Models, Prompting & Context Engineering (13% of the exam)

Free sampleClaude Models, Prompting & Context Engineeringmedium

A software company runs two Claude-based review paths. An editor plug-in suggests fixes as developers type and must return its first token within 1.5 seconds. A nightly job reviews every merged pull request that touches concurrency code, and engineers read its findings the next morning. On 200 held-out concurrency pull requests, enabling extended thinking on the nightly reviewer raised the share of seeded defects found from 58 to 81 per cent and added about 40 seconds per review, and the engineering director has approved the higher cost for that job. Where should extended thinking be enabled?

  • AOn both paths, so that developers and the nightly job receive reasoning of the same depth.
  • BOn the editor plug-in only, because interactive suggestions reach developers the soonest.
  • COn the nightly concurrency review only, leaving the editor plug-in to answer without it. Correct
  • DOn neither path, moving the nightly job to the fastest tier to offset the cost of the reviews.
Enable extended thinking where measured reasoning gains matter and latency is slack, and keep it off paths bound by a tight interactive latency budget. Extended thinking lets the model spend additional tokens reasoning before it produces the final answer, which improves results on multi-step problems such as concurrency defects but lengthens time to the answer and raises token spend. The decision is therefore made per path: an offline job with an approved budget and a measured quality gain is where the trade pays, while an interactive path with a 1.5 second first-token budget is where it does not.

Why A is wrong: Consistency across tools is an appealing principle, and the nightly gain suggests the plug-in might improve too. It is wrong because extended thinking adds generation time before the answer, which the 1.5 second first-token budget on the editor path cannot absorb, and the plug-in's quick fix suggestions were never shown to need deeper reasoning.

Why B is wrong: It is tempting to put the best reasoning where users see it first. It is wrong on both counts: the plug-in has the tight latency budget that extended thinking would break, and the measured 23-point gain in defects found belongs to the nightly concurrency review, which this option leaves without it.

Why C is correct: This is correct because extended thinking earns its latency where a task needs multi-step reasoning and the consumer is not waiting: the nightly review has no one reading it until morning, a measured gain from 58 to 81 per cent, and approved cost. The editor plug-in keeps its 1.5 second budget by answering directly.

Why D is wrong: Cost reduction is a reasonable instinct for a job that runs on every merged pull request. It is wrong because the director has already approved the higher cost, and dropping both extended thinking and model capability on hard concurrency reasoning gives up the measured improvement in defect detection that the job exists to deliver.

Free sampleClaude Models, Prompting & Context Engineeringmedium

A national records office answers freedom of information requests within a statutory deadline. For each request, a coordinator agent splits the request into search tasks, subagents each screen a batch of about 300 documents for relevance, and the coordinator then decides which statutory exemptions apply and drafts the reasoning. A trial on 150 closed requests found the fastest tier agreed with archivists on relevance screening 96 per cent of the time but chose the correct exemptions only 61 per cent of the time, while the most capable tier reached 93 per cent on exemptions. The office requires 90 per cent agreement on exemptions, and running every step on the most capable tier exceeds its cost ceiling per request. Which design best meets the requirement?

  • AFastest tier for every step, with an instruction asking the coordinator to reason with more care.
  • BMost capable tier for the screening subagents, fastest tier for the coordinator to keep drafts quick.
  • CMid tier for every step, so that cost and exemption accuracy both fall between the two results.
  • DMost capable tier for the coordinator's exemption decisions, fastest tier for screening subagents. Correct
Within one multi-agent system, route high-volume simple subtasks to a fast tier and reserve the most capable tier for the reasoning step that needs it. Model selection does not have to be a single choice per system. When a workflow mixes many simple, high-volume calls with a few calls that need complex judgement, assigning tiers per step lets each step meet its own measured requirement. Because document screening dominates the number of calls, moving it to the fastest tier is what makes room in the cost ceiling for the most capable tier on the exemption decisions.

Why A is wrong: This is tempting because it keeps cost at its lowest and prompt changes are cheap to try. It is wrong because a 61 per cent result on applying statutory exemptions is a capability gap on complex reasoning, not a missing instruction, and nothing in the trial suggests a wording change would close a 29-point shortfall against the requirement.

Why B is wrong: It can seem sensible to put the strongest model on the step that touches every document. It is wrong because it inverts the measured fit: screening already meets the bar on the fastest tier, while the exemption decisions that failed at 61 per cent stay on that tier, and paying the higher rate across the high-volume step raises cost rather than lowering it.

Why C is wrong: A uniform middle tier is a common compromise when two figures pull in opposite directions. It is wrong because the mid tier's exemption accuracy was never measured, so there is no evidence it clears the 90 per cent requirement, and it still pays more than necessary for screening work the fastest tier already handles.

Why D is correct: This is correct because it matches each tier to the work it was measured on: high-volume relevance screening, where the fastest tier already agrees with archivists 96 per cent of the time, and the low-volume legal reasoning on exemptions, where only the most capable tier clears the 90 per cent requirement. Most document-level calls move to the cheaper tier, which is what brings the request back under the cost ceiling.

Free sampleClaude Models, Prompting & Context Engineeringmedium

A pizza chain is replacing its phone ordering line with a voice assistant. Callers expect a reply within about one second of finishing a sentence, and the script already includes a pause of up to four seconds when the assistant says it is checking the order before confirming it. On 400 recorded calls, the fastest tier kept every conversational turn under one second but misread 12 per cent of orders with several modifications, such as half-and-half toppings; the mid tier averaged 1.6 seconds per turn and misread 4 per cent; the most capable tier misread 1 per cent at 2.5 seconds per turn. Head office will not accept a design that relies on callers spotting mistakes. Which design best meets the requirement?

  • AFastest tier for conversational turns, most capable tier to parse the full order during the check pause. Correct
  • BMost capable tier on every turn, playing a short holding phrase to cover the longer response time.
  • CFastest tier on every turn, with a prompt telling it to read the whole order back before confirming.
  • DMid tier on every turn as a compromise, accepting slower replies in exchange for fewer misread orders.
Route stages of one interaction to different tiers when each stage has a different latency budget and a different reasoning demand. Latency budgets often differ between stages of the same interaction. Here every conversational turn has a one second budget that only the fastest tier meets, while the confirmation step already carries a four second pause that can absorb the most capable tier's response time. Placing the demanding parsing work in that slack lets the system meet both the turn-taking budget and the accuracy requirement, which no single tier achieves.

Why A is correct: This is correct because it routes each stage to the tier that fits its constraint: the fastest tier keeps conversational turns under one second, and the most capable tier does the hard multi-modification parsing inside the existing four second checking pause, where its 2.5 second response fits and its 1 per cent misread rate removes reliance on callers.

Why B is wrong: This is tempting because it gives the best order accuracy measured and the holding phrase seems to hide the delay. It is wrong because a 2.5 second wait after every sentence breaks the one second conversational budget, and filler audio on every turn makes the call feel slow rather than fixing the latency.

Why C is wrong: Reading the order back is a good conversational habit and keeps every turn fast. It is wrong because the read-back is a compensating control that only works if callers notice the error, which head office has ruled out, and the final order is still parsed by the tier that misreads 12 per cent of complex orders.

Why D is wrong: A single middle tier looks like a sensible balance between the two extremes. It is wrong because 1.6 seconds per turn still breaks the one second budget on every sentence, and a 4 per cent misread rate on complex orders still leaves errors that only callers could catch.

Developer Productivity & Operational Enablement (7% of the exam)

Free sampleDeveloper Productivity & Operational Enablementmedium

A state revenue agency is rolling Claude Code out to 140 engineers who work across 40 repositories. Its information security branch requires that Claude Code's built-in web search and web fetch tools be unavailable in every session on agency machines, that neither a repository's maintainers nor an individual engineer can remove that restriction, and that it be in force on the first day of rollout without a change to each of the 40 repositories. What should the architect recommend?

  • ACommit a deny rule for web tools to each repository's shared project settings and require code review on any change to it
  • BAdd an instruction to every repository's CLAUDE.md telling Claude Code that use of the web search and fetch tools is prohibited
  • CAllow the web tools, log every fetched address to a central store, and have the security branch review that log each week
  • DDeploy a centrally managed policy to agency machines that denies the web tools and takes precedence over project and user settings Correct
When a tooling restriction must hold across every repository and resist removal by maintainers or engineers, enforce it through centrally managed policy rather than per-repository settings. Claude Code merges settings from several layers, and a centrally managed policy deployed by administrators takes precedence over both the project settings committed to a repository and each engineer's personal settings. That makes it the only layer here that repository maintainers and individual engineers cannot override, and because it is installed on machines it covers all 40 repositories at once with no repository changes. Project settings remain the right place for team policy that a repository's own maintainers are trusted to own.

Why A is wrong: This is tempting because committed project settings are shared and reviewable, which is the right home for most team policy. It fails two stated constraints: it needs a change in all 40 repositories before rollout, and a repository's maintainers can approve a later commit that removes the rule.

Why B is wrong: Shared project instructions do reach every session, which makes this look like a team-wide control. An instruction is guidance to the model rather than an enforced permission, so it is a weak sole control for a security requirement, and it still needs edits to all 40 repositories that maintainers could later revert.

Why C is wrong: Central logging gives the security branch visibility, which feels like governance. It is a compensating control after the fact: the requirement is that the tools be unavailable, and a weekly review only detects use that has already happened.

Why D is correct: An administrator-managed policy sits above project and user settings in Claude Code's precedence order, so no repository commit or personal setting can loosen it. Because it is deployed to machines rather than to repositories, it applies to all 40 repositories from day one without touching any of them.

Free sampleDeveloper Productivity & Operational Enablementmedium

A government digital service team of 12 engineers configured Claude Code to ask for approval before every file edit and every shell command on its benefits-claims repository. Each engineer now approves around 180 prompts a day, the median approval takes under a second, and a review found two production database migrations that were approved without being read. The service owner requires that approval continue for actions that are hard to reverse, such as pushing, deploying and running migrations, and that the policy be the same for every engineer. What should the architect recommend?

  • ACommit a shared policy that pre-approves edits and the test runner, and keeps approval on pushes, deploys and migrations Correct
  • BSwitch the team to a mode that skips permission prompts, and rely on pull request review to catch any mistakes
  • CKeep approval on every action, and add a session log so reviewers can audit which approvals were granted each week
  • DLet each engineer build their own allow list of routine commands in personal settings as repeated prompts appear
Place human approval on irreversible, high-impact actions and pre-approve routine ones in a shared policy, so approval fatigue does not erode the gate that matters. Approving every action made the prompts so frequent that approval became reflexive, which is how two migrations went through unread. Moving routine, easily reversed actions such as file edits and test runs into a shared pre-approved set leaves prompts only for pushes, deploys and migrations, where a human check is worth its cost. Defining that policy in committed project settings keeps it consistent across the team and subject to review.

Why A is correct: Pre-approving the routine, reversible actions removes most of the daily prompts, so the ones that remain are the hard-to-reverse actions the service owner named and are more likely to be read. Committing the policy to the repository makes it identical for all 12 engineers and reviewable when it changes.

Why B is wrong: This removes the approval fatigue at once, and pull request review is a real control. It removes the gate from deploys and migrations too, which act on production before any pull request is reviewed, so it trades away the stated requirement.

Why C is wrong: A log adds accountability, which seems to answer the unread migrations. It is a compensating control that leaves the cause in place: 180 prompts a day still train engineers to approve without reading, so the high-impact prompts stay buried.

Why D is wrong: Personal allow lists do cut the prompt count for each engineer, which is the visible symptom. They produce 12 different policies, breaking the same-for-everyone requirement, and nothing stops an engineer from allowing a deploy or migration command.

Free sampleDeveloper Productivity & Operational Enablementmedium

An education technology company maintains its learning-analytics product on two long-lived branches: main, and a release branch that keeps an older test command and a different database migration tool for a school district contract. New engineers spend about three days setting up Claude Code from a wiki page, and Claude Code regularly runs the main-branch test command during release-branch work. The engineering manager wants a new hire's first session to work correctly on either branch with no manual setup. What should the architect recommend?

  • ABuild a standard laptop image that installs the main-branch instructions and settings into each user-level configuration
  • BCommit the shared instructions, permissions and tool definitions to each branch, so the configuration checked out matches the code Correct
  • CKeep the wiki page, add a section for the release branch, and ask engineers to tell Claude which branch they are working on
  • DWrite a setup script that copies the current branch's settings into each engineer's personal configuration when run
Commit team configuration with the code so it is versioned per branch and a fresh clone gives a working, branch-correct setup. The defect is that configuration and code are versioned separately: anything held at user level, whether installed by an image or a script, stays the same when the checked-out branch changes. Committing the shared instructions, permission policy and tool definitions to the repository puts them under the same version control as the code, so each branch carries its own test command and migration instructions and a new hire gets the right setup from the clone itself.

Why A is wrong: A standard image removes most of the three-day setup, which matches the onboarding goal. User-level configuration is the same whichever branch is checked out, so release-branch work would still receive the main-branch test command.

Why B is correct: Project configuration committed to the repository is versioned with the code, so checking out the release branch also checks out its own test command and migration tool instructions. A new hire gets a working setup simply by cloning, with nothing to install by hand.

Why C is wrong: Better documentation addresses the confusion between branches in a low-cost way. It keeps the manual setup the manager wants removed, and it relies on every engineer restating the branch in each session, which is the kind of step people forget.

Why D is wrong: A script automates setup and can read branch-specific values, which looks branch-aware. The copied personal settings do not change when the engineer switches branch unless the script is run again, so the wrong test command returns after the first switch.

Want the full bank?

318 CCAR-P questions, every one with an explanation of why every option is right or wrong. No sign-up to start.

Practise CCAR-P free

Frequently asked questions

Are these CCAR-P practice questions free?

Yes. Every CCAR-P question on this page is free to read with no sign-up, and each one explains why the right answer is right and why every other option is wrong. The full bank of 318 questions is on Examworthy.

Do the questions explain why the wrong answers are wrong?

Yes, and that is the point. Each option, correct or not, has its own rationale, so you learn to rule out the tempting wrong answer, not just recognise the right one. That is the reasoning the CCAR-P tests.

Are these real CCAR-P exam questions?

No. These are original, blueprint-aligned practice questions written to the public Anthropic content outline. We never reproduce live exam items. They mirror the format and difficulty of the real exam.

How many questions are on the real CCAR-P?

The CCAR-P is 63 questions in 120 minutes, with a pass mark of 720 / 1000. For the full domain-by-domain breakdown and a study plan, read the study guide.

Examworthy is not affiliated with or endorsed by Anthropic. All questions are original, blueprint-aligned practice material. We never reproduce live exam items. CCAR-P and related marks belong to their respective owners.