How to pass Claude Certified Architect - Professional (CCAR-P)
25 min read7 domains coveredFree practice, no sign-up
Claude Certified Architect - Professional (CCAR-P) is an exam for the person who owns a Claude system end to end. It assumes you can already build with Claude, and asks a harder question: given a business problem, a set of constraints and several stakeholders who want different things, which design decision is right, and why does each alternative fall short. The seven domains run from framing the problem and choosing an architectural pattern, through integration, retrieval and evaluation, to governance, regulatory fit, stakeholder communication and the long tail of operating a system after handoff.
The items are short applied situations. A hospital agent holds tools its role never calls. A retrieval assistant starts answering wrongly the week after a document refresh. A team wants a larger model when the real fault is a stale index. Each stem describes a working system under pressure and asks for the single best move. Several options will be sensible engineering in the abstract; only one answers the requirement as stated, at the cost and risk the scenario will tolerate.
It suits senior engineers, solutions architects and technical leads who design production AI systems and have to defend those designs to security, legal, finance and executive audiences. If your experience is mostly hands-on implementation, expect the governance and stakeholder domains to feel unfamiliar: together they carry real weight, and they reward judgement about people and process as much as about prompts and tools.
The habit the exam rewards above all is choosing the control that changes the structure of the system over the one that watches or asks nicely. Removing a capability beats confirming its use. Scoping a credential server-side beats instructing the model to stay in scope. Fixing the index beats upgrading the model. When a stem offers both kinds of answer, the structural one is usually correct.
CCAR-P tests whether you can pick the single best architectural decision for a stated situation, preferring a structural fix over a compensating control and diagnosing from what changed.
Difficulty
Advanced
Best for
Senior engineers, solutions architects and technical leads who design, govern and hand over production Claude systems, and who explain trade-offs to security, legal, finance and executive stakeholders.
Prerequisites
Check Anthropic's published exam guide for current eligibility requirements. In practice you want experience shipping at least one Claude-based system to production, comfort with retrieval and tool integration, and some exposure to security review, evaluation and stakeholder sign-off.
63
Questions
120 min
Time allowed
720 / 1000
Pass mark
$175
Exam cost (USD)
318
Practice questions
How this exam thinks
Four habits carry most of the marks on CCAR-P, and none of them is recall.
First, every question is an applied situation. There is a named organisation, a system in a particular state, a constraint such as a deadline, a latency target or an audit finding, and a stakeholder asking for a recommendation. Read the last sentence first so you know whether you are being asked for a recommendation, a diagnosis or the most likely cause, then read the stem for the constraints that rule options out. A clinician who will abandon a pilot over extra interruptions has just eliminated every answer that adds a confirmation step.
Second, you are choosing the single best decision, not a defensible one. Most distractors are reasonable things a competent architect might do on another day. They fail because they answer a different requirement, add cost the scenario cannot bear, or postpone the decision the stakeholder asked you to make. Judge each option against the stated requirement and nothing else.
Third, prefer a structural fix over a compensating control. A structural fix removes the cause: take away the tool the role never uses, scope the credential to the session on the server, re-chunk the documents, change the ordering so a stable prefix can be cached. A compensating control leaves the cause in place and manages it: log the call, alert on it, ask a human to confirm it, add a line to the system prompt forbidding it. Compensating controls have their place where the capability is genuinely needed, but when the stem shows that the role does not need the thing at all, the answer that removes it beats every answer that watches it.
Fourth, when something regresses, find what changed before tuning anything. A quality drop after a document refresh with the model, prompt and code unchanged points at retrieval, not at the model. An accuracy fall the week a shared tool bundle replaced a scoped tool set points at the tool set. A cost rise with every user-side figure flat points at what was added to each request. The exam repeatedly offers a tempting tuning move, such as a larger model, a longer prompt or a bigger evaluation set, alongside the option that investigates the one change in the window. Investigate the change first.
What each domain tests and how to study it
The CCAR-P blueprint is split across 7 domains. Weights are the official share of the exam; see the official exam guide for the authoritative breakdown.
What you must be able to do. Frame a business problem before choosing technology, pick the simplest architectural pattern that meets the requirement, and design an end-to-end system with a working feedback loop.
In one sentenceStart from the outcome the business needs, choose the least autonomous pattern that delivers it, and make sure the design has input checks, output checks and a feedback loop.
Recall check: answer these from memory first
Name the three architectural patterns in increasing order of autonomy, and say what each extra step of autonomy costs.
List the four stages of an end-to-end Claude system and name the one most often missing from a proposed design.
Give two signals in a scenario that a multi-agent design is not worth its coordination cost.
Name the five business value pillars and say how you tell which one a stem prioritises.
What it tests. Whether you can turn a business problem into a Claude-based design without reaching for the most impressive option. It covers separating the outcome from a proposed mechanism, deciding whether a language model belongs in the solution at all, designing the full path from input through processing and output to feedback, choosing between an augmented LLM call, a code-defined workflow and a model-directed agent, multi-agent orchestration and its coordination cost, decomposition into independently checkable steps, and tying the design to the value pillar a stakeholder actually cares about.
How to study it. Take two systems you know and draw each as four boxes: input, processing, output and feedback. The box you struggle to fill is usually the one the exam will ask about, and it is most often feedback. Then practise the pattern choice out loud: for each scenario, say why a single augmented call is not enough before you allow a workflow, and why a workflow is not enough before you allow an agent. For multi-agent items, ask what the extra agents buy and what they cost in coordination, latency and failure handling.
Easy to confuse
A workflow versus an agent. In a workflow, code fixes the sequence of steps and the model fills in each one; in an agent, the model decides the next step itself. If the steps can be listed before the task starts, a workflow is cheaper, faster and more predictable, and the exam treats the agent as the over-engineered answer.
The business outcome versus the proposed mechanism. A stakeholder who asks for a chatbot is describing a mechanism; the outcome might be fewer repeat support contacts. Designs judged against the mechanism tend to miss the requirement, so the best option is usually the one whose success can be measured against the original problem.
Parallel specialised agents versus one well-scoped agent. Parallel agents earn their cost when sub-tasks are genuinely independent and large enough to benefit from separate context. When the work is small or tightly coupled, the added orchestration, result merging and failure containment cost more than they return, and a single agent wins.
A retail bank's head of complaints asks for a customer-facing chatbot. Discovery shows that 18 per cent of complaints breach the regulator's five-business-day acknowledgement deadline, and case timing shows most of the delay sits in staff manually sorting each complaint into one of 12 regulatory categories before it can be routed. The team has eight weeks and must show the sponsor a measurable improvement in that breach rate. What should the architect recommend?
ABuild the customer-facing chatbot as requested, measured by how many complaints customers submit through it rather than by email each week
BFine-tune a model on five years of past complaints so it learns the 12 categories, and begin routing only once that training has finished
CDeploy an agent that investigates each complaint, decides the outcome and sends the final response letter to the customer without staff review
DUse Claude to classify each new complaint into the 12 categories and route it, measured by breach rate and by accuracy on a labelled samplecheck_circle Correct
Separate the outcome a sponsor needs from the mechanism they proposed, and aim the solution at the measured bottleneck with a matching success metric. The business problem is the acknowledgement breach rate, and the evidence places the delay in manual categorisation. Framing the solution as classification plus routing attacks that cause directly, fits an LLM's strength with unstructured text, and lets the team prove success with the breach rate and a labelled accuracy sample within the deadline. A chatbot answers the proposed mechanism rather than the problem.
Why A is wrong: It is tempting because it delivers exactly what the sponsor asked for. It is wrong because the measured delay sits in internal categorisation, not in how complaints arrive, so a new intake channel leaves the breach rate untouched and its success metric says nothing about the original problem.
Why B is wrong: It is tempting because historical labelled complaints look like ideal training data. It is wrong because fine-tuning comes before any prompt-based classification has been tried, adds data preparation and evaluation work that threatens the eight-week deadline, and may not beat a well-prompted model on a 12-category task.
Why C is wrong: It is tempting because it promises to clear the whole backlog at once. It is wrong because it is far larger than the stated problem, removes human judgement from a high-impact regulated decision, and the acknowledgement deadline only requires faster categorisation and routing.
Why D is correct: This targets the step the timing data identified as the bottleneck, uses a language model for a text classification task it suits, and ties success to the breach rate the sponsor cares about plus an accuracy check that catches misrouting.
What you must be able to do. Match a model tier to a workload by trade-off, write system prompts and templates that know their own limits, and treat context as a budget with both a cost and a quality impact.
In one sentenceChoose the model by capability against latency and cost, steer it with the lightest prompting technique that works, and spend context only on what each step needs.
Recall check: answer these from memory first
Say when you would route a step to a faster, cheaper model tier and what you would check before doing so.
Explain why the order of static and dynamic content in a prompt decides whether caching helps.
Name a requirement that a prompt-level guardrail can meet and one that it cannot.
What it tests. The model and prompt layer, judged as design decisions rather than tricks. It tests choosing between more capable and faster, cheaper tiers and routing between them inside one system, writing system prompts that set role, scope and constraints, building templates with clear slots for variable input, and knowing when a prompt-level guardrail is enough and when a control outside the model is required. It covers zero-shot, few-shot and chain-of-thought prompting and their costs, context budgeting through trimming, summarisation and selective retrieval, and prompt reuse through caching a stable prefix, modular prompts and Skills.
How to study it. Learn the model choice as a trade-off triangle rather than a list of names, because the items never depend on which models are current. For every prompt you have written, mark which parts are static and which change per request, then reorder it so the static material comes first; that single exercise covers most caching items. Finally, list every guardrail you have ever put in a system prompt and ask which of them would survive a hostile input. The ones that would not belong in code.
Easy to confuse
Prompt caching versus shortening the prompt. Caching reuses a stable prefix across requests, so it only pays when the static content comes first and is repeated. Shortening removes content altogether. If the stem describes a long, identical preamble ahead of changing user input, the fix is ordering for the cache, not cutting material the model needs.
Few-shot examples versus chain-of-thought. Examples steer format and show how to handle edge cases; step-by-step reasoning helps multi-step problems at a cost in tokens and latency. Match the technique to the failure: wrong shape or inconsistent edge handling wants examples, faulty multi-step logic wants reasoning.
A more capable model versus a better context. A larger model cannot answer from information that never reached it. When the stem shows missing, stale or buried context, upgrading the model is the distractor and fixing what goes into the window is the answer.
A software company runs two Claude-based review paths. An editor plug-in suggests fixes as developers type and must return its first token within 1.5 seconds. A nightly job reviews every merged pull request that touches concurrency code, and engineers read its findings the next morning. On 200 held-out concurrency pull requests, enabling extended thinking on the nightly reviewer raised the share of seeded defects found from 58 to 81 per cent and added about 40 seconds per review, and the engineering director has approved the higher cost for that job. Where should extended thinking be enabled?
AOn both paths, so that developers and the nightly job receive reasoning of the same depth.
BOn the editor plug-in only, because interactive suggestions reach developers the soonest.
COn the nightly concurrency review only, leaving the editor plug-in to answer without it.check_circle Correct
DOn neither path, moving the nightly job to the fastest tier to offset the cost of the reviews.
Enable extended thinking where measured reasoning gains matter and latency is slack, and keep it off paths bound by a tight interactive latency budget. Extended thinking lets the model spend additional tokens reasoning before it produces the final answer, which improves results on multi-step problems such as concurrency defects but lengthens time to the answer and raises token spend. The decision is therefore made per path: an offline job with an approved budget and a measured quality gain is where the trade pays, while an interactive path with a 1.5 second first-token budget is where it does not.
Why A is wrong: Consistency across tools is an appealing principle, and the nightly gain suggests the plug-in might improve too. It is wrong because extended thinking adds generation time before the answer, which the 1.5 second first-token budget on the editor path cannot absorb, and the plug-in's quick fix suggestions were never shown to need deeper reasoning.
Why B is wrong: It is tempting to put the best reasoning where users see it first. It is wrong on both counts: the plug-in has the tight latency budget that extended thinking would break, and the measured 23-point gain in defects found belongs to the nightly concurrency review, which this option leaves without it.
Why C is correct: This is correct because extended thinking earns its latency where a task needs multi-step reasoning and the consumer is not waiting: the nightly review has no one reading it until morning, a measured gain from 58 to 81 per cent, and approved cost. The editor plug-in keeps its 1.5 second budget by answering directly.
Why D is wrong: Cost reduction is a reasonable instinct for a job that runs on every merged pull request. It is wrong because the director has already approved the higher cost, and dropping both extended thinking and model capability on hard concurrency reasoning gives up the measured improvement in defect detection that the job exists to deliver.
What you must be able to do. Scope every agent to the capabilities its role needs, enforce identity and authorisation outside the model, design retrieval that fits the data, and choose the integration mechanism and context strategy that suit the system.
In one sentenceThe heaviest domain: least privilege on tools and credentials, retrieval matched to data shape, observability that catches silent degradation, and the right way for Claude to reach each external system.
Recall check: answer these from memory first
An agent holds write tools its read-only role never calls. Name the best fix and two weaker alternatives the exam will offer.
Say where a customer's account scope must come from in an integration, and why it must not come from the model.
Give two query patterns where semantic vector search is the wrong retrieval choice, and what to use instead.
Say when an MCP server is worth building instead of a direct API call.
What it tests. How a Claude system connects to everything around it. It tests reviewing a tool set for capability bloat and removing what the role does not use, analysing whose identity a system acts with and whether authorisation is enforced server-side or merely requested of the model, justifying a configuration against an accuracy and latency requirement, choosing observability that surfaces quality problems as well as outages, designing RAG pipelines with sensible chunking, metadata and index refresh, matching retrieval to the data and the questions asked, choosing between an MCP server, a direct API or CLI integration and agent-to-agent communication, and weighing progressive discovery against loading everything up front.
How to study it. Spend the most time here. Start with least privilege, because it resolves a large share of items: for any agent in a stem, list what its role actually does, then treat every capability beyond that as the problem. Next, trace identity: for each call, ask whose credential is used and where the customer or tenant scope comes from, and distrust any design where the model supplies that scope. For retrieval, build a small pipeline and break it deliberately with oversized chunks, missing metadata and a stale index, so you can recognise each failure by its symptoms. Learn when vector search is the wrong tool: exact identifiers and aggregations want keyword or structured queries.
Easy to confuse
Removing a capability versus gating it with confirmation or logging. If the role never uses the capability, removal eliminates the risk outright, while confirmation dialogs add friction and decay into reflexive clicks, and logging only records harm after it happens. Gating is the right answer only when the role genuinely needs the capability.
Authentication versus authorisation. Authentication establishes who the user is; authorisation decides what that identity may touch. A signed-in user with a shared, broad service credential behind the integration is authenticated but not properly authorised, and the fix is scope enforced in the integration layer from the session.
Progressive discovery versus monolithic context. Loading every tool, document and instruction up front costs tokens on every request and distracts tool selection; progressive discovery loads them as a task needs them but risks something never being found. The stem's tell is whether most of the loaded material goes unused.
An MCP server versus a direct API integration. An MCP server pays off when several clients or agents will reuse the same tools and someone owns its maintenance. A single integration used by one system is often simpler as a direct call, so the exam weighs reuse and ownership against simplicity.
Worked example from the CCAR-P bank
lock_openFree sampleIntegrationmedium
A hospital network is piloting a discharge-summary agent that reads a patient's notes, medications and results and drafts a summary for the attending clinician to edit. The agent was built on the hospital's shared clinical tool gateway and inherited every tool on it, including place_medication_order and cancel_appointment, which the summary role never calls. Clinicians have said they will abandon the pilot if it adds interruptions to their workflow, and go-live is in three weeks. The security lead asks for a recommendation on the two write tools. What should the architect recommend?
ARemove both write tools from the agent's configuration so its tool set covers only the read operations the summary role usescheck_circle Correct
BKeep both tools but require the clinician to approve each order or cancellation in a confirmation dialog before it runs
CKeep both tools and log every invocation to the security monitoring platform, with an alert on any call from the agent
DAdd a system-prompt instruction forbidding medication orders and cancellations, and verify it with a red-team test suite
When an agent holds a capability its role never uses, removing that capability beats confirming, logging or instructing against its use. Least privilege is a structural control: a tool that is not in the agent's configuration cannot be invoked by any prompt, error or injected instruction. Confirmations, logs and prompt rules all leave the capability reachable and either add friction or act after the harm. Because the summary role is read-only, removing the write tools costs nothing in function while removing the risk entirely.
Why A is correct: Correct. The summary role never places orders or cancels appointments, so removing those tools eliminates the risk outright rather than detecting or gating it. It adds no clinician interruptions, shrinks the attack surface a manipulated prompt could reach, and is a configuration change that fits the three-week timeline.
Why B is wrong: A confirmation step feels safe because a human sees every write before it lands. It is wrong because it is a compensating control on a capability the role does not need, it adds exactly the workflow interruptions clinicians said would end the pilot, and approval fatigue tends to turn confirmations into reflexive clicks.
Why C is wrong: Logging with alerting is tempting because it adds no clinician friction and gives the security team visibility. It is wrong because it is a detective control: an erroneous medication order or cancelled appointment has already happened by the time the alert fires, and the capability itself remains.
Why D is wrong: An instruction plus red-team testing looks rigorous and costs nothing at runtime. It is wrong because a prompt instruction is a soft control that injected content or an unusual input can override, and passing a test suite does not prove the model will refuse every case when the tool is still available to call.
What you must be able to do. Define metrics tied to requirements before changing anything, build evaluation sets that combine methods sensibly, diagnose regressions from what changed, and optimise cost and latency without breaking quality.
In one sentenceMeasure first against a representative set, change one thing at a time, find what changed when quality drops, and verify quality holds after every optimisation.
Recall check: answer these from memory first
Answers degrade after a document refresh while the model and prompt are unchanged. Name the first thing to investigate.
Name the three evaluation methods you would combine and where each is reliable.
Say why a model grader needs its own validation and what it is validated against.
Give three cost or latency levers and the quality check that should follow each.
What it tests. Whether you can tell if a Claude system is working and make it better without guessing. It tests choosing metrics for accuracy, latency, cost, safety and security and spotting when one headline number hides a failure elsewhere, designing evaluation sets from real traffic and its edge cases, combining code-based checks, model-graded evaluation and human review, A/B testing against a baseline, diagnosing whether a failure is a prompt problem, a retrieval problem, a hallucination or a mismatched model, optimising token use and latency through caching, routing, batching and trimming, and monitoring production for drift.
How to study it. Make the diagnostic habit automatic. For any regression scenario, write down what changed in the window before you look at the options, then pick the option that investigates that change. Build one small evaluation set yourself with a code-based check and a model grader, and validate the grader against a handful of human judgements so you feel why that step matters. For optimisation items, name the constraint the stem states, such as cost, latency or throughput, then choose the lever that addresses it and the check that confirms quality held.
Easy to confuse
A retrieval failure versus a model failure. If the model, prompt and code were unchanged and quality fell after the documents or index changed, the cause is almost certainly retrieval. Swapping or upgrading the model is the exam's favourite distractor for this pattern, because it tunes a component that did not change.
Sampling noise versus a real regression. A small swing on a small set can be noise, but a large drop on a fixed set with the failures sharing one pattern is a real change. Enlarging the evaluation set before acting is the wrong move when the evidence already points at a cause.
Optimising cost versus preserving quality. Every cost or latency lever, from a cheaper model tier to trimmed context, can quietly cut accuracy. The correct option pairs the optimisation with a quality check against the same evaluation set; an option that reports savings without that check is incomplete.
A hospital group is piloting Claude to draft discharge summaries from inpatient records, and the clinical safety officer has set one requirement in writing: no summary may leave out a documented drug allergy or a medication change. The team's evaluation plan scores each draft against a clinician-written reference summary using an overall similarity score, and the pilot average is 0.87. About one record in six carries an allergy or a medication change. What should the architect recommend the team measure before the pilot widens?
AThe overall similarity score, with its pass threshold raised from 0.87 to 0.92 across the pilot
BThe share of summaries the model itself rates as complete, from a self-check added to each draft
CRecall of documented allergies and medication changes per summary, on a labelled set, reported apartcheck_circle Correct
DClinician satisfaction with each draft, as a five-point rating from the ward doctors who sign it
Derive the metric from the stated requirement, measuring recall on safety-critical items separately so an aggregate quality score cannot hide omissions. The requirement is defined by a failure on specific items, so the metric has to count those items. An overall similarity score averages over all text and over records that contain no allergy or medication change, which lets a 0.87 mean coexist with missed allergies. Item-level recall on a labelled set, with a threshold tied to the safety officer's requirement and reported separately, exposes exactly the failure the requirement forbids.
Why A is wrong: Raising the bar is tempting because it looks like a stricter safety gate. It is wrong because the similarity score averages over the whole narrative and over the five in six records with no allergy or medication change, so a dropped allergy line barely moves it and a higher threshold still does not target the stated requirement.
Why B is wrong: A self-check is tempting because it is automatic and cheap to run on every draft. It is wrong because a model's own judgement of completeness is not a ground-truth accuracy signal; the same model that omitted an allergy can rate its draft complete.
Why C is correct: The requirement names a specific failure, the omission of an allergy or medication change, so the metric must count those items: recall against records labelled with them, reported separately from any aggregate so a good average cannot hide a miss.
Why D is wrong: Clinician ratings are tempting because the raters are the domain experts. It is wrong because a satisfaction score reflects overall impression and readability; a busy signer does not check every line against the record, so omissions go unmeasured.
What you must be able to do. Layer safety controls with enforcement outside the model, name the failure mode a scenario exposes, place human review in proportion to risk, and map regulatory requirements to architectural controls.
In one sentenceA control in code is a guarantee and an instruction to the model is not: layer the guarantees, put people where the risk is, and turn each regulation into a concrete design choice.
Recall check: answer these from memory first
Name two controls enforced outside the model and two that are only requested of it.
Say where human review belongs in a system and what evidence shows a review step has stopped working.
Given a data residency requirement, name the architectural decision it constrains.
What it tests. The controls and judgement that keep a Claude system safe and compliant. It tests layering input screening, output filtering and restricted permissions with enforcement outside the model, recognising characteristic LLM failure modes such as hallucination, prompt injection through untrusted content, data leakage and inconsistent output, placing human review before irreversible or high-impact actions and on a risk-based sample rather than everywhere or nowhere, mapping obligations such as data minimisation, residency, health data handling and audit trails to architecture, and choosing measurable steps on bias, fairness and transparency over statements of intent.
How to study it. Sort every control you know into two piles: enforced in code and requested of the model. Practise until you can place any control in a stem instantly, because many items turn on that line. For human review, learn the failure of confirmation fatigue: a reviewer who approves almost everything in seconds is no longer a control. For compliance, drill the translation step, from a stated obligation to the design choice that meets it, rather than memorising regulations.
Easy to confuse
Human-in-the-loop as a control versus a rubber stamp. Review works when it is rare enough and consequential enough that people actually read what they approve. Near-universal approval at a few seconds per decision shows the step has become reflex, and the structural fix is removing the capability rather than redesigning the dialog.
Prompt injection versus hallucination. Prompt injection is untrusted content steering the model into actions it should not take; hallucination is the model inventing content with no grounding. Injection is contained by limiting what the agent can do, hallucination by grounding and validation, so naming the right failure picks the right control.
A measurable fairness step versus a statement of intent. Testing outcomes across groups on a representative set, and disclosing AI involvement to users, are actions that can be checked. A policy line promising fairness is not, and the exam treats it as the weaker option whenever a measurable one is available.
A retail bank gives its 600 relationship managers an assistant that looks up client accounts through a tool. The model supplies the account number as a tool argument, and the tool handler calls the core banking service with a shared service account that can read every client. The core banking service already holds each manager's book and enforces it when called with that manager's own credential. A conduct review found 14 lookups in one month of clients outside the requesting manager's own book, each after the manager typed a name that matched a different client. The regulator requires that a manager be prevented, not merely detected, from viewing a client outside their book, and managers must keep looking clients up by name. Which TWO changes meet the requirement? Select TWO.
AAdd the manager's list of client account numbers to the system prompt and instruct the model to refuse lookups of any other account.
BHave the tool handler check each requested account against the signed-in manager's book, taken from the session, and reject any mismatch.check_circle Correct
CLog every lookup with the manager's identity and send compliance a daily report of the accounts accessed outside the manager's book.
DMake the assistant ask the manager to confirm the client's full name and date of birth before it calls the lookup tool for any account.
EReplace the shared service account with a delegated credential for the signed-in manager, so the banking service applies their access rights.check_circle Correct
Enforce data-access scope from the authenticated session and delegated identity, never from model-supplied arguments or a shared all-access credential. The model chooses the account number, so any control that relies on the model choosing correctly can fail. Attaching the manager's identity from the session and checking the requested account against their book in the tool handler, and calling the banking service with the manager's delegated credential instead of an all-access service account, both put the boundary in code and systems the model cannot argue with. Lookup by name still works; only the out-of-book result is refused.
Why A is wrong: Tempting because it gives the model the information it needs to stay inside the book. It is wrong because the restriction still depends on the model obeying an instruction, and a name collision or a crafted message can lead it to request another account; the regulator asked for prevention, which an instruction cannot guarantee.
Why B is correct: Correct. The authorisation decision moves out of the model and into code that attaches the manager's identity server-side from the session, so an account outside the book is refused whatever argument the model supplies, while lookup by name keeps working.
Why C is wrong: Tempting because an audit trail is useful and conduct teams expect one. It is wrong because logging and reporting detect a breach after the client data has been shown, and the regulator explicitly requires prevention rather than detection.
Why D is wrong: Tempting because a confirmation step would catch some mistaken name matches. It is wrong because the manager making the request is the person confirming, the step is still carried out by the model, and nothing stops a confirmed lookup of a client outside the book.
Why E is correct: Correct. A shared credential that can read every client gives the tool far more reach than any one manager holds; calling the banking service as the manager means its own entitlement checks refuse out-of-book accounts, enforcing the boundary outside the model.
What you must be able to do. Run discovery before design, explain trade-offs so each audience can make the decision it owns, set expectations a probabilistic system can meet, and carry a solution through handoff with clear ownership.
In one sentenceGather the requirement before designing, frame each trade-off for the person who has to decide it, agree SLAs the system can keep, and never hand over a system without an owner for its ongoing evaluation.
Recall check: answer these from memory first
Name four things structured discovery must establish before design begins.
A sponsor expects the system to be right every time. Say how you reset that expectation and what you agree instead.
List what handoff documentation must contain for another team to operate and change the system.
What it tests. The work around the system that decides whether it succeeds. It tests structured discovery of stakeholders, current processes, data sources, constraints and success criteria, spotting the requirement that has not been gathered yet, communicating a design choice and what it gives up to technical, security, legal and executive audiences, managing expectations and SLAs for a system whose output is probabilistic, building feedback loops from stakeholders back to the system, documenting architecture, decision rationale, prompt and model versions and runbooks, and supporting each lifecycle phase from discovery through design, handoff, monitoring and iteration.
How to study it. Treat this domain as seriously as the technical ones, because its weight is real and its distractors are subtle. For each stem, identify who is asking and what decision they own: an executive owns the investment and risk appetite, a security lead owns the control, a legal team owns the compliance reading. The best answer gives that person what they need to decide. Practise rewriting one technical trade-off three ways, for an engineer, a risk owner and a sponsor. Then walk a past project through the five lifecycle phases and name the risk at each transition, especially handoff.
Easy to confuse
Designing now versus finishing discovery first. When a stem shows a missing success criterion, an unknown data source or an unconsulted stakeholder, the best move is to close that gap before committing to a design. Options that start building to save time carry the risk of solving the wrong problem.
Explaining a trade-off versus making the decision for the stakeholder. The architect's job is to frame options, costs and risks so the owner of the decision can choose. An option where the architect quietly picks a risk position that belongs to legal, security or the sponsor is usually wrong, however technically sound it is.
Handoff of the code versus handoff of ownership. A system handed over without a named owner for ongoing evaluation, monitoring and prompt or model updates will drift unnoticed. The exam rewards the option that assigns that ownership explicitly, not the one that delivers the most complete repository.
Worked example from the CCAR-P bank
lock_openFree sampleStakeholder Communication & Lifecycle Managementmedium
A retail bank wants Claude to draft replies to customer complaints. Discovery has established that the team handles about 4,000 complaints a month, that the complaints policy library holds 2,300 documents which the compliance team revises weekly, and that the regulator requires each reply to cite the policy clause, and its version, that was in force on the date of the complaint. The sponsor proposes fine-tuning a model on two years of approved replies. What should the architect recommend?
AFine-tune a model on the two years of approved replies so policy positions and house style are carried in the model weights.
BLoad the whole policy library into the context window on each request so that retrieval cannot miss a relevant policy clause.
CPrompt the model with a summary of current policy positions, and have a reviewer add the clause citations before each reply is sent.
DRetrieve clauses from a versioned policy index filtered by complaint date, and have each draft cite the retrieved clause and version.check_circle Correct
When discovery reveals frequently revised source material and a version-specific citation duty, the architecture should retrieve from a versioned index rather than fine-tune. Discovery answers about data freshness and traceability decide the architecture. Fine-tuning bakes knowledge into weights that go stale on every revision and cannot attribute an answer to a source. Retrieval keeps knowledge outside the model, so a weekly revision is live once indexed, a date filter selects the version in force on the complaint date, and the retrieved text is what the reply cites to satisfy the regulator.
Why A is wrong: This is tempting because the sponsor proposed it and two years of approved replies look like ideal training data for tone and policy. It is wrong because a weekly revision cycle would leave the weights stale within days, and a fine-tuned model cannot point to which clause version it relied on, so the citation requirement cannot be met.
Why B is wrong: This is tempting because it removes retrieval misses as a failure mode. It is wrong because 2,300 documents on each of 4,000 requests a month is costly and slow, the model still has to pick the version in force on the complaint date from a mass of competing versions, and long-context stuffing is the over-engineered answer where targeted retrieval fits.
Why C is wrong: This is tempting because it keeps a human in the loop on a regulated output and avoids building an index. It is wrong because a summary of current positions cannot reflect the version in force on an older complaint date, and moving the citation duty to a manual step is a compensating control that adds cost to every reply where a structural fix exists.
Why D is correct: Correct. The discovered facts (weekly revisions, a large library, and a date-specific citation duty) all point to retrieval over a versioned index: updates reach answers as soon as they are indexed, the date filter selects the version in force, and the retrieved clause gives the draft something concrete to cite.
What you must be able to do. Configure Claude tooling for a team rather than an individual, apply AI assistance to developer workflows with review proportional to risk, and debug production issues from traces rather than guesses.
In one sentenceTeam settings live in shared, committed configuration, generated work still gets reviewed where it matters, and production faults are isolated from telemetry before anyone changes a prompt or a model.
Recall check: answer these from memory first
A convention works for one engineer and nobody else. Name the most likely cause and the fix.
Say where human review of generated code is still required and where it can be lighter.
Describe how a trace tells you whether a fault lies in the integration or in the model output.
What it tests. Enabling teams to work well with Claude. It tests setting up tools such as Claude Code with shared project configuration and instructions committed alongside the code, consistent permissions and shared tool integrations, distinguishing team-wide configuration from personal settings, applying AI-assisted tooling to code generation, review, testing and documentation while keeping human review where risk warrants it, and using traces and logs to tell whether a production fault lies in the integration or in the model's output.
How to study it. This is the smallest domain, so a focused pass is enough, but do not skip it. Ask two questions of every configuration item: who receives this setting, and is it shared through the repository or held privately by one person. For debugging items, practise reading a trace in order, request, tool calls, tool results and response, and naming the first point where it diverges from what should have happened. That divergence point separates an integration fault from a model fault.
Easy to confuse
Team configuration versus personal settings. Settings held in a personal configuration never reach teammates, so behaviour one person sees and others do not is nearly always a scope fault. Anything the team must share belongs in project configuration committed with the repository.
An integration fault versus a model fault. If the trace shows a tool returning an error, an empty or malformed payload, or the wrong record, the model was given bad input and the fix is in the integration. Only when the inputs were correct and the response still went wrong is the model output the place to look.
A state revenue agency is rolling Claude Code out to 140 engineers who work across 40 repositories. Its information security branch requires that Claude Code's built-in web search and web fetch tools be unavailable in every session on agency machines, that neither a repository's maintainers nor an individual engineer can remove that restriction, and that it be in force on the first day of rollout without a change to each of the 40 repositories. What should the architect recommend?
ACommit a deny rule for web tools to each repository's shared project settings and require code review on any change to it
BAdd an instruction to every repository's CLAUDE.md telling Claude Code that use of the web search and fetch tools is prohibited
CAllow the web tools, log every fetched address to a central store, and have the security branch review that log each week
DDeploy a centrally managed policy to agency machines that denies the web tools and takes precedence over project and user settingscheck_circle Correct
When a tooling restriction must hold across every repository and resist removal by maintainers or engineers, enforce it through centrally managed policy rather than per-repository settings. Claude Code merges settings from several layers, and a centrally managed policy deployed by administrators takes precedence over both the project settings committed to a repository and each engineer's personal settings. That makes it the only layer here that repository maintainers and individual engineers cannot override, and because it is installed on machines it covers all 40 repositories at once with no repository changes. Project settings remain the right place for team policy that a repository's own maintainers are trusted to own.
Why A is wrong: This is tempting because committed project settings are shared and reviewable, which is the right home for most team policy. It fails two stated constraints: it needs a change in all 40 repositories before rollout, and a repository's maintainers can approve a later commit that removes the rule.
Why B is wrong: Shared project instructions do reach every session, which makes this look like a team-wide control. An instruction is guidance to the model rather than an enforced permission, so it is a weak sole control for a security requirement, and it still needs edits to all 40 repositories that maintainers could later revert.
Why C is wrong: Central logging gives the security branch visibility, which feels like governance. It is a compensating control after the fact: the requirement is that the tools be unavailable, and a weekly review only detects use that has already happened.
Why D is correct: An administrator-managed policy sits above project and user settings in Claude Code's precedence order, so no repository commit or personal setting can loosen it. Because it is deployed to machines rather than to repositories, it applies to all 40 repositories from day one without touching any of them.
A study plan that works
Read the exam guide and map the seven domains
Day 1
Read Anthropic's published exam guide for CCAR-P end to end and write the seven domains on one page with their weights. Note that Integration carries the most weight, and that governance and stakeholder communication each carry as much as the model and prompting domain. Book a provisional exam date so the plan has an end.
Draw two systems end to end
Week 1
Take two Claude systems you know and draw each as input, processing, output and feedback, then name the architectural pattern each uses and justify it against the simpler alternative. Add the model tier choice, what is cached and what enters the context window at each step. This covers most of Domains 1 and 2 in practice rather than on paper.
Audit integration: tools, identity and retrieval
Week 2
For one agent, list every tool and credential it holds and cut each one its role does not use. Trace where the user or tenant scope comes from on every call. Then build or inspect a small retrieval pipeline and break it with poor chunking and a stale index, so you recognise each failure by its symptoms. This is the heaviest domain, so give it a full week.
Practise evaluation and the what-changed diagnosis
Week 3
Build a small evaluation set with a code-based check, a model grader and a few human judgements to validate the grader. Run one change against a baseline, one variable at a time. For every regression scenario you meet in practice, write down what changed before reading the options, until the habit is automatic.
Cover governance and the stakeholder domains
Week 3 to 4
Sort controls into enforced-in-code and requested-of-the-model, and practise mapping a regulatory obligation to the design choice that meets it. Then rewrite one technical trade-off for an engineer, a risk owner and a sponsor, and walk a past project through discovery, design, handoff, monitoring and iteration, naming the risk at each transition.
Practise applied scenarios and read every rationale
Week 4
Move to full scenario sets and read the explanation on every option, including the questions you got right. Track which distractor family catches you: a compensating control where a structural fix was available, tuning a component that did not change, or an option that answers a different stakeholder's requirement.
Close weak domains, then sit a timed mock
Week 5
Use per-domain accuracy to drill the two weakest domains rather than re-reading strong ones. Then sit at least one full timed run to rehearse pacing, and on multiple-response items read how many answers to select before evaluating the options.
Know when you're ready
Readiness for CCAR-P is a measured score on applied scenarios you have not seen before, not a sense that the ideas are familiar. Every principle in this guide is easy to agree with in isolation. Removing an unused tool is obviously better than logging it. Applying that under a stem with a deadline, an anxious stakeholder and four sensible-sounding options is a different skill, and only practice shows whether you have it.
The common trap at this level is uneven readiness. Strong engineers clear Integration and Evaluation early, then find in the mock that the governance and stakeholder domains are costing them marks because the distractors there are about who owns a decision rather than how a system works. Judge yourself per domain, on unseen items, across more than one session, and treat any domain that only clears the line on a good day as unfinished.
A second signal worth trusting is how you handle regression stems. If you still find yourself drawn to the larger model or the longer prompt before asking what changed, the diagnostic habit is not yet automatic, and the exam will find that.
This guide gives you the map. The practice bank is where you find out whether you can navigate it, with an explanation of why the right answer is right and every wrong one is wrong. Readiness scoring tells you when you are there. Not before.
Ready to put this into practice?
Free CCAR-P questions, every answer explained. No sign-up.
Read the last sentence of the stem first. It tells you whether you are recommending, diagnosing or naming a cause, which changes what a correct option looks like.
List the constraints before reading the options. A deadline, a latency target, a staff workflow that cannot take interruptions or an audit finding will eliminate most distractors on its own.
When one option removes a cause and another watches or confirms it, prefer the removal unless the stem shows the capability is genuinely needed.
On any regression stem, find what changed in the window before you consider tuning. An option that upgrades an unchanged component is usually the trap.
Distrust any design where the model supplies its own scope, permission or identity. Those decisions belong in the system, enforced outside the model.
Ask who owns the decision in stakeholder items. The best option equips that person to decide rather than deciding for them.
Check how many responses a multiple-response item asks for, and flag and move on from any item that is eating time.
Frequently asked questions
Is CCAR-P hard?
It is a professional-level exam aimed at senior architects. There is little to memorise, but every item is an applied situation where several options are reasonable and only one is the best decision for the stated constraints. The difficulty is judgement, so scenario practice with full explanations matters far more than reading documentation.
What is the difference between CCAR-P and CCAR-F?
CCAR-F is the builder's exam: it tests implementing agents and tools, the agentic loop, tool and MCP design, structured output and Claude Code configuration and mechanics. CCAR-P is the architect's exam: it tests system-level trade-offs across a whole solution, from pattern choice and integration to governance, regulatory fit, stakeholder communication and lifecycle management after handoff. CCAR-F asks how to make a mechanism work; CCAR-P asks which design to choose, what it gives up and who needs to agree.
How long should I study for CCAR-P?
Four to six weeks of focused study is typical for someone already designing Claude systems in production. If your background is mainly implementation, budget extra time for the governance and stakeholder domains, which reward a different kind of judgement from the technical material.
Which domain should I focus on?
Integration carries the most weight and underpins many items elsewhere, so give it the most time, with solution design and evaluation close behind. Do not treat governance or stakeholder communication as soft topics: together they carry substantial weight and their distractors are subtle.
Do I need to take CCAR-F first?
Check Anthropic's published exam guide for current eligibility requirements. In practice, the builder-level material is assumed knowledge: if you cannot already reason about tools, context and agent behaviour, the architect-level trade-offs will be harder to judge.
Do I need to memorise model names or limits?
No. The model domain tests the trade-off between capability, latency and cost and how to route between tiers, not which models are current or what their limits are. Learn the trade-off and the reasoning, not a list that will change.
Why does the exam keep preferring removing a capability over monitoring it?
Because a capability an agent does not hold cannot be misused by any prompt, error or injected instruction, while monitoring, confirmations and prompt rules leave it reachable and act during or after the harm. When the stem shows the role does not need the capability, removal is the structural fix and the other options are compensating controls.
How many practice questions should I do before booking?
Enough that every domain clears the pass line with margin on questions you have not seen before, across more than one session, and that a full timed run feels comfortable on pacing. Review quality beats volume: read the explanation on every option, because knowing why a sensible answer is not the best one is the skill being tested.
Examworthy is not affiliated with or endorsed by Anthropic. This guide is original study material based on the public exam blueprint. We never reproduce live exam items. CCAR-P and related marks belong to their respective owners.