What you will be able to do
- Name the three escalation triggers the exam treats as legitimate: an explicit request for a human, a policy exception or gap, and no meaningful progress
- Choose between escalating at once and offering to resolve, based on whether the customer explicitly asked for a human
- Explain why sentiment scores and the model's self-reported confidence are poor proxies for how complex a case really is
- Design agent behaviour for multiple customer matches: ask for another identifier, never pick one by heuristic
- Write escalation criteria with few-shot examples into a system prompt
- Explain the distinction between escalating immediately on an explicit demand for a human and offering to resolve a straightforward issue when no such demand was made
- Explain why multiple customer matches require clarification by requesting additional identifiers rather than heuristic selection
Key concept
Explicit escalation criteria — The system prompt should name the observable conditions under which the agent hands off to a human: the customer asked for one, the policy does not cover the request, or the agent cannot make progress. It should show examples of each case. The agent should not be left to infer when to escalate from tone, apparent difficulty or its own sense of confidence.
1.What should actually trigger a handoff
A customer-facing agent makes one decision on every turn, whether or not you design for it: handle this myself, or hand it to a person? If you leave that decision implicit, the agent fills the gap with whatever seems reasonable at the time. Sometimes it escalates a simple refund because the customer used capital letters. Sometimes it keeps a customer stuck in a loop because it was never told that going in circles is a reason to stop.
The exam expects you to know three triggers that hold up in production:
1. The customer explicitly asks for a human. This is the customer's choice, not the agent's judgement call. 2. A policy exception or gap. The request falls outside what the written policy covers, or the policy is ambiguous or silent about it. This trigger is about *policy*, not about how hard the case is. A long, fiddly case that policy fully covers is still the agent's job. 3. No meaningful progress. The agent has tried the tools and paths it has and is no longer moving the case forward.
What these triggers share is that someone other than the model can check them. A reviewer can look at a transcript and agree whether the customer asked for a person, whether the policy addressed the request, or whether the last several turns got anywhere. That property is what makes them good triggers, and later sections contrast it with signals that lack it.
The honest answer is no, because none of the three triggers has fired. The customer is angry but has not asked for a person, the policy covers the request, and the agent can make progress. You find this kind of case by mapping your support team's current practice before you automate it, not by guessing afterwards.
Sources1
2.When the customer asks for a human, escalate now
The first trigger has one rule that the exam tests directly: honour an explicit request for a human immediately, without first investigating. Do not look up the account to see whether the problem is easy. Do not try a quick fix first. Do not ask "are you sure? I can probably help with that." Even if the underlying request is trivial, such as a password reset the agent could finish in one tool call, the customer has told you how they want to be served.
The mistake is to treat the request as information the agent should weigh against its own assessment of the case. It is an instruction. The agent can still be useful during the handoff: confirm that it is transferring the customer, and pass along what the customer has already said so they do not have to repeat it. What it must not do is make the handoff wait on its own diagnosis.
This only works reliably when the rule is written out. "Be helpful" and "escalate complex cases" both give the model room to try one more thing first. "If the customer asks for a human agent, transfer immediately; do not attempt to resolve the issue first" does not.
Keep the distinction sharp, because the exam builds scenarios on both sides of it. Escalating immediately is correct when the customer explicitly demands a human, however straightforward the issue is. Offering to resolve is correct when the issue is straightforward, sits within the agent's capability, and the customer has not demanded a person: the agent offers to fix it, and escalates only if the customer then explicitly asks for a human. What decides between the two is whether the customer explicitly demanded a human, not how simple the issue is and not how upset the customer sounds. A straightforward issue is a reason to offer resolution only when there is no explicit demand; once there is one, the simplicity of the issue stops mattering.
A retail customer asks a support agent to match a lower price they found on a competitor's website. The store's documented policy only describes price adjustments when the store's own website lowers a price within 14 days of purchase; it does not mention competitor pricing at all. How should the agent proceed?
Correct answer: A — Escalate the request, since the documented policy is silent on competitor price matching entirely
- A. Correct - the policy only addresses own-site adjustments and is silent on competitor pricing, so this gap should be escalated rather than resolved by inference.
- B. Incorrect - treating silence as an implicit denial fabricates a policy position that was never actually stated.
- C. Incorrect - approving by analogy invents an extension of policy the agent isn't authorized to make.
- D. Incorrect - collecting proof and then independently approving still involves making a policy decision the agent isn't authorized to make on an unaddressed scenario.
3.Frustrated but not asking: acknowledge, then resolve
The opposite case is a customer who is clearly annoyed but has not asked for a person, and whose request the agent can handle. Here the correct pattern has three steps:
1. Acknowledge the frustration briefly and sincerely. Don't be defensive, and don't pile on apologies. 2. Resolve the issue, or offer to resolve it right away, because it is within the agent's capability. 3. Escalate only if the customer repeats that they want a human. You can mention escalation as an option, but the agent does not force it.
How do you know what counts as "within the agent's capability"? You define it. Before you build, list the tasks the agent is expected to perform and the tools that support each one. A routine exchange under a standard return window is on that list. An exception to the return window is not.
The contrast with the previous section is the whole lesson. An explicit request for a human means the agent escalates at once. Frustration with a request the agent can handle means the agent acknowledges it, helps, and escalates only if the customer asks again. Tone alone never tips the decision.
An insurance customer asks whether their policy covers a rental car while their electric vehicle's battery is being replaced under a separate manufacturer recall. The policy documentation addresses rental reimbursement only for collision and comprehensive claims, and does not mention manufacturer recall repairs. What should the agent do?
Correct answer: A — Escalate the question, since the policy documentation does not address this specific recall-related scenario
- A. Correct - the policy is silent on rental coverage during recall repairs, a gap that should be escalated rather than resolved by the agent's own judgment.
- B. Incorrect - treating an unaddressed scenario as automatically excluded asserts a coverage decision the policy never actually makes.
- C. Incorrect - approving reimbursement by analogy to comprehensive claims extends coverage beyond what the documented policy states.
- D. Incorrect - redirecting the customer to the manufacturer avoids answering the coverage question rather than resolving the policy gap appropriately.
Transfer right away. The customer has now explicitly asked for a human, so the first trigger applies. Making a second attempt to resolve the issue would be the investigate-first mistake from the previous section.
An airline support agent searches for a passenger by name and flight route to check a baggage claim. The lookup tool returns three passenger records with the same name on that route, each with a different booking reference. What should the agent do?
Correct answer: A — Ask the customer for their booking confirmation number or another identifier to pinpoint the correct record
- A. Correct - when multiple records match, the agent should request an additional identifier from the customer rather than heuristically guessing which record applies.
- B. Incorrect - picking the most recent flight date is a heuristic guess that could easily point to the wrong passenger record.
- C. Incorrect - profile completeness has no bearing on which record actually belongs to the customer being helped.
- D. Incorrect - merging details from multiple unverified records risks combining information from the wrong customer's account.
Sources4
4.Why sentiment and self-reported confidence are poor proxies
Two shortcuts keep appearing in escalation designs, and the exam expects you to reject both.
Sentiment-based escalation routes a case to a human when the customer sounds negative. But sentiment measures how the customer feels, not how hard the case is. A furious customer with a wrong-size t-shirt has a simple case. A calm, polite customer asking for something the policy never contemplated has a hard one. Escalating on sentiment sends easy cases to humans and lets the policy gaps through, which is the reverse of what you want. Surface language is also a noisy signal in its own right. Anthropic's content moderation guide makes this point about phrases like "killed it", which read as violent but are not.
Self-reported confidence asks the model to rate its own certainty ("escalate if confidence < 0.7"). The number feels precise, but nothing ties it to whether the agent is actually right. A model can be confidently wrong on a policy gap, because it applies the nearest rule as if it covered the case, and it can report low confidence on a routine task. A threshold on an uncalibrated number is a heuristic that only looks measured.
The fix is not a better proxy. It is to go back to the triggers from the first section, which describe the case itself: what the customer asked for, what the policy says, and whether progress is happening.
Sources5
5.Policy gaps: escalate when the policy is silent
The second trigger needs a closer look, because it is where agents most often act too readily. The standard example: the price-adjustment policy says the store will match a lower price *on its own site* within a set window. A customer asks the agent to match a *competitor's* price.
The policy doesn't forbid that, and it doesn't allow it either. It is silent. A helpful-seeming agent will reason by analogy ("we match our own prices, so a competitor match is probably fine") or by contrast ("it only mentions our own site, so no"). Either way the agent has invented policy. The correct move is to recognise the gap and escalate to someone who has the authority to decide.
A well-designed commerce agent works on the same principle: store terms come from the business's policy systems, not from the model's inference. If the policy the agent can see does not answer the question, the agent has nothing to state. That is exactly the kind of situation the escalation path is for.
This is also why the trigger is "policy exceptions and gaps" and not "complex cases". A competitor price match is not complex. It is simply not covered.
6.Multiple matches: ask for an identifier, don't guess
Ambiguity also comes from tool results. A customer gives a name, and the lookup tool returns three customers with that name. The tempting heuristics all sound sensible: pick the most recently active account, the one in the customer's apparent region, or the one with an open order. Each of them will sometimes choose the wrong person. When the next step is a refund, a cancellation or disclosing account details, a wrong guess is a data-exposure incident, not a small error.
The correct pattern is to ask the customer for an additional identifier that separates the matches, such as an email address, order number or postcode, and continue only once a single record remains. This is ordinary information gathering. A well-designed support flow already asks the customer questions to collect what a downstream tool needs before calling it.
The principle underneath is the same one that keeps the agent from inventing policy: act only on records the tools have unambiguously identified. Grounded commerce agents enforce a version of this in the harness itself, where writes accept only IDs that a tool actually returned. Your escalation instructions should cover the identity case explicitly: "If a lookup returns more than one match, do not choose. Ask the customer for another identifier."
Put plainly: multiple customer matches require clarification, not heuristic selection. Requesting additional identifiers costs the customer one short reply. Heuristic selection, however plausible the rule behind it, risks acting on another customer's account, and nothing in the tool result tells the agent which match is the right one. Clarification is also not the same as escalation: an ambiguous lookup is not a reason to hand off to a human. The agent asks the customer for the additional identifier, looks up again, and carries on once exactly one customer matches.
7.Writing escalation criteria into the system prompt
Everything above ends up in the system prompt. A vague instruction ("escalate when appropriate") leaves the model to fill the gap with the proxies the exam warns against. What works is an explicit set of criteria, each with the reason behind it, plus a few worked examples that show when to escalate and when to resolve on your own.
The reasons matter because no list of rules covers every case. If the prompt explains *why* a competitor price match escalates (policy is silent, and only a human can create policy), the model can apply the same logic to a request type you never listed.
The examples matter because models copy the details of examples closely. Choose pairs that separate the cases the exam cares about. Include a frustrated customer the agent resolves next to a calm one who asks for a human and is transferred at once. Include a routine exchange handled end to end next to a request that falls in a policy gap and is escalated. Include a single lookup match that proceeds next to a triple match that prompts a request for an email address. Leave out any example where the reason for escalating is tone, or the model will learn that pattern too.
Escalate to a human when (1) the customer asks for a human: transfer immediately, without investigating first; (2) the request is not addressed by, or is ambiguous under, the written policy, because only staff can decide exceptions; (3) you have tried your available tools and cannot make further progress. Do NOT escalate only because the customer is upset: acknowledge it, resolve what you can, and transfer if they then ask for a person. If a lookup returns multiple customers, ask for another identifier; never pick one. Then add three to six short example exchanges, with contrasting pairs showing escalate and resolve.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Even when the customer asks for a human, the agent should first check whether the issue is simple enough to fix quickly, because that serves the customer faster.Why is that wrong?
An explicit request for a human is honoured immediately, without investigation. How easy the issue is does not matter once the customer has asked for a person.
2.Strongly negative sentiment is a reliable signal that a case is complex and should go to a human.Why is that wrong?
Sentiment reflects how the customer feels, not how complex the case is. Surface wording misleads, and angry customers with routine requests should be acknowledged and helped, not escalated automatically.
Covered in Why sentiment and self-reported confidence are poor proxies
3.If the policy covers own-site price adjustments, the agent can reasonably extend it to competitor price matches.Why is that wrong?
When the policy is silent or ambiguous about a request, the agent should escalate rather than invent terms by analogy. Store terms must come from the policy systems.
4.When a lookup returns several customers, the agent should pick the most likely one, such as the most recently active account.Why is that wrong?
Multiple matches call for asking the customer for another identifier. The agent should act only on a record the tools have unambiguously identified, never on a heuristic guess.
Covered in Multiple matches: ask for an identifier, don't guess
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The more you know about how humans handle certain cases, the better you can work with Claude to do the task.”
↩︎ What should actually trigger a handoff“How are edge cases or ambiguous tickets handled?”
↩︎ What should actually trigger a handoff - 2.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practicesOfficial docs
“Claude responds well to clear, explicit instructions.”
↩︎ When the customer asks for a human, escalate now“Claude is smart enough to generalize from the explanation.”
↩︎ Writing escalation criteria into the system prompt“Think of Claude as a brilliant but new employee who lacks context on your norms and workflows.”
↩︎ Key concept - 3.
“provide straightforward pathways to human specialists when needed”
↩︎ When the customer asks for a human, escalate now“provide straightforward pathways to human specialists when needed”
↩︎ Exam trap 1 - 4.
“break down your ideal customer interaction into every task you want Claude to be able to perform”
↩︎ Frustrated but not asking: acknowledge, then resolve“Ask a set of questions to determine the appropriate quote, adapting to their responses”
↩︎ Multiple matches: ask for an identifier, don't guess - 5.
“the content moderation system needs to recognize that "killed it" is a metaphor, not an indication of actual violence”
↩︎ Why sentiment and self-reported confidence are poor proxies“the content moderation system needs to recognize that "killed it" is a metaphor, not an indication of actual violence”
↩︎ Exam trap 2 - 6.
“The agent states products, prices, availability, and store terms only from tool results in the conversation”
↩︎ Policy gaps: escalate when the policy is silent“cart writes accept only product IDs that a catalog or order tool returned in that session”
↩︎ Multiple matches: ask for an identifier, don't guess“The agent states products, prices, availability, and store terms only from tool results in the conversation”
↩︎ Exam trap 3“cart writes accept only product IDs that a catalog or order tool returned in that session”
↩︎ Exam trap 4 - 7.https://claude.com/blog/building-ai-agents-in-financial-servicesSecondary source
“Clear escalation pathways for complex or ambiguous financial situations”
↩︎ Policy gaps: escalate when the policy is silent - 8.https://claude.com/blog/best-practices-for-prompt-engineeringSecondary source
“Claude 4.x and similar advanced models pay very close attention to details in examples.”
↩︎ Writing escalation criteria into the system prompt