What you will be able to do
- Define measurable success criteria with the team that owns a process before Claude takes on any of its steps
- Design approval checkpoints so that consequential actions apply only after a person approves them
- Explain where human review stays in a workflow that Claude now supports
1.Define what 'good' looks like before handing work over
A step handed to Claude without a definition of success can't be judged. Nobody can say whether the new process is better than the old one, or whether it's safe to expand. Anthropic's support guide makes defining success a job for the people who already run the process: the support team, working with whoever builds the solution. They set criteria and measurable targets before the solution goes live.
Targets depend on how much an error in that step costs. In the insurance example, understanding what the customer is asking has a target of 95%, while generating a quote has a target of 100%, because the quote is information the customer relies on. Several of the checks work by reviewing a sample of real conversations. Starting with a small, reviewable sample is the practical way to find out whether Claude is meeting the bar.
| What is measured | Target |
|---|---|
| Correctly understanding customer inquiries | 95% or higher |
| Relevance of each response to the customer's issue | 90% or above |
| Accuracy of general company and product information | 100% |
| Staying on topic | 95% of responses |
| Knowing when to generate a quote, and the quote's accuracy | 100% |
Some outputs have no single right answer, which makes a definition of good even more important. The legal summarization guide notes that there is no one correct summary of a document, so the team has to name in advance exactly what each summary must extract. That list is the specification Claude works to and the checklist reviewers score against.
details_to_extract = [
"Parties involved (sublessor, sublessee, and original lessor)",
"Property details (address, description, and permitted use)",
"Term and rent (start date, end date, monthly rent, and security deposit)",
"Responsibilities (utilities, maintenance, and repairs)",
"Consent and notices (landlord's consent, and notice requirements)",
"Special provisions (furniture, parking, and subletting restrictions)",
]An insurer's manual claims process involves five separate teams handing a claim off through email at each stage. Leadership decides to rebuild the process so a single Claude-driven agent triages, requests documents, and routes each claim end to end. Which term best describes this change?
Correct answer: A — Redesigning the workflow, since the multi-team handoff process is being replaced by an agent-driven process
- A. Correct. Replacing the multi-team, email-driven handoff structure with a single agent that owns triage, document requests, and routing end to end is a redesign of the workflow rather than a small addition to it.
- B. Incorrect. The five-team handoff structure is being removed, not preserved, so this is not augmentation of the existing process.
- C. Incorrect. Nothing in the scenario describes grouping claims for batch processing; the change is about who owns each step of the process.
- D. Incorrect. The scenario does not describe reusing prior claim decisions as cached context; it describes restructuring who performs each step.
2.Build approval into the workflow, not the chat
Some tasks in a process change something real: a price, stock levels, a published listing, a marketing budget. Claude can still do most of the work on these tasks. The design question is where the person's decision sits. Anthropic's merchant agent is a clear example. It explains performance, flags low stock, and drafts marketing campaigns. It can also recommend price changes. But it never applies any of these changes itself.
Instead, every write the agent proposes becomes a staged change that the operator sees as a preview. The store's limits are checked twice, once when the change is staged and again when it is applied. The change applies only when a person approves it through a real control outside the conversation: a button in the merchant portal, a confirmation prompt, or a permission policy that always asks. Typing 'yes' in the chat doesn't count as approval.
Because approving is the one step that has to stay with a person, it needs a control that belongs to that person, such as a portal button or a confirmation prompt. Text typed into a conversation is not that kind of control. Putting approval outside the chat means only a person's deliberate action can apply the change.
Sources3
3.Keep human review where accountability lives
Some outputs carry accountability that a person has to own, even when they don't trigger an action in any system. Legal summaries are the example in Anthropic's guide. Mistakes can create liability for the organisation or its clients, so the guide recommends saying clearly that the summaries are AI-generated and that legal professionals should review them. That is also how you describe the new workflow honestly to the people who depend on it: Claude produces the draft, and a named professional is still responsible for it.
Human review is also how the workflow gets better over time. The guide warns that prompts usually need testing and refinement before they are ready for production. It recommends a systematic evaluation, both quantitative and qualitative, against the success criteria the team defined earlier. The reviewers' findings feed back into the prompts, so the handoff keeps improving instead of being set once and left alone.
A hospital wants to integrate Claude into its clinical documentation workflow to help physicians summarize patient charts, where accuracy is far more important than response speed or cost. Which starting approach best fits this integration?
Correct answer: A — Start with the most capable model, since accuracy outweighing cost considerations favors this choice
- A. Correct. Applications where accuracy outweighs cost considerations, such as clinical documentation, are a case for starting with the most capable model rather than a faster, cheaper one.
- B. Incorrect. Prioritizing speed and low cost does not match a scenario where the stated priority is accuracy over speed or cost.
- C. Incorrect. Optimizing for output token pricing contradicts the scenario's explicit priority of accuracy over cost.
- D. Incorrect. Client-side tool support is unrelated to the summarization accuracy this clinical documentation workflow needs.
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.If Claude asks for confirmation in the chat and the user types 'yes', the change has been properly approved.Why is that wrong?
In Anthropic's merchant blueprint, a change applies only after a person approves it outside the conversation, through a portal button, a confirmation prompt, or an always-ask permission policy. A 'yes' typed in chat doesn't count.
2.Success criteria can be worked out after Claude is live, once you see what it produces.Why is that wrong?
The team that owns the process defines measurable benchmarks before the solution goes live, and the evaluation is then run against them.
Covered in Define what 'good' looks like before handing work over
3.Once Claude's summaries are accurate enough, professional review of them can be dropped.Why is that wrong?
Where errors carry liability, the guide recommends labelling the output as AI-generated and having legal professionals review it.
Covered in Keep human review where accountability lives
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Work with your support team to define success criteria and write detailed evaluations with measurable benchmarks and goals.”
↩︎ Define what 'good' looks like before handing work over“Aim for a comprehension accuracy of 95% or higher.”
↩︎ Define what 'good' looks like before handing work over“Target 100% accuracy, as this is vital information for a successful customer interaction.”
↩︎ Define what 'good' looks like before handing work over“Measure this by reviewing a sample of conversations”
↩︎ Define what 'good' looks like before handing work over“Work with your support team to define success criteria and write detailed evaluations with measurable benchmarks and goals.”
↩︎ Exam trap 2 - 2.
“There is no single correct summary for any given document.”
↩︎ Define what 'good' looks like before handing work over“To achieve optimal results, identify the specific information you want to include in the summary.”
↩︎ Define what 'good' looks like before handing work over“Provide disclaimers or legal notices clarifying that the summaries are generated by AI and should be reviewed by legal professionals.”
↩︎ Keep human review where accountability lives“Prompting often requires testing and optimization for it to be production ready.”
↩︎ Keep human review where accountability lives“Creating a strong empirical evaluation based on your defined success criteria allows you to optimize your prompts.”
↩︎ Keep human review where accountability lives“Provide disclaimers or legal notices clarifying that the summaries are generated by AI and should be reviewed by legal professionals.”
↩︎ Exam trap 3 - 3.
“Draft marketing campaigns with audiences, placements, and budget.”
↩︎ Build approval into the workflow, not the chat“is a staged change with a server-generated ID that the operator sees as a preview card.”
↩︎ Build approval into the workflow, not the chat“are checked when the change is staged and again when it is applied”
↩︎ Build approval into the workflow, not the chat“The change applies only after a person approves it outside the conversation”
↩︎ Build approval into the workflow, not the chat“An approval typed in chat approves nothing.”
↩︎ Exam trap 1