What you will be able to do
- Decide whether a business problem suits Claude rather than a rules-based or traditional ML approach
- Break a business interaction down into the separate tasks Claude has to perform
- Turn a vague requirement such as 'summarise this' into specific outputs Claude can produce
- Set a measurable success target for each task before building
Key concept
Task-level decomposition of the business problem — A business request, such as 'automate support', is really a bundle of separate tasks. You design, prompt and evaluate Claude one task at a time, so the first step is to list every task the solution has to perform.
1.Start by checking that the problem suits Claude
Business problems don't arrive labelled "LLM task". Before you design anything, check whether the problem has the traits Claude handles well. Anthropic's use-case guides list the signals for several common problems. For customer support, the signals are high volumes of similar questions, answers that draw on large knowledge bases, a need for round-the-clock coverage, sudden spikes in demand, and a consistent brand voice. The guide sums up the first signal this way: Claude "excels at handling a large number of similar questions efficiently, freeing up human agents for more complex issues." Automation here takes over repetitive work and leaves the harder cases to people, rather than replacing the whole team.
For classification problems such as ticket routing, the comparison is usually with a traditional ML pipeline, and the differences favour Claude on several counts.
| Concern | Traditional ML | Claude |
|---|---|---|
| Training data | Requires massive labeled datasets | Can classify with a few dozen labeled examples |
| Changing classes | Laborious and data-intensive to change | Adapts to new or changed class definitions without extensive relabeling |
| Unstructured input | Needs extensive feature engineering | Classifies based on content and context |
| Rule-defined classes | Relies on bag-of-words or simple pattern matching | Understands and applies underlying rules |
| Explainability | Little insight into decisions | Human-readable explanations for each decision |
| Ambiguous input | Often misclassified or sent to a catch-all category | Interprets context and nuance, potentially reducing misrouted tickets |
Recognising the kind of problem also points you to the right design and evaluation methods. The moderation guide says it directly: "Content moderation is a classification problem." So you can measure a moderation system with the same accuracy techniques you would use for any classifier. Its main argument for Claude over keyword rules is that moderation needs a nuanced grasp of language. In "The main actor really killed it!", "killed it" is a metaphor. A veiled threat, on the other hand, should be flagged even though it never mentions violence.
2.Break the interaction down into tasks
Once the problem looks like a good fit, describe the ideal experience from end to end. The support guide suggests you "Outline an ideal customer interaction to define how and when you expect the customer to interact with Claude," because the outline will "help to determine the technical requirements of your solution."
In the guide's car-insurance example, a single chat covers a greeting, questions about electric-vehicle coverage, off-topic questions, and a request for a quote. For the quote, Claude asks follow-up questions, sends the collected details to a quote-generation API tool, and turns the tool's response into a natural reply. That one conversation breaks down into four task groups:
| Task group | What Claude must do |
|---|---|
| Greeting and general guidance | Greet the customer; give general information about the company and the interaction |
| Product information | Explain electric-vehicle coverage, answer follow-up questions, offer links to sources |
| Conversation management | Stay on topic (car insurance); steer off-topic questions back |
| Quote generation | Ask eligibility questions, adapt them to the answers, submit to the quote API, present the quote |
The breakdown shapes the design. Answering from the knowledge base needs company information in Claude's context. Quote generation needs a tool connected to a back-end API. The breakdown also tells you which test cases to write, because the guide says decomposition lets you "prompt and evaluate Claude for every task."
Sources1
3.Turn vague requirements into specific outputs
Some business requests sound specific until you try to build them. "Summarise our contracts" is a common example. The legal summarisation guide warns that "There is no single correct summary for any given document." Without direction, Claude has no way of knowing which details matter to the business. The fix is to identify the specific information the summary must contain and write it down as a list:
details_to_extract = [
"Parties involved (sublessor, sublessee, and original lessor)",
"Property details (address, description, and permitted use)",
"Term and rent (start date, end date, monthly rent, and security deposit)",
"Responsibilities (utilities, maintenance, and repairs)",
"Consent and notices (landlord's consent, and notice requirements)",
"Special provisions (furniture, parking, and subletting restrictions)",
]The moderation guide takes the same approach from a different direction. You start with examples, not categories. You "first create examples of content that should be flagged and content that should not be flagged," include edge cases, and only then draw up the list of moderation categories. You can also adjust the categories to the business. A site that wants to keep minors from posting could add an "Underage Posting" category, for example. In both guides, the business requirement ends up as something concrete that Claude can be checked against.
4.Set a measurable target for each task
The last step before building is to agree what success means with the people who own the problem. The support guide says to "define success criteria and write detailed evaluations with measurable benchmarks and goals." Targets are set per task, and they differ according to what an error would cost the business:
| Criterion | Target |
|---|---|
| Query comprehension accuracy | 95% or higher |
| Response relevance (LLM-based grading for scale) | 90% or above |
| Accuracy of general company and product information | 100% |
| Relevant sources offered where helpful | 80% of interactions |
| Staying on topic | 95% of responses |
| Knowing when to generate a quote, and the quote's accuracy | 100% |
The targets follow the business cost of getting it wrong. A missing link is a small miss. A wrong quote damages the transaction itself, and the guide calls it 'vital information for a successful customer interaction'.
Some outputs have no objective metric. The legal guide notes that evaluating summaries "often lacks clear-cut, objective metrics." It recommends combining quantitative and qualitative methods, grounded in the success criteria you defined.
A retail company is building a customer-support chat widget that must handle very high message volume during flash sales, respond within a few hundred milliseconds, and stay within a tight per-conversation cost budget, while still handling multi-turn troubleshooting reasoning correctly. Which model should the architecture team select as the primary model for this workload?
Correct answer: B — Deploy Claude Haiku 4.5, since it delivers near-frontier reasoning at the lowest per-token cost with the fastest response times available.
- A. Incorrect. Opus 4.8 is positioned for complex agentic coding and enterprise work with moderate latency, not for the lowest-latency, cost-sensitive, high-volume support workload described here.
- B. Correct. Haiku 4.5 is described as the fastest model with near-frontier intelligence at the most economical price point, matching the flash-sale volume, latency, and cost constraints.
- C. Incorrect. Sonnet 5 is a strong general-purpose choice for coding and agentic workflows, but its cost and latency profile is not the best fit when Haiku 4.5 already meets the reasoning bar at lower cost and higher speed.
- D. Incorrect. Fable 5 targets long-running agentic intelligence at premium pricing and slower comparative latency, which conflicts directly with the sub-second, cost-sensitive requirement.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Using Claude for classification still requires a large labeled training set, just as a traditional ML model does.Why is that wrong?
The ticket-routing guide says a few dozen labeled examples are enough, and that class definitions can change without extensive relabeling.
2.A single overall accuracy target is enough to decide whether a Claude solution is ready.Why is that wrong?
Targets are set per task and follow business impact. In the support example, relevant links are targeted at 80% of interactions, while quote accuracy must be 100%.
Covered in Set a measurable target for each task
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Claude excels at handling a large number of similar questions efficiently, freeing up human agents for more complex issues.”
↩︎ Start by checking that the problem suits Claude“Outline an ideal customer interaction to define how and when you expect the customer to interact with Claude.”
↩︎ Break the interaction down into tasks“This outline will help to determine the technical requirements of your solution.”
↩︎ Break the interaction down into tasks“Work with your support team to define success criteria and write detailed evaluations with measurable benchmarks and goals.”
↩︎ Set a measurable target for each task“break down your ideal customer interaction into every task you want Claude to be able to perform.”
↩︎ Key concept“Target 100% accuracy, as this is vital information for a successful customer interaction.”
↩︎ Exam trap 2 - 2.
“Content moderation is a classification problem.”
↩︎ Start by checking that the problem suits Claude“first create examples of content that should be flagged and content that should not be flagged.”
↩︎ Turn vague requirements into specific outputs - 3.
“Claude can easily adapt to changes in class definitions or new classes without extensive relabeling of training data.”
↩︎ Start by checking that the problem suits Claude“Claude's pre-trained model can effectively classify tickets with just a few dozen labeled examples, significantly reducing data preparation time and costs.”
↩︎ Exam trap 1 - 4.
“There is no single correct summary for any given document.”
↩︎ Turn vague requirements into specific outputs“Evaluating the quality of summaries is a notoriously challenging task.”
↩︎ Set a measurable target for each task