CertSafari
    CLAUDE-CERTIFIED-ASSOCIATE-FOUNDATIONS-CCAO-F-VAR5 · Lessons

    Domain 4 · Lesson 15/30

    Turning a Claude Use Case into Testable Requirements

    Apply Claude to analyze requirements and use cases

    6 min read
    3.2% of exam
    4 sources
    Published 28 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Break a use case into the separate tasks Claude must perform
    • Specify categories, examples and extraction targets, including edge cases
    • Agree measurable success criteria with the owning team before building

    1.Outline the ideal interaction, then split it into tasks

    Once a process has passed screening, the next job is to state precisely what Claude is meant to do. Anthropic's customer-support guide begins with an outline of an ideal interaction, which it says helps determine the technical requirements of the solution. Its worked example is car-insurance support. The customer opens a chat, asks about cover for a new electric car, asks follow-up questions, drifts off topic, asks for a quote, and is finally guided to next steps. Even the quote step is spelled out in detail: Claude asks questions that adapt to the answers, sends the collected details to a quote-generation API tool, and turns the tool's response into a natural reply.

    The outline is then broken into tasks, because a support chat is really several jobs in one conversation: answering questions, retrieving information, and acting on requests. Breaking it down means each task can be prompted and evaluated separately, and it shows the range of interactions the test cases need to cover.

    Task groups derived from the insurance support example
    Task groupWhat Claude must do
    Greeting and general guidanceGreet the customer and give general information about the company
    Product informationExplain electric vehicle coverage, answer follow-ups, offer links to sources
    Conversation managementStay on car insurance and redirect off-topic questions
    Quote generationAsk eligibility questions, submit them to the quote API, present the quote

    Sources1

    2.Pin down categories, examples and what to extract

    Decomposition says what the tasks are. The next step is to make each one concrete enough to judge. For classification, the ticket-routing guide stresses that Claude can only route as well as your categories are defined, and it gives example intent categories from hardware problems to billing enquiries to urgent security issues.

    The content-moderation guide begins from examples rather than categories. Before building, collect content that should be flagged and content that should not, including edge cases and difficult scenarios. Then review the examples to derive a well-defined category list. Working this way tends to bring hidden requirements to the surface, such as the difference between a metaphor and a real threat.

    For generation tasks, the equivalent is naming the fields. The legal-summarisation guide warns that there is no single correct summary of a document, and that without clear direction Claude struggles to know which details to include. Its fix is an explicit list of the information to extract:

    Requirements for a sublease summary, written as an explicit extraction listpython
    details_to_extract = [
        "Parties involved (sublessor, sublessee, and original lessor)",
        "Property details (address, description, and permitted use)",
        "Term and rent (start date, end date, monthly rent, and security deposit)",
        "Responsibilities (utilities, maintenance, and repairs)",
        "Consent and notices (landlord's consent, and notice requirements)",
        "Special provisions (furniture, parking, and subletting restrictions)",
    ]

    Sources23

    3.Define "good" with the owning team before you build

    The last requirement is the one most often skipped: a measurable definition of success, agreed with the people who own the process. Both support guides give the same instruction, to work with the support team on success criteria with measurable benchmarks and goals. The insurance example sets targets for each task. Comprehension of what the customer meant should reach 95% or higher, measured by reviewing a sample of conversations. Response relevance should reach 90% or above, scored with LLM-based grading so it can scale. The ticket-routing guide proposes targets like these:

    Example success criteria for LLM ticket routing
    CriterionHow to measureSuggested target
    Classification consistencyPeriodically test with a set of standardized inputs95% or higher
    Adaptation to new categoriesIntroduce new ticket types and measure time to satisfactory accuracyWithin 50–100 sample tickets
    Multilingual handlingCompare routing accuracy across languagesNo more than a 5–10% drop for non-primary languages
    Edge casesBuild a test set of edge cases and measure routing accuracyAt least 80%
    Bias across customer groupsRegularly audit routing decisionsAccuracy within 2–3% across groups
    Explanation qualityHuman raters score explanations, for example 1–5Average of 4 or higher

    Targets should follow the stakes. In the insurance example, general company and product information is judged against what Claude was given in context, with a target of 100% accuracy. Quote generation also targets 100%, because the guide calls it vital information for a successful interaction. Some outputs resist a single number: the legal guide admits that judging summaries is a notoriously difficult task and that different readers value different things. That is a reason to agree the criteria with stakeholders in advance, not to skip them.

    Grounding is part of what makes a criterion testable. When Claude produces requirements with no internal documents to work from, there is nothing in context to check its accuracy against. Its draft may read well and still not describe your organisation.

    A product manager is using Claude to help define the minimum viable product (MVP) for a new feature. Which approach should they use?

    Sources143

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A customer support chatbot is one task, so it needs one prompt and one overall quality check.Why is that wrong?

      A support chat bundles question answering, information retrieval and action-taking. Breaking it into tasks lets each one be prompted and evaluated.

      Covered in Outline the ideal interaction, then split it into tasks

    2. 2.Requirements examples only need to cover the typical cases the system will see most often.Why is that wrong?

      The guides say to include edge cases and difficult scenarios from the start, and to test edge-case accuracy separately.

      Covered in Pin down categories, examples and what to extract

    3. 3.You can judge whether Claude is good enough once you have seen what it produces.Why is that wrong?

      Success criteria, with measurable benchmarks and goals, are defined with the team that owns the process before building starts.

      Covered in Define "good" with the owning team before you build

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Outline an ideal customer interaction to define how and when you expect the customer to interact with Claude.”
      ↩︎ Outline the ideal interaction, then split it into tasks
      “This outline will help to determine the technical requirements of your solution.”
      ↩︎ Outline the ideal interaction, then split it into tasks
      “This ensures you can prompt and evaluate Claude for every task”
      ↩︎ Outline the ideal interaction, then split it into tasks
      “Work with your support team to define success criteria and write detailed evaluations with measurable benchmarks and goals.”
      ↩︎ Define "good" with the owning team before you build
      “Evaluate a set of conversations and rate the relevance of each response (using LLM-based grading for scale).”
      ↩︎ Define "good" with the owning team before you build
      “based on the information provided to Claude in context. Target 100% accuracy in this introductory information.”
      ↩︎ Define "good" with the owning team before you build
      “Customer support chat is a collection of multiple different tasks, from question answering to information retrieval to taking action on requests”
      ↩︎ Exam trap 1
      “Work with your support team to define success criteria and write detailed evaluations with measurable benchmarks and goals.”
      ↩︎ Exam trap 3
    2. 2.
      “first create examples of content that should be flagged and content that should not be flagged.”
      ↩︎ Pin down categories, examples and what to extract
      “review your examples to create a well-defined list of moderation categories.”
      ↩︎ Pin down categories, examples and what to extract
      “Ensure that you include edge cases and challenging scenarios that may be difficult for a content moderation system to handle effectively.”
      ↩︎ Exam trap 2
    3. 3.
      “Without clear direction, it can be difficult for Claude to determine which details to include.”
      ↩︎ Pin down categories, examples and what to extract
      “Evaluating the quality of summaries is a notoriously challenging task.”
      ↩︎ Define "good" with the owning team before you build
    4. 4.
      “Create a test set of edge cases and measure the routing accuracy, aiming for at least 80% accuracy on these challenging inputs.”
      ↩︎ Define "good" with the owning team before you build

    Ready to test yourself?

    Practise the 27 questions on this subdomain.