What you will be able to do
- Run a discovery conversation that identifies the specific users, the outcome they need, and the real constraints before any AI design choice is made
- Break an AI use case down into the discrete tasks the system must perform, so each one can be prompted and evaluated
- Turn vague stakeholder language into measurable requirements and acceptance tests that code can pass or fail
- Spot where a requirement was lost between the user's need and the built system, and tell which link to go back to
Key concept
The Description Chain — Requirement gathering is a chain of translations: from what the user says, to a requirement, to a technical spec, to instructions for the AI, to tests. Discovery is where the chain starts. Anything the architect fails to capture there cannot be recovered by better prompting later.
1.Discovery comes before any AI decision
Stakeholders usually arrive with a solution already in mind: "we need a chatbot", "we want an agent that handles tickets". Structured discovery exists to put off that decision until the problem is understood. The architect's first job is to decompose the problem. Only then do they decide what role, if any, a model plays at each step.
The order matters because the same request can lead to very different architectures. It might need a single well-prompted model call, a fixed workflow, or an autonomous agent. Anthropic's guidance on building agents recommends starting with the simplest solution that works and adding complexity only when it clearly pays for itself. You can't apply that principle until you know what the problem actually is. A discovery session that opens with "which model should we use?" has skipped the step that answers it.
2.The four questions a discovery session must answer
Anthropic's builder curriculum frames the pre-build work as four plain-language questions. Each one exists to surface a different kind of requirement, and skipping any of them leaves a predictable gap. "Patients" is not a specific enough user group. It hides the staff who will update the data, the people without smartphones, and whoever is accountable for privacy. "A wait time checker" is a feature, not an outcome. The outcome might be "avoid sitting in a waiting room with a sick child", and that changes what a good answer looks like.
| Discovery question | What it surfaces | Gap if skipped |
|---|---|---|
| Who are the users? Be specific. | Every group that touches the system, including staff and secondary users | Designing for one persona while others break it |
| What do they actually need? | The outcome the user hopes for, not the feature they named | Building exactly what was asked for and still missing the need |
| What does a great solution feel like to use? | Experience expectations: speed, effort, context of use | Technically correct output that users abandon |
| What are the real constraints? | Budget, technical skill, device access, privacy requirements | An architecture the organisation cannot run or is not allowed to run |
Feature lists repeat the solution the stakeholder has already imagined. Descriptions of the experience reveal the underlying requirements, such as "checked from a phone in a parking lot in under a minute". Those constrain the design far more usefully and leave the architect free to pick the right mechanism.
Sources1
3.Decompose the use case into testable tasks
Once the need is clear, break it into the individual tasks the system must perform. Anthropic's customer-support guide does this before any build. One "support chat" turns out to contain greeting, product questions, staying on topic, and quote generation that calls an external API. Each task has its own prompt, its own failure modes, and its own evaluation.
Decomposition also shows where the AI boundaries are. Some tasks are pure language work. Others need a tool call or data from a system of record. Some, like presenting a price, may need a human or a deterministic system rather than a model. Each of these is a requirement to capture now, not a surprise to find in the pilot.
During a discovery call, a customer support team says they want to route 50,000 short ticket-triage requests per day, need sub-second responses to keep chat wait times low, and have a limited annual budget for the pilot. Which model selection approach best matches these discovery findings?
Correct answer: A — Begin implementation with Claude Haiku 4.5, test the triage prompts against production ticket samples, and upgrade only if it fails to reach the required accuracy threshold
- A. Correct. Anthropic's guidance recommends starting with a fast, cost-effective model like Haiku 4.5 for high-volume, straightforward, latency-sensitive, cost-sensitive tasks, then upgrading only if a capability gap is found through testing.
- B. Incorrect. Opus 4.8 at xhigh effort is positioned for complex agentic coding and enterprise work requiring the deepest reasoning, not high-volume, sub-second, budget-constrained ticket triage.
- C. Incorrect. Extended thinking on every request adds latency and cost that conflicts with the stated sub-second response requirement and limited budget.
- D. Incorrect. Claude Fable 5's large context window addresses processing large volumes of text per request, not the discovery findings of low latency and tight budget for short triage prompts.
Sources3
4.Make every requirement measurable
Stakeholders speak in adjectives: fast, accurate, simple, safe. Each adjective is a decision that hasn't been made yet. Structured requirement gathering turns each one into something you could defend and check. "Fast" becomes a latency target. "Accurate" becomes an error rate on a named test set. "Simple" becomes a description of the user's actual situation.
The strongest form of this is the acceptance test, written before implementation, that code can pass or fail. Tests give the stakeholders, the engineering team and the AI the same definition of done. Good test sets include edge cases the stakeholder didn't mention, such as no data, a closed service, or a zero value. Asking about those cases during discovery is often what uncovers the hidden requirements. For AI systems, work with the owning team to set success criteria and evaluations with measurable benchmarks, so that "is it good enough for production?" has an agreed answer before the pilot starts.
A pharmaceutical research team describes a discovery requirement: multi-step literature synthesis across scientific papers where accuracy outweighs response cost, and the workflow will run as a long, high-autonomy agent that plans its own steps. Which model selection approach best fits these gathered requirements?
Correct answer: A — Start with Claude Opus 4.8, since complex reasoning and advanced agentic work are named as strengths for high-autonomy, accuracy-critical tasks
- A. Correct. The model selection matrix lists complex agentic coding and enterprise work, along with complex reasoning and advanced research, as fits for Opus 4.8, matching the accuracy-first, high-autonomy discovery findings.
- B. Incorrect. Starting with the fastest, cheapest model is recommended when cost or latency dominate, but here the stated requirement is that accuracy outweighs cost considerations.
- C. Incorrect. Reducing effort trades intelligence for latency and cost, which works against a requirement where accuracy is explicitly prioritized over cost.
- D. Incorrect. The Batch API is a cost- and throughput-optimization mechanism for asynchronous volume, not a lever Anthropic ties to improving reasoning accuracy.
5.Validate the chain with real users, not just tests
Measurable requirements are necessary but not sufficient. Tests only check what you wrote down. If discovery captured the wrong intent, the tests will pass happily while users stay unhappy. That mismatch is a diagnostic signal: it tells you the break is upstream, between what the user said and what was written as a requirement.
This is why discovery doesn't end when the requirements document is signed off. Put an early build in front of a real user without explaining it, and compare what you intended with what they experienced. When the gap appears, trace it back to the link where the requirement was lost instead of patching the prompt or the code.
A solutions architect is running the first discovery session with a new prospect before recommending any specific Claude model. According to Anthropic's model-selection guidance, which factors should the architect establish as key criteria during this session?(Select 3)
Correct answers: A, B, C — The specific capabilities the model must have to meet the prospect's needs; How quickly the model must respond within the prospect's application; The prospect's available budget for development and production usage
- A. Correct. Capabilities is one of the key criteria Anthropic's guidance lists to establish before choosing a model.
- B. Correct. Speed, meaning how quickly the model must respond, is explicitly named as a key criterion to evaluate first.
- C. Correct. Cost, meaning the budget for development and production usage, is explicitly named as a key criterion to evaluate first.
- D. Incorrect. Headcount in the IT department is not one of the documented model-selection criteria and does not inform capability, speed, or cost decisions.
- E. Incorrect. Marketing channel usage is unrelated to model capability, speed, or cost requirements and is not a discovery criterion for model selection.
- F. Incorrect. A third-party ticketing software version is an integration detail, not one of the capability, speed, or cost criteria used to scope model selection.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Discovery starts by selecting the model or agent framework that best fits the stakeholder's request.Why is that wrong?
The problem is decomposed first. Only then is AI's role decided at each step, and the answer may be a simple workflow rather than an agent.
Covered in Discovery comes before any AI decision
2.Identifying the primary end user is enough to define the user requirements.Why is that wrong?
Discovery has to name every group that touches the system, such as staff, secondary users and constrained users. Otherwise requirements owned by those groups are silently missed.
Covered in The four questions a discovery session must answer
3.Qualitative requirements like "fast" or "simple" are fine to carry into design and can be refined during the build.Why is that wrong?
Every descriptive word in a requirement has to become a concrete, defensible decision, such as a threshold or a defined user context, before it can be designed or tested against.
Covered in Make every requirement measurable
4.If all acceptance tests pass, the requirements were gathered correctly.Why is that wrong?
Tests only confirm what was written down. Passing tests alongside unhappy users means the intent was captured wrongly upstream, during discovery.
Covered in Validate the chain with real users, not just tests
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://academy.claude.com/courses/ai-fluency-for-builders/delegation-the-builder-s-toolkitOfficial docs
“Delegation means decomposing the problem first, then deciding what AI handles at each step.”
↩︎ Discovery comes before any AI decision“What are the real constraints? Think budget, technical skill at the clinic, patient access to devices, privacy requirements.”
↩︎ The four questions a discovery session must answer“Describe the experience, not the features.”
↩︎ The four questions a discovery session must answer“Write acceptance tests before code. They give you and AI a shared definition of done.”
↩︎ Make every requirement measurable“Delegation means decomposing the problem first, then deciding what AI handles at each step.”
↩︎ Exam trap 1“Who are the users? Be specific.”
↩︎ Exam trap 2 - 2.https://academy.claude.com/courses/ai-fluency-for-builders/description-building-great-thingsOfficial docs
“AI cannot hear what the user did not say.”
↩︎ Discovery comes before any AI decision“Every adjective should be a decision you can defend.”
↩︎ Make every requirement measurable“A passing test with an unhappy user means you described the wrong intent.”
↩︎ Validate the chain with real users, not just tests“The Description Chain connects user voice to requirement to technical spec to AI instruction. Prompt engineering is only one link.”
↩︎ Key concept“Every adjective should be a decision you can defend.”
↩︎ Exam trap 3“A passing test with an unhappy user means you described the wrong intent.”
↩︎ Exam trap 4 - 3.
“Before you start building, break down your ideal customer interaction into every task you want Claude to be able to perform.”
↩︎ Decompose the use case into testable tasks“This ensures you can prompt and evaluate Claude for every task”
↩︎ Decompose the use case into testable tasks“Work with your support team to define success criteria and write detailed evaluations with measurable benchmarks and goals.”
↩︎ Make every requirement measurable - 4.
“A failure at any link cascades downstream.”
↩︎ Validate the chain with real users, not just tests