What you will be able to do
- Explain why a few concrete input/output examples beat a longer prose description when Claude's results keep varying
- Write two or three examples that are relevant, diverse and clearly separated from instructions
- Use the interview pattern to surface design considerations before Claude implements anything in an unfamiliar domain
- Have Claude write a test suite covering expected behaviour, edge cases and performance requirements first, then iterate by sharing the failures
- Decide whether to send several issues in one detailed message or fix them one at a time, based on whether the fixes interact
Key concept
Concrete signals over prose — Iterative refinement works when you give Claude something concrete to aim at and check against, such as examples, test cases or answers to its questions, instead of a longer description. Stating that target before Claude starts costs less than correcting it afterwards.
1.Why a prose description comes back different each time
Words like 'consistently', 'clean' or 'normalised' feel precise to the person who writes them. Each one still leaves a decision open, and Claude can settle that decision differently on different runs. Adding more adjectives just adds more words to interpret. Anthropic's prompting guidance says to show the target instead of describing it: "Examples are one of the most reliable ways to steer Claude's output format, tone, and structure." A few examples give Claude actual inputs and outputs to work from, so it no longer has to guess what your wording meant.
The ticket-routing guide in Anthropic's docs makes the same point from a production angle. It describes examples as the most effective lever it has, "Despite providing examples being the most effective way to improve performance". When Claude misreads a particular kind of input, it recommends explicit instructions or examples showing how that edge case should be handled. It also notes that adding the reasoning behind a tricky example helps Claude generalise the logic to cases you did not show.
2.Writing two or three examples that pin down the transformation
The docs say "a few" examples, which lines up with the exam's two or three. How many matters less than how the examples are chosen. The prompting guide asks for three properties. First, relevance: "Mirror your actual use case closely." Second, diversity: "Cover edge cases and vary enough that Claude doesn't pick up unintended patterns." Third, structure: wrap them in <example> tags so Claude can tell them apart from your instructions.
Diversity is the property people most often get wrong. Suppose all three sample inputs are well-formed and about the same length. Claude may learn that incidental similarity and not the rule you meant. Make at least one example the awkward case, such as a malformed value or an empty field. That one example resolves the ambiguity your prose left open.
| Vague request | With concrete cases | What the concrete version adds |
|---|---|---|
| implement a function that validates email addresses | write a validateEmail function. example test cases: user@example.com is true, invalid is false, user@.com is false. run the tests after implementing | Three input/output pairs, one of them a near-miss, plus an instruction to check the result |
| add tests for foo.py | write a test for foo.py covering the edge case where the user is logged out. avoid mocks. | A named edge case and a stated testing preference |
It looks almost like a valid address, and 'validate email addresses' in prose doesn't say whether it should pass. Including it with an expected result of false settles a question the description never answered.
Sources1
3.The interview pattern: have Claude ask questions before it builds
Examples work when you already know what the output should look like. In an unfamiliar domain you often don't. Say you are adding a cache for the first time. You may not have decided when entries become stale or what should happen if the backing store fails, and you may not know those questions need answering. The interview pattern reverses the usual direction: before writing any code, Claude asks you questions. Claude Code's best-practices page gives the prompt:
I want to build [brief description]. Interview me in detail using the AskUserQuestion tool.
Ask about technical implementation, UI/UX, edge cases, concerns, and tradeoffs. Don't ask obvious questions, dig into the hard parts I might not have considered.
Keep interviewing until we've covered everything, then write a complete spec to SPEC.md.Two parts of that prompt do most of the work. "dig into the hard parts I might not have considered" tells Claude to raise things you have missed, not to confirm what you already said. The final line turns the answers into a written spec, which the implementation can then follow. An Anthropic field guide explains why this is worth doing before implementation: learning about unknowns during implementation "can be (relatively) expensive", because a small change to the spec can lead to very different code. The same guide suggests pointing the interview at the decisions that matter most: "prioritize questions where my answer would change the architecture." In the caching example, invalidation strategy and failure behaviour are exactly those questions.
4.Test-driven iteration: write the test suite first, then share the failures
Test-driven iteration is the executable form of the same idea. Examples pin down a transformation in words; a test suite makes the target something Claude can run and check its own work against. The order is the whole technique: you write the test suite first, before any implementation exists, and then iterate by sharing test failures until the suite passes. Anthropic's guidance on Claude Code on the web describes it directly: "Let Claude write tests that define the expected behavior, then implement the code to make those tests pass." The suite written up front should cover three things: the expected behaviour, the edge cases, and any performance requirements. The rate-limiter prompt in that guidance shows how to state them before implementation. It gives the expected behaviour (allow 100 requests per minute per API key), the edge cases (return 429 when the limit is exceeded, reset after 60 seconds, track keys independently), and it turns the limit itself into a measurable requirement a test can assert. It then closes with "Use a TDD approach: write comprehensive tests first, then implement the rate limiting logic to pass them." A latency or throughput bound would be written into the same suite in the same way, as a test that fails until the implementation meets it.
Once the suite exists, iterating means sharing test failures rather than re-describing the problem. Each round of improvement is guided by the failing tests: the guidance says "You don't need to monitor Claude's progress since the tests will catch issues and guide iteration toward a working solution." When Claude runs the suite itself, it "can iterate on the implementation without your supervision, using test failures to identify and fix problems." When you run the suite yourself, you do the same thing by hand: paste the failing test's output into the conversation. Each failure names the input, the expected output and the actual output, which is exactly the concrete signal this lesson is about, and each round of sharing failures moves the implementation one step closer. This is progressive improvement in the literal sense: the suite stays fixed, and the number of failing tests goes down with each iteration. The same guidance recommends keeping a suite in the repository for this reason: "Consider adding a test suite to your repository so Claude more easily verify that it has successfully completed a task". Claude Code's best-practices page applies the pattern to bug fixes too: describe the symptom, point at the likely location, and ask Claude to "write a failing test that reproduces the issue, then fix it".
Sharing a failure also fixes edge-case handling after the fact. Suppose a migration script mishandles rows where a column is null, and a prose description of the problem produced a fix that covers one code path but not another. Rather than describing the bug again, give Claude a specific test case: the example input row with the null value and the exact expected output. The support article's example does this for a validation step: "Users with no email are crashing the validation step. Make it handle that gracefully and add a test." The test turns the edge case into something Claude can run, not just read, and the next iteration is guided by whether that test passes.
Tests written first define the target before any code exists, so the implementation is checked against your intent rather than the tests being fitted to whatever Claude built. They also give Claude failures to iterate on without you re-describing the problem each round.
5.One message or one fix at a time: interacting versus independent issues
After a test run or a review you often hold several problems at once. The question is whether to send them all in one message or take them one at a time. The support guidance describes the sequential case: "If the first answer is off, you do not need to rephrase the whole request. Simply say what is wrong", and Claude "will keep everything else and adjust only that point." That is the right approach when the issues are independent, such as a naming fix, a missing log line and an unrelated typo. Each correction leaves the rest of the work untouched, so nothing is lost by taking them in turn, and each small change is easy to review.
Interacting issues are different. If the fix for one changes the shape of the fix for another, say a cache invalidation bug and a stale-read bug in the same code path, fixing them one at a time invites a patch that quiets the first symptom and makes the second worse. Claude Code's best-practices page asks for the root cause instead: paste the error, then "fix it and verify the build succeeds. address the root cause, don’t suppress the error". When fixes interact, put every issue into a single detailed message, with the full error output and a sentence on how the problems relate. Claude can then design one change that covers them together, and the same page notes you can ask it to run the check and iterate in the same message.
| Situation | How to send it | Why |
|---|---|---|
| Independent issues (a rename, a missing log line, an unrelated typo) | One at a time: say what is wrong and let Claude adjust only that point | Each fix leaves the rest untouched, so small sequential changes are easy to review |
| Interacting issues (two bugs in the same code path where one fix changes the other) | All at once, in a single detailed message with the pasted errors and how they relate | A one-at-a-time patch can suppress a symptom instead of fixing the root cause |
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.If Claude keeps misreading a transformation, rewrite the description in more detail.Why is that wrong?
More prose adds more to interpret. Two or three concrete input/output examples, including an edge case, give Claude a target it can generalise from.
Covered in Why a prose description comes back different each time
2.The interview pattern means Claude asks a few clarifying questions about requirements you already stated.Why is that wrong?
The point is to bring up things you had not thought about, such as edge cases, failure modes and tradeoffs, before any code is written.
Covered in The interview pattern: have Claude ask questions before it builds
3.Have Claude implement the feature first, then ask it to add tests that cover what it built.Why is that wrong?
Test-driven iteration writes the suite first so it defines the expected behaviour, edge cases and performance requirements, and then uses test failures to drive each round of improvement.
Covered in Test-driven iteration: write the test suite first, then share the failures
4.Always report issues one at a time so each change stays small and reviewable.Why is that wrong?
Sequential fixes suit independent issues. When fixes interact, a one-at-a-time patch can suppress a symptom rather than fix the root cause, so those issues belong in a single detailed message.
Covered in One message or one fix at a time: interacting versus independent issues
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practicesOfficial docs
“Examples are one of the most reliable ways to steer Claude's output format, tone, and structure.”
↩︎ Why a prose description comes back different each time“Diverse: Cover edge cases and vary enough that Claude doesn't pick up unintended patterns.”
↩︎ Writing two or three examples that pin down the transformation“Relevant: Mirror your actual use case closely.”
↩︎ Writing two or three examples that pin down the transformation“A few well-crafted examples (known as few-shot or multishot prompting) improve accuracy and consistency.”
↩︎ Exam trap 1 - 2.
“Despite providing examples being the most effective way to improve performance”
↩︎ Why a prose description comes back different each time - 3.https://code.claude.com/docs/en/best-practicesOfficial docs
“Ask about technical implementation, UI/UX, edge cases, concerns, and tradeoffs.”
↩︎ The interview pattern: have Claude ask questions before it builds“write a failing test that reproduces the issue, then fix it”
↩︎ Test-driven iteration: write the test suite first, then share the failures“fix it and verify the build succeeds. address the root cause, don’t suppress the error”
↩︎ One message or one fix at a time: interacting versus independent issues“In one prompt: ask Claude to run the check and iterate in the same message, as in the table above.”
↩︎ One message or one fix at a time: interacting versus independent issues“Don't ask obvious questions, dig into the hard parts I might not have considered.”
↩︎ Exam trap 2“address the root cause, don’t suppress the error”
↩︎ Exam trap 4 - 4.
“prioritize questions where my answer would change the architecture”
↩︎ The interview pattern: have Claude ask questions before it builds“finding them out during implementation can be (relatively) expensive”
↩︎ The interview pattern: have Claude ask questions before it builds - 5.
“Let Claude write tests that define the expected behavior, then implement the code to make those tests pass.”
↩︎ Test-driven iteration: write the test suite first, then share the failures“Use a TDD approach: write comprehensive tests first, then implement the rate limiting logic to pass them.”
↩︎ Test-driven iteration: write the test suite first, then share the failures“You don't need to monitor Claude's progress since the tests will catch issues and guide iteration toward a working solution.”
↩︎ Test-driven iteration: write the test suite first, then share the failures“can iterate on the implementation without your supervision, using test failures to identify and fix problems.”
↩︎ Test-driven iteration: write the test suite first, then share the failures“Consider adding a test suite to your repository so Claude more easily verify that it has successfully completed a task”
↩︎ Test-driven iteration: write the test suite first, then share the failures“Let Claude write tests that define the expected behavior, then implement the code to make those tests pass.”
↩︎ Exam trap 3 - 6.https://support.claude.com/en/articles/14553240-give-claude-context-claude-md-and-better-promptsOfficial docs
“Users with no email are crashing the validation step. Make it handle that gracefully and add a test.”
↩︎ Test-driven iteration: write the test suite first, then share the failures“If the first answer is off, you do not need to rephrase the whole request. Simply say what is wrong”
↩︎ One message or one fix at a time: interacting versus independent issues“Claude will keep everything else and adjust only that point.”
↩︎ One message or one fix at a time: interacting versus independent issues“Stating acceptance criteria up front is more efficient than several rounds of revision.”
↩︎ Key concept