What you will be able to do
- Write success criteria for a Claude-assisted solution that are specific, measurable, achievable and relevant, before any build starts
- Choose the simplest design that can meet those criteria: a single call, a fixed workflow, or an agent
- Decide when a design needs tools and when it does not
- Brief Claude the way you would brief a capable new colleague
Key concept
Success criteria before solution — Before you design or build anything with Claude, write down what a good result looks like, in terms you can measure. Every later choice, from which architecture to use to whether a revision helped, is judged against those criteria.
1.Start by defining what "good" looks like
When you use Claude to help design a solution, the first thing to produce is not a prompt or a pipeline. It is a statement of what a good outcome is. Anthropic's prompt engineering guide takes this for granted: it assumes you already have a clear definition of success for your use case and some way to test against it, and it tells you to set those up first if you have not. Without them you have nothing to judge a design choice by, and later you have no way to tell whether a change made things better or worse.
Anthropic lists four properties of good criteria. They are specific ("accurate sentiment classification", not "good performance"). They are measurable, using numbers or well-defined qualitative scales such as Likert ratings or expert rubrics. They are achievable, based on benchmarks, earlier experiments or expert knowledge rather than wishful thinking. And they are relevant to the application's purpose. For example, citation accuracy may be critical in a medical app and matter much less in a casual chatbot.
| Version | Criterion | What it gives you |
|---|---|---|
| Bad | The model should classify sentiments well | Nothing to measure, and no way to tell a pass from a fail |
| Good | F1 score of at least 0.85 on a held-out test set of 10,000 diverse Twitter posts, a 5% improvement over the current baseline | A number (measurable), a named task (specific), representative data (relevant) and a realistic target (achievable) |
| Good, multidimensional | F1 ≥ 0.85; 99.5% of outputs non-toxic; 90% of errors cause inconvenience, not egregious error; 95% response time < 200ms | Quality, safety, error severity and speed are each checked separately |
The third row matters because one number rarely covers everything. Anthropic's list of dimensions to consider includes task fidelity and edge cases, consistency across similar inputs, relevance and coherence, whether tone and style suit the audience, how personal or sensitive information is handled, how well the model uses the context it is given, latency, and cost per call at your expected volume. Choose the ones your stakeholders actually care about and write a threshold for each.
Over three weeks a team refined the instructions for a Claude-drafted internal newsletter until the editor was consistently happy with every draft. Nobody asked the readers. What is the greatest risk?
Correct answer: B — The drafts are now tuned to one editor's taste, and no one has checked that they serve the readers.
- A. Long instructions are an onboarding inconvenience that can be tidied up. They make the arrangement harder to hand over, but they do not make the newsletter less effective for the people reading it.
- B. Correct. Iteration improves whatever it is measured against, and here the only measure was one editor's approval. Three weeks of refinement have optimised for a proxy, so the drafts could satisfy the editor completely while serving readers worse than before.
- C. The refinement effort is a sunk cost that will be repaid if the arrangement works. Time spent iterating is expected; the problem is what the iteration was aimed at.
- D. Concentrating judgment in one person is a genuine weakness in the design, and it follows from the same root cause. It is a consequence of tuning to one taste rather than the primary danger, which is that readers were never consulted.
2.Choose the simplest design that meets the criteria
Once the criteria exist, Claude can help you compare design options. Anthropic's guidance on agentic systems is to find the simplest solution that works and add complexity only when it is needed, and that can mean not building an agentic system at all. Its reasoning is that agentic systems often trade latency and cost for better task performance, so you should add them only when that trade is worth making against your criteria.
| Design | What it is | When it fits |
|---|---|---|
| Single optimised LLM call | One call, improved with retrieval and in-context examples | Many applications: Anthropic says this is usually enough |
| Workflow | LLMs and tools orchestrated through predefined code paths, e.g. prompt chaining (outline → check outline → write document) or routing by input type | Well-defined tasks that split cleanly into fixed subtasks, where you need predictability and consistency |
| Agent | The LLM directs its own process and tool use | When flexibility and model-driven decision-making are needed at scale |
A second design question is whether the solution needs tools. Anthropic's tool-use documentation says tools fit when the task needs something text alone cannot provide: actions with side effects (sending an email, updating a record), fresh or external data, outputs that must have a guaranteed structure, or calls into existing systems such as databases and internal APIs. Tools do not fit when the model can answer from its training alone (summarisation, translation), when the exchange is one-shot Q&A with nothing to execute, or when the extra round trip would take longer than the work itself. The model never executes anything itself. It emits a structured request, and your code runs the operation.
The decision should have been a tool call. Anthropic calls regex extraction of a decision from free text a clear sign that the structure belongs in a tool schema, which enforces the shape rather than hoping the prose contains it.
3.Brief Claude as you would a capable new colleague
Whatever design you pick, Claude's contribution depends on the brief it gets. Anthropic's prompting guidance suggests treating Claude like a brilliant new employee who does not yet know your norms and workflows: the more precisely you explain what you want, the better the result. Its golden rule is a practical test. Show your prompt to a colleague who has little context on the task and ask them to follow it. If they would be confused, Claude will be too.
Two habits follow from this. First, be explicit about ambition. "Create an analytics dashboard" gets less than a request that asks for as many relevant features and interactions as possible, going beyond the basics. If you want Claude to go above and beyond, you have to ask for it. Second, give the reason behind a constraint. "NEVER use ellipses" works less well than explaining that the response will be read aloud by a text-to-speech engine that cannot pronounce them, because Claude can generalise from the explanation. When order or completeness matters, give the instructions as numbered steps.
Sources5
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Success criteria can be worked out after the pilot, once you have seen what Claude produces.Why is that wrong?
Anthropic's guidance assumes the criteria and a way to test against them already exist before prompt work starts. Without them you cannot judge whether a design or a revision is an improvement.
Covered in Start by defining what "good" looks like
2.A flexible autonomous agent is the safest default because it can handle anything the process throws at it.Why is that wrong?
The recommendation is the simplest solution that meets the need. Fixed workflows give predictability for well-defined tasks, and for many applications a single well-built call is enough.
Covered in Choose the simplest design that meets the criteria
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Specific: Clearly define what you want to achieve.”
↩︎ Start by defining what "good" looks like“Most use cases need multidimensional evaluation along several success criteria.”
↩︎ Start by defining what "good" looks like“Building a successful LLM-based application starts with clearly defining your success criteria and then designing evaluations to measure performance against them.”
↩︎ Key concept - 2.
“Some ways to empirically test against those criteria”
↩︎ Start by defining what "good" looks like“A clear definition of the success criteria for your use case”
↩︎ Exam trap 1 - 3.https://www.anthropic.com/engineering/building-effective-agentsSecondary source
“we recommend finding the simplest solution possible, and only increasing complexity when needed.”
↩︎ Choose the simplest design that meets the criteria“For many applications, however, optimizing single LLM calls with retrieval and in-context examples is usually enough.”
↩︎ Choose the simplest design that meets the criteria“workflows offer predictability and consistency for well-defined tasks, whereas agents are the better option when flexibility and model-driven decision-making are needed at scale.”
↩︎ Exam trap 2 - 4.
“if you're writing a regex to extract a decision from model output, that decision should have been a tool call.”
↩︎ Choose the simplest design that meets the criteria“The model never executes anything on its own.”
↩︎ Choose the simplest design that meets the criteria - 5.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practicesOfficial docs
“Think of Claude as a brilliant but new employee who lacks context on your norms and workflows.”
↩︎ Brief Claude as you would a capable new colleague“Golden rule: Show your prompt to a colleague with minimal context on the task and ask them to follow it.”
↩︎ Brief Claude as you would a capable new colleague“Claude is smart enough to generalize from the explanation.”
↩︎ Brief Claude as you would a capable new colleague