What you will be able to do
- Name the kind of failure behind a poor output before changing the prompt
- Fix a generic or unhelpful answer by adding context, motivation and explicit instructions
- Fix a response that is too brief by retooling the prompt instead of blaming the context window
- Recognise long-conversation drift and re-supply the context Claude has lost
Key concept
Diagnose before you fix — A poor output is a symptom. Before you change anything, work out what kind of failure you are looking at, such as a vague ask, missing context, a conversation that has drifted, or a knowledge gap. Each cause has a different fix, and the wrong fix leaves the problem in place.
1.Name the failure before you change the prompt
When an output disappoints, the obvious move is to reword the request and try again. That only works if wording was the problem. Anthropic's course on AI capabilities and limitations treats troubleshooting as diagnosis. It says most real failures come from two underlying properties interacting, not one, and that once you can name which two, you know which fix to use.
The course names two pairs that come up again and again. The first is next token prediction combined with knowledge, which produces hallucinated specifics: fluent, confident details that are wrong. The second is working memory combined with steerability, which produces long-conversation drift: a thread that slowly stops following what you asked. The fixes it lists match the causes. You verify specifics, re-supply context, offload work to code execution, or invite pushback. Rewording on its own is not in that list.
| What you see | Likely cause | Targeted fix |
|---|---|---|
| Generic answer that could fit almost any situation | The prompt lacks context and specifics | Add background, audience and motivation; state the output you want |
| Answer is too short or stops partway | The prompt did not ask for depth | Retool the prompt; give Claude the previous prompt and response so it can continue |
| A long thread mixes up details and contradicts earlier turns | Working Memory + Steerability (long-conversation drift) | Re-supply the context instead of piling on corrections |
| Confident, specific details that turn out to be wrong | Next Token Prediction + Knowledge (hallucinated specifics) | Verify specifics against a source |
Sources1
2.Generic answers come from generic prompts
The most common cause of an unhelpful answer is a prompt that leaves Claude guessing. Anthropic's help article on unhelpful answers gives four principles. Explain your ask simply and clearly. Include as much context as possible. Break complex requests into substeps. Give feedback in follow-up turns. The key test on context is to imagine you are briefing someone who knows nothing about your situation.
The prompting best-practices guide gives Claude the same framing: treat it as a brilliant but new employee who lacks context on your norms and workflows. Its golden rule is a quick diagnostic you can run on any prompt that underperformed. Show the prompt to a colleague with minimal context and ask them to follow it. If they would be confused, Claude will be too. The guide also warns that Claude will not infer ambition from a thin request. If you want output that goes beyond the basics, you have to ask for it.
Create an analytics dashboard. Include as many relevant features and interactions as possible. Go beyond the basics to create a fully-featured implementation.The reason behind an instruction matters as much as the instruction itself. The guide's example rule is "never use ellipses". It gets better results when the prompt explains that the response will be read aloud by a text-to-speech engine that cannot pronounce them. Claude generalises from the explanation, so it also handles related cases the bare rule never mentioned.
Context and specifics. That means who the event is for, what it is, and why it matters, plus the format and tone you need. Think of the new employee who has never heard of your event. Everything they would have to ask you belongs in the prompt.
A support-ticket triage prompt gives Claude a single sentence: "Categorize this ticket." Categories come back inconsistent between runs, and the team wants a fix that addresses the root cause rather than adding post-processing. What should they change first?
Correct answer: C — Rewrite the prompt to list the exact categories, the desired output format, and the criteria for choosing between them.
- A. Incorrect. Temperature affects sampling randomness, not whether the model understands what categories exist or how to choose between them; it does not address an underspecified prompt.
- B. Incorrect. Swapping models does not resolve ambiguity in the instructions themselves; an underspecified prompt will still under perform on a larger model.
- C. Correct. Claude responds well to clear, explicit instructions; a vague one-line prompt lacking the category list, output format, and selection criteria is the classic underperformance pattern, and the documented fix is to be specific about desired output and constraints.
- D. Incorrect. Manual relabeling is a workaround applied after generation, not a fix to the prompt causing the inconsistency.
3.When the response is too brief
A short or partial answer is easy to blame on the model or on capacity. Anthropic's help article on brief responses corrects the capacity idea directly: the context window applies to the prompts you provide, not to the output Claude generates. So a three-sentence answer is not a sign that Claude ran out of room. The article's advice is to retool the prompt.
In practice, retooling means stating the depth you need, just as the prompting guide advises being specific about the desired output format and constraints. If you need several paragraphs for a named audience, say so. If the order or completeness of parts matters, list them as numbered steps. When a response stops partway, the article recommends giving Claude the previous prompt and response in your next prompt so it can pick up where it left off. That is better than regenerating from scratch and hoping for a longer answer.
4.When the conversation itself is the problem
Sometimes the prompt is fine and the thread is not. A conversation that has run for days across several related tasks accumulates details from all of them. The course calls this long-conversation drift, the collision of working memory and steerability. The symptoms are distinctive. Claude blends in details from an earlier task, loses instructions you gave near the start, or contradicts something it said earlier.
The targeted fix, from the course's list, is to re-supply context. Adding more corrections to a thread that has already lost its focus deals with the symptom, not the cause. Restate what matters now, clearly and in one place, and separate it from the unrelated work that is confusing the thread. The help article's advice to break complex requests into substeps points the same way: one coherent task at a time is easier to steer than several tangled together.
A marketing writer's prompt asks Claude to "write engaging product descriptions." The outputs are technically correct but read as flat and generic compared to what the writer had in mind. Which change is most likely to close that gap?
Correct answer: A — Explicitly describe the desired tone, voice, and level of creative flourish rather than relying on Claude to infer "engaging" from one adjective.
- A. Correct. Anthropic's guidance is that if you want "above and beyond" behavior, you should explicitly request it rather than relying on the model to infer this from vague prompts; naming the specific tone and style closes the gap between what was asked and what was wanted.
- B. Incorrect. Randomly sampling from multiple generic outputs does not add the missing specificity about desired tone and is not a reliable improvement technique.
- C. Incorrect. Shortening the token budget constrains length, not voice or creative quality, and does not address the underlying vagueness of "engaging."
- D. Incorrect. Politeness markers or punctuation do not change how specifically the task is defined and are not a documented technique for shaping tone.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A short or cut-off answer means Claude's context window is full, so the fix is to shorten the prompt.Why is that wrong?
The context window applies to your prompts, not to Claude's output. A brief answer calls for retooling the prompt, for example by stating the depth needed or by passing back the previous prompt and response.
Covered in When the response is too brief
2.Claude will work out that you want a thorough, above-and-beyond answer without being told.Why is that wrong?
Claude responds to explicit instructions. If you want more than the basics, you have to request it.
Covered in Generic answers come from generic prompts
3.When a long thread starts mixing up details, the best fix is to keep adding corrections in the same conversation.Why is that wrong?
Blending and contradiction in a long thread is drift, where working memory and steerability collide. The targeted fix is to re-supply the context, not to stack more corrections on a thread that has lost focus.
Covered in When the conversation itself is the problem
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://academy.claude.com/courses/ai-capabilities-and-limitations/when-properties-collideOfficial docs
“Most real-world AI failures aren't one property acting up. They're two properties meeting at the same time.”
↩︎ Name the failure before you change the prompt“Next Token Prediction + Knowledge (hallucinated specifics)”
↩︎ Name the failure before you change the prompt“Naming the properties at play points you straight to the fix: verify specifics, re-supply context, offload to code execution, or invite pushback.”
↩︎ When the conversation itself is the problem“Naming the properties at play points you straight to the fix: verify specifics, re-supply context, offload to code execution, or invite pushback.”
↩︎ Key concept“Working Memory + Steerability (long-conversation drift)”
↩︎ Exam trap 3 - 2.https://support.claude.com/en/articles/7996857-my-prompt-isn-t-giving-me-a-helpful-answerOfficial docs
“Pretend you are giving these instructions to someone with no background knowledge about what you are asking.”
↩︎ Generic answers come from generic prompts“Break down complex requests into substeps.”
↩︎ When the conversation itself is the problem - 3.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practicesOfficial docs
“Think of Claude as a brilliant but new employee who lacks context on your norms and workflows.”
↩︎ Generic answers come from generic prompts“Golden rule: Show your prompt to a colleague with minimal context on the task and ask them to follow it.”
↩︎ Generic answers come from generic prompts“Providing context or motivation behind your instructions, such as explaining to Claude why such behavior is important, can help Claude better understand your goals”
↩︎ Generic answers come from generic prompts“Be specific about the desired output format and constraints.”
↩︎ When the response is too brief“If you want "above and beyond" behavior, explicitly request it rather than relying on the model to infer this from vague prompts.”
↩︎ Exam trap 2 - 4.https://support.claude.com/en/articles/8114518-claude-s-response-to-my-prompt-is-too-briefOfficial docs
“If its responses are too brief or only partially complete, consider retooling your prompts.”
↩︎ When the response is too brief“We recommend giving Claude the previous prompt and response when writing your next prompt to pick up where it left off.”
↩︎ When the response is too brief“context window applies to the prompts you provide but not the output it generates.”
↩︎ Exam trap 1