What you will be able to do
- Explain why a session that generated code is a weak reviewer of that same code
- Set up a writer/reviewer split where the reviewer never sees the generator's reasoning
- Tell a truly independent review instance apart from one that only looks independent, such as a resumed session
- Configure a read-only reviewer subagent with a scoped, criteria-driven review brief
Key concept
Independent review instance — A reviewer that starts from a clean context and sees only the artifact and a review brief. It never sees the reasoning, rejected alternatives or corrections that produced the artifact. Its value comes from what it does not know about how the code was written.
1.Why the author is a poor reviewer
Suppose Claude has just written a feature and you ask the same session to review it. The reviewer is not reading the code fresh. The approach it chose, the alternatives it dropped and every correction you gave it along the way are all still in its context. The exam guide states the consequence directly: a model that keeps the reasoning from generation is less likely to question its own decisions in that session. The assumptions that caused a subtle bug are the same assumptions it now reads the code through.
Anthropic's multi-agent guidance calls the general problem context pollution: when an agent's context accumulates information from one subtask that is irrelevant to subsequent subtasks, context pollution occurs, diluting attention and reducing response quality. Claude Code's best-practices guide describes a version you may have seen in long sessions. After Claude has been corrected again and again, "Context is polluted with failed approaches." The fix it gives is not a firmer instruction inside that context. You clear the session and start again with a better prompt.
Reviewing is a subtask with exactly this shape. The generator's reasoning is large, and a reviewer judging whether the result is correct does not need it. So the fix is architectural: give the review to a context that never held that reasoning.
Nothing about its context. It still holds the whole generation history, including the reasoning that produced any mistakes. You changed the instruction but not the vantage point, and the vantage point is the problem.
2.A second instance without the generator's context
Claude Code's best-practices guide shows the basic pattern as two sessions running side by side. Session A writes. Session B never saw how the code was built. It gets a pointer to the file and a specific brief about what to look for. Its findings then go back to the writer to fix.
| Step | Session A (Writer) | Session B (Reviewer) |
|---|---|---|
| 1 | Implement a rate limiter for our API endpoints | — |
| 2 | — | Review the rate limiter implementation in @src/middleware/rateLimiter.ts. Look for edge cases, race conditions, and consistency with our existing middleware patterns. |
| 3 | Here's the review feedback: [Session B output]. Address these issues. | — |
Look at what B receives: the artifact and the criteria (edge cases, race conditions, consistency with existing patterns). It does not get A's chain of reasoning. This is also the evaluator-optimizer workflow from Anthropic's agent patterns, where one LLM call generates a response while another provides evaluation and feedback in a loop. That pattern fits best when there are clear evaluation criteria, which is why B's brief names what to look for instead of saying "check it".
A team built a pipeline where the same Claude instance that generates code is then asked, within the same conversation, to 'review your own work for bugs before finishing.' QA later finds subtle issues the model missed during that self-review step. Which architectural change is most effective at catching those issues going forward?
Correct answer: C — Spawn a second, independent Claude instance with no access to the generation session's history to review the code fresh
- A. Incorrect. More reasoning tokens inside the same session still operate on top of the reasoning that already justified the original decisions, so the underlying self-review limitation persists.
- B. Incorrect. Stricter wording in the system prompt doesn't remove the session's retained reasoning context, which is the actual cause of missed self-review findings.
- C. Correct. A fresh instance with no prior reasoning context is not anchored to the generator's original decisions, making it more likely to catch subtle issues than self-review.
- D. Incorrect. Repeating the review within the same session still carries forward the same reasoning that produced the code, so the blind spot remains.
3.Packaging the reviewer as a subagent
You don't need to open a second terminal to get independence. Claude Code subagents run in their own context, so a reviewer defined as a subagent starts clean every time it is invoked. The documentation's example reviewer is a short Markdown file with frontmatter:
---
name: code-reviewer
description: Reviews code for quality and best practices
tools: Read, Glob, Grep
model: sonnet
---
You are a code reviewer. When invoked, analyze the code and provide
specific, actionable feedback on quality, security, and best practices.Two design choices are worth noticing. The tools are limited to Read, Glob and Grep, so the reviewer can inspect code but cannot edit it. And the definition contains no generation history at all. When you invoke it, give it a spec to check against, not a story about how the code was written:
Use a subagent to review the rate limiter diff against PLAN.md. Check that
every requirement is implemented, the listed edge cases have tests, and
nothing outside the task's scope changed. Report gaps, not style preferences.The closing line, "Report gaps, not style preferences", matters as much as the isolation. A fresh reviewer with no criteria will fill its report with whatever it happens to notice. A fresh reviewer given a plan and a definition of a gap gives you findings you can act on.
Independence has a cost, so choose where to spend it. Anthropic reports that multi-agent implementations typically use 3-10x more tokens than single-agent approaches for equivalent tasks. It also warns against teams who build separate agents for planning, execution, review and iteration and then lose context at each handoff. A dedicated reviewer earns its keep where a missed bug is expensive, not as a reflex on every change.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Resuming the session or subagent that generated the code and asking it to review its diff counts as an independent review, because review is a separate step.Why is that wrong?
What makes a review independent is its context, not the order of steps. A resumed session still holds the generator's reasoning. An independent instance starts from a clean context and sees only the artifact and the brief.
Covered in A second instance without the generator's context
2.Extra reviewer agents are free insurance, so every workflow should add separate planning, review and iteration agents.Why is that wrong?
Each extra agent duplicates context and adds coordination overhead. Anthropic measures that overhead at several times the tokens of a single agent, so a separate reviewer should be used where its independence is actually worth that cost.
Covered in Packaging the reviewer as a subagent
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://code.claude.com/docs/en/best-practicesOfficial docs
“Context is polluted with failed approaches.”
↩︎ Why the author is a poor reviewer“Look for edge cases, race conditions, and consistency with our existing middleware patterns.”
↩︎ A second instance without the generator's context“Here's the review feedback: [Session B output]. Address these issues.”
↩︎ A second instance without the generator's context“Report gaps, not style preferences.”
↩︎ Packaging the reviewer as a subagent - 2.
“When an agent's context accumulates information from one subtask that is irrelevant to subsequent subtasks, context pollution occurs.”
↩︎ Why the author is a poor reviewer“diluting attention and reducing response quality.”
↩︎ Why the author is a poor reviewer“multi-agent implementations typically use 3-10x more tokens than single-agent approaches for equivalent tasks”
↩︎ Packaging the reviewer as a subagent“Subagents provide isolation, with each operating in its own clean context focused on its specific task.”
↩︎ Key concept“Subagents provide isolation, with each operating in its own clean context focused on its specific task.”
↩︎ Exam trap 1“multi-agent implementations typically use 3-10x more tokens than single-agent approaches for equivalent tasks”
↩︎ Exam trap 2 - 3.
“one LLM call generates a response while another provides evaluation and feedback in a loop.”
↩︎ A second instance without the generator's context - 4.https://code.claude.com/docs/en/sub-agentsOfficial docs
“Reviews code for quality and best practices”
↩︎ Packaging the reviewer as a subagent