What you will be able to do
- Split a large multi-file review into focused per-file passes and script them at scale
- Add a separate integration pass for cross-file issues and reconcile findings across passes
- Design a verification pass that filters false positives and routes findings by severity, while knowing what the docs do and don't say about self-reported confidence
- Choose between the managed Code Review service and /code-review ultra for verified, multi-agent review
1.Why one giant review pass underperforms
A clean reviewer context solves one problem. Size creates another. Give a single reviewer a forty-file diff and it has to hold every file's details while judging each one. The exam guide names two failure modes. The first is attention dilution: local issues in file 30 get less scrutiny than those in file 3. The second is contradictory findings, where the same pattern is accepted in one place and flagged in another.
Anthropic's agent-pattern guidance gives the general principle: for complex tasks with many considerations, LLMs generally perform better when each consideration is handled by a separate LLM call, allowing focused attention on each specific aspect. Its name for splitting work this way is sectioning: breaking a task into independent subtasks run in parallel. Claude Code's best-practices guide gives the same advice about scope. Instead of letting one context read hundreds of files, scope investigations narrowly or use subagents so the exploration doesn't consume your main context.
2.Per-file passes, scripted
When the unit of work is one file, give each file its own fresh invocation. The best-practices guide shows the pattern for a large migration. First generate a task list, then loop over it with non-interactive claude -p calls, one per file:
for file in $(cat files.txt); do
claude -p "Migrate $file from Python 2 to Python 3. Return OK or FAIL." \
--allowedTools "Edit,Bash(git commit *)"
doneThe same shape works for review. Swap the prompt for a per-file review brief and drop the edit tools. Three details carry over. Each call starts with its own context, so file 150 gets the same attention as file 1. The prompt asks for a fixed, machine-readable result (Return OK or FAIL.), so the outputs can be combined by a script instead of read by a person. And the guide's rollout advice applies: test on a few files, then run on all of them. That way you tune the brief before paying for 200 calls.
An engineer configures the generation session with a high extended thinking effort and instructs the model to reflect deeply on flaws before submitting its code. Bugs still slip through review. Why is extended thinking insufficient as a substitute for an independent review instance here?
Correct answer: B — Extended thinking still runs in the same session that produced the code, so it retains the reasoning that justified those decisions
- A. Incorrect. The limitation is not about token cost or reasoning depth; extended thinking can genuinely deepen reasoning, it just does so from within a biased vantage point.
- B. Correct. Extended thinking within the same session does not remove the retained reasoning context, which makes the model less likely to question its own prior decisions.
- C. Incorrect. The limitation isn't a budget ceiling; it's that the reflection happens inside the same reasoning context that produced the code.
- D. Incorrect. Extended thinking does not disable tool use, and the core issue is retained reasoning context, not file access.
Sources1
3.The integration pass and reconciling findings
Per-file passes are blind by design. A function that returns a nullable value can look fine in its own file, and so can the caller that dereferences it without a check. The bug only exists across both files. The exam guide's answer is a separate integration pass whose job is cross-file data flow: what one file produces and another consumes. The per-file passes keep local depth. The integration pass restores the view across files.
Anthropic's managed Code Review service shows the general shape. It describes a fleet of specialized agents that examine the code changes in the context of your full codebase, with multiple agents analysing the diff and surrounding code in parallel, each looking for a different class of issue. The sources don't give a specific prompt for an integration pass. What they establish is the principle of separate, focused passes, and that those passes must see the surrounding code, not just the diff.
Several passes produce overlapping and sometimes conflicting reports, so you need a reconciliation step at the end. The managed service's version: findings are deduplicated, ranked by severity, and posted as inline comments on the specific lines where issues were found. Anthropic's parallelization guidance describes a related variant, voting, for security review, where several different prompts review and flag the code if they find a problem. Agreement between independent passes is itself evidence.
Then the integration pass becomes the single giant pass again, with the same attention dilution. Each kind of pass has one job: per-file passes go deep on local issues, and the integration pass looks only at interactions between files.
4.Verification passes and routing findings
Reviewers produce false positives, and a report full of them trains people to ignore it. So the last stage checks the findings themselves. In the managed service, a verification step checks results against actual code behavior to filter out false positives. The ultrareview mode goes further: every reported finding is independently reproduced and verified, so the results focus on real bugs rather than style suggestions.
After verification, routing decides what reaches a person and how urgently. The managed service tags each finding with a severity:
| Marker | Severity | Meaning |
|---|---|---|
| 🔴 | Normal | A bug that should be fixed before merging |
| 🟡 | Nit | A minor issue, worth fixing but not blocking |
| 🟣 | Pre-existing | A bug that exists in the codebase but wasn't introduced by this PR |
Each finding also includes a collapsible extended reasoning section that explains why Claude flagged the issue and how it verified the problem. That gives a human reviewer evidence to check, not just a verdict.
5.Managed options: Code Review and /code-review ultra
You don't always have to build this pipeline yourself. The managed Code Review service runs on GitHub pull requests. It can trigger when a PR opens, on every push, or on manual request, and it posts inline comments. It is in research preview for Team and Enterprise plans and is billed through usage credits. Reviews scale in cost with PR size and complexity, completing in 20 minutes on average.
/code-review ultra is the on-demand version for a branch or PR: a larger fleet of reviewer agents explores the change in parallel, which surfaces issues that a local review can miss, and it verifies each finding before reporting it. It runs remotely: "No local resource use: the review runs entirely in a cloud sandbox, so your terminal stays free for other work while it runs". It also has limits you should know:
| Case | What happens |
|---|---|
| Diff too large | Refused. By default a branch review covers up to 500 changed files and 8,000 changed lines, and the refusal names the limits in effect |
| Nothing to review | Refused when the diff against the base is empty, with a suggested way out such as switching branch or passing a different base |
| First commit | Every file is reviewed after you confirm in the launch dialog. Non-interactive claude ultrareview and claude -p refuse and point to an interactive session |
A refactor touches 60 files. A single reviewer instance given the entire diff at once produces contradictory findings between files. What change to the review architecture best addresses this?
Correct answer: A — Split the review into per-file passes for local issues, plus an integration pass for cross-file consistency
- A. Correct. Splitting into per-file local passes plus a separate cross-file integration pass avoids the attention dilution and contradictory findings that come from reviewing everything in one undifferentiated pass.
- B. Incorrect. Reviewing only the latest commit in isolation loses the cross-file view needed to catch integration issues across the full 60-file change.
- C. Incorrect. Running the same overloaded single pass twice does not fix attention dilution; it just repeats the same structural problem.
- D. Incorrect. A larger context window doesn't resolve attention dilution across unrelated files or eliminate contradictory findings between them.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.If the whole multi-file diff fits in the context window, one review pass is as good as focused per-file passes.Why is that wrong?
Fitting in context is not the same as getting focused attention. Separate calls per consideration let each pass concentrate on one thing, which is why per-file passes beat one combined pass on local issues.
Covered in Per-file passes, scripted
2.Once every file has passed its own per-file review, the change as a whole has been reviewed.Why is that wrong?
Per-file passes can't see data flow between files. A separate pass has to look at the change in the context of the surrounding code to catch issues that only exist across files.
3./code-review ultra can take any diff, so it replaces a scripted per-file review for a very large migration.Why is that wrong?
Ultrareview has default size limits and refuses diffs over them. A migration larger than those limits still needs to be split or scripted.
Covered in Managed options: Code Review and /code-review ultra
Practise it for real
Build a small multi-pass review of your own branch: independent per-file passes, one integration pass, then compare against a managed verified review.
1.Create
.claude/agents/code-reviewer.mdwith the frontmatter from the sub-agents docs (tools: Read, Glob, Grep).Why: A read-only reviewer that starts from a clean context every time it is invoked.
You should see: The subagent is available. If this is the first agent file in a new agents directory, restart the session so it loads.
2.Write the changed file paths to
files.txtand adapt thefor file in $(cat files.txt)loop so eachclaude -pcall reviews one file and ends withReturn OK or FAIL.. Leave out--allowedTools "Edit,...".Why: One fresh context per file avoids attention dilution and gives output a script can aggregate.
You should see: One OK/FAIL line per file, with no edits to your tree.
3.Run the loop on two or three files first, adjust the prompt, then run it on the full list.
Why: The docs advise testing on a few files before scaling.
You should see: A stable prompt before you spend calls on every file.
4.In an interactive session, ask a subagent to review the whole diff against your plan, using the 'Report gaps, not style preferences' brief, with the focus on cross-file interactions.
Why: This is the integration pass the per-file loop cannot do.
You should see: Findings about interactions between files that the per-file loop did not report.
5.Run
/code-review ultraon the same branch and compare its verified findings with your pipeline's.Why: Shows what independent reproduction and a larger reviewer fleet add.
You should see: A cloud-run review with verified findings, or a named refusal if the diff exceeds the size limits.
Stuck? Get a nudge
If the per-file loop and the integration pass report the same issue, treat the agreement as a signal. That is the voting variant of parallelization.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://code.claude.com/docs/en/best-practicesOfficial docs
“Scope investigations narrowly or use subagents so the exploration”
↩︎ Why one giant review pass underperforms“Test on a few files, then run on all of them”
↩︎ Per-file passes, scripted - 2.https://www.anthropic.com/engineering/building-effective-agentsSecondary source
“LLMs generally perform better when each consideration is handled by a separate LLM call, allowing focused attention on each specific aspect.”
↩︎ Why one giant review pass underperforms“Sectioning: Breaking a task into independent subtasks run in parallel.”
↩︎ Why one giant review pass underperforms“Reviewing a piece of code for vulnerabilities, where several different prompts review and flag the code if they find a problem.”
↩︎ The integration pass and reconciling findings“LLMs generally perform better when each consideration is handled by a separate LLM call, allowing focused attention on each specific aspect.”
↩︎ Exam trap 1 - 3.
“A fleet of specialized agents examine the code changes in the context of your full codebase”
↩︎ The integration pass and reconciling findings“Findings are deduplicated, ranked by severity, and posted as inline comments on the specific lines where issues were found.”
↩︎ The integration pass and reconciling findings“then a verification step checks results against actual code behavior to filter out false positives.”
↩︎ Verification passes and routing findings“Findings include a collapsible extended reasoning section you can expand to see why Claude flagged the issue and how it verified the problem.”
↩︎ Verification passes and routing findings“Reviews scale in cost with PR size and complexity, completing in 20 minutes on average.”
↩︎ Managed options: Code Review and /code-review ultra“A fleet of specialized agents examine the code changes in the context of your full codebase”
↩︎ Exam trap 2 - 4.https://code.claude.com/docs/en/ultrareviewOfficial docs
“every reported finding is independently reproduced and verified, so the results focus on real bugs rather than style suggestions”
↩︎ Verification passes and routing findings“No local resource use: the review runs entirely in a cloud sandbox, so your terminal stays free for other work while it runs”
↩︎ Managed options: Code Review and /code-review ultra“a larger fleet of reviewer agents explores the change in parallel, which surfaces issues that a local review can miss”
↩︎ Managed options: Code Review and /code-review ultra“Diff too large: a branch review can include up to 500 changed files and 8,000 changed lines by default.”
↩︎ Exam trap 3