What you will be able to do
- Rewrite a vague review instruction as a testable, categorical criterion
- Explain why "be conservative" or "only report high-confidence findings" moves where the model stops reporting without saying which findings you want
- Write report and skip lists that define what a review surfaces without relying on the model's confidence
- Explain how a high false positive category erodes developer trust in the accurate categories, and when to disable it temporarily while its prompt is reworked
- Define severity levels with explicit criteria and concrete code examples so classification stays consistent across runs
Key concept
Explicit categorical criteria — A review prompt says which kinds of issue to report and which to skip, and places the bar with concrete conditions. It does not ask the model to decide for itself what counts as important or confident enough.
1.Why vague instructions produce vague findings
Take two instructions for a review agent. "Check that comments are accurate" leaves the model to decide what "accurate" means. A stale phrase, an imprecise word, a comment that simplifies what the function does: each could be a finding, and nothing in the prompt rules any of them out. "Flag comments only when claimed behavior contradicts actual code behavior" names one condition you can test. Either the comment says the function returns null on failure and the code throws, or it doesn't. The second instruction leaves no room for the reviewer's taste, and taste is where false positives come from.
Anthropic's general prompting guidance starts from the same idea: "Claude responds well to clear, explicit instructions. Being specific about your desired output can help enhance results." It suggests picturing the model as "a brilliant but new employee who lacks context on your norms and workflows." A new hire told to "check that comments are accurate" would ask what you mean by accurate. The model can't ask, so it guesses.
This matters more on current models because they read instructions literally. The Opus 4.8 guide says the model "does not silently generalize an instruction from one item to another, and it does not infer requests you didn't make." So the criterion you write is the criterion the model applies. A vague criterion gets applied loosely, and a precise one gets applied precisely. The same guidance recommends giving the reason behind a rule, because "Claude is smart enough to generalize from the explanation." A criterion that comes with its reason is easier to apply to cases you didn't list.
The exam guide draws the contrast as two ways of tightening a prompt. General instructions like "be conservative" or "only report high-confidence findings" fail to improve precision the way specific categorical criteria do, because a general instruction never says which category of finding is the false positive. Anthropic's model guides list exactly these phrases: a review prompt that says "only report high-severity issues," "be conservative," or "don't nitpick" makes the model report less, but it still chooses for itself what to drop. A categorical criterion names the unwanted type, so precision improves where you wanted it to and nowhere else. The next section walks through what the general instruction actually does to the results.
2.Why "be conservative" doesn't set a precision bar
The obvious fix for noise is to ask the model to hold back: "be conservative", "only report high-severity issues", "don't nitpick". Anthropic's model guides describe what that actually does. The model "may investigate the code just as thoroughly, identify the bugs, and then not report findings it judges to be below your stated bar." The result, in the guide's words: "Precision typically rises, but measured recall can fall even though the model's underlying bug-finding ability has improved."
Notice what the instruction doesn't do. It never says which findings are unwanted. The model still decides what counts as "high severity" or "conservative". You only moved the point where it stops reporting, and you can't see which real bugs fell below that point. A wording nit the model happens to rate highly still gets through, and a real bug it rates low gets dropped. A categorical criterion fixes both problems, because it rules the wording nit out by type and rules the bug in by type.
| Qualitative instruction | What the model may do with it | Concrete alternative from the guidance |
|---|---|---|
| "only report high-severity issues" | Finds the bug, then doesn't report it if it judges it below your stated bar | "report any bugs that could cause incorrect behavior, a test failure, or a misleading result" |
| "be conservative" | Reports fewer findings, especially lower-severity bugs; recall falls | Report everything at the finding stage and filter in a separate verification step |
| "don't nitpick" | Leaves the model to decide what a nit is | "only omit nits like pure style or naming preferences" |
When the pipeline has more than one stage, the guidance goes further and takes confidence filtering out of the finding step altogether. The Opus 5 guide says: "ask it to report everything and filter in a separate pass instead." This is the recommended wording for the finding stage:
Report every issue you find, including ones you are uncertain about or consider low-severity. Do not filter for importance or confidence at this stage - a separate verification step will do that. Your goal here is coverage: it is better to surface a finding that later gets filtered out than to silently drop a real bug. For each finding, include your confidence level and an estimated severity so a downstream filter can rank them.The criterion. "Be conservative" leaves the model to decide what counts as a finding. It may report fewer comments, but nothing tells it to drop wording nits in particular, and it can also drop real contradictions. "Flag only when claimed behavior contradicts actual code behavior" rules out wording issues by type and keeps every real mismatch.
Why does a noisy category matter beyond the wasted comments? The exam guide's point is that the false positive rate sets developer trust for the whole review, not only for the category producing the noise. When one category, say comment accuracy, is mostly wrong, developers learn to skim past review comments, and the accurate categories, such as an unscoped query or PII in a log line, get skipped along with it. High false positive categories undermine confidence in the accurate categories. Anthropic's own Code Review is built around that risk: after the finding agents run, "a verification step checks results against actual code behavior to filter out false positives." The content moderation guide adds the measurement side. Track precision and recall per category, and "Use this data to iteratively refine your moderation prompts, keywords, and assessment criteria." You cannot manage trust in a category whose false positive rate you never measured.
When a category's false positive rate is high and its prompt cannot be fixed today, take the category out of the review instead of letting it drag the others down. In Claude Code Review the switch is REVIEW.md, which the docs describe as the place to "say what your team wants flagged, at what severity, and how findings are reported." Add the noisy category to its "Do not report" list. The review keeps posting the categories developers trust, you rework the criteria for the disabled category against your evals, and you remove the entry once its precision is acceptable. Temporarily disabling a high false-positive category restores developer trust faster than tuning it live, because every wrong comment posted while you iterate costs trust that the accurate findings then have to win back.
Severity labels have the same weakness as "important": without criteria, the model decides what an Important finding is, and two runs classify the same bug differently. The documented REVIEW.md example fixes this by defining the top level with concrete code examples and by stating what falls to Nit:
## What Important means here
Reserve Important for findings that would break behavior, leak data,
or block a rollback: incorrect logic, unscoped database queries, PII
in logs or error messages, and migrations that aren't backward
compatible. Style, naming, and refactoring suggestions are Nit at
most.Claude Security documents the same pattern for its High, Medium and Low levels. Each level gets a condition the reviewer can check plus a typical example, and the docs note that "Severity is assigned per finding based on exploitability in your codebase, not the category itself". A finding is classified by the criteria it meets, not by the bucket its category usually lands in. That is what makes classification consistent: the same code pattern meets the same criteria on every run.
| Severity | Criteria | Typical example |
|---|---|---|
| High | Exploitable by an unauthenticated remote attacker against a default deployment, with no meaningful preconditions | Unauthenticated command injection in a public API endpoint |
| Medium | Exploitable behind authentication, or needs 1–2 realistic preconditions (specific role, known identifier, user interaction) | SQL injection behind auth requiring knowledge of table schema |
| Low | Needs 3+ preconditions, local-only access, or lacks a concrete demonstrated attack path | Timing side-channel requiring network proximity and thousands of requests |
3.Writing report and skip lists
If a single pass has to filter its own output, the guidance tells you to state the bar in concrete terms: "report any bugs that could cause incorrect behavior, a test failure, or a misleading result; only omit nits like pure style or naming preferences." That sentence contains both halves of a good criterion: what always gets reported, defined by its consequence, and what always gets skipped, defined by its type.
Claude Code Review turns this into a file. REVIEW.md holds review-only instructions, and the docs say to "Use it to say what your team wants flagged, at what severity, and how findings are reported." The documented example keeps an explicit skip list and an explicit must-check list side by side:
## Do not report
- Anything CI already enforces: lint, formatting, type errors
- Generated files under `src/gen/` and any `*.lock` file
- Test-only code that intentionally violates production rules
## Always check
- New API routes have an integration test
- Log lines don't include email addresses, user IDs, or request bodies
- Database queries are scoped to the caller's tenantNeither list mentions confidence. Each entry is a category a reviewer can recognise in the code: a lock file, a log line containing a user ID, an unscoped query. Every skip entry also comes with its reason, such as CI already enforcing it, the file being generated, or the violation being intentional in tests. The exam guide adds minor style and local patterns to the usual skip categories, and puts bugs and security issues on the always-report side. When a flagged issue is really a local convention and not a defect, the fix is a criterion that says so, not a higher confidence threshold.
An architect is structuring a long review prompt that defines separate criteria for security, correctness, and style categories, each with its own inclusion rules and severity examples. Which structuring approach best helps Claude apply the right criteria to the right category without cross-contamination?
Correct answer: D — Wrap each category's criteria and examples in its own uniquely named XML tag, such as <security_criteria> and <correctness_criteria>, so the boundaries between categories are unambiguous.
- A. Incorrect. A single continuous paragraph without structural markers is more prone to the model blending or misapplying criteria across categories compared to explicitly tagged sections.
- B. Incorrect. Relying on bullet order alone to imply category membership is fragile and ambiguous; there is no explicit marker tying a given bullet to a specific category.
- C. Incorrect. Duplicating every category's full criteria in every section adds unnecessary length and redundancy without addressing the actual boundary-clarity problem, which structured tags solve directly.
- D. Correct. Wrapping each category's rules and examples in its own descriptively named XML tag gives the model clear, unambiguous boundaries between categories, which is the recommended way to structure prompts that mix multiple sets of instructions.
A team already tried adding "only report issues you are confident about" to a noisy category and saw no improvement in precision. An architect now wants to redesign the category's scope entirely rather than continuing to tune confidence language. Which redesign reflects the correct lesson from the earlier failed attempt?
Correct answer: B — Replace the confidence instruction with a list of the specific issue types that qualify for this category, and explicitly state which related issue types should be skipped.
- A. Incorrect. Message placement is not the reason a vague confidence instruction fails; the lesson from the earlier attempt is that the content of the instruction, not its location, needs to change.
- B. Correct. This replaces the confidence-based filter with categorical criteria defining exactly which issue types are in scope and which are excluded, addressing the root cause of the earlier failure rather than adjusting the phrasing of the same ineffective approach.
- C. Incorrect. Adding emphasis to the same confidence-based instruction does not change its fundamental nature; it still fails to define what makes an issue reportable, so the same failure mode would likely recur.
- D. Incorrect. A numeric threshold still asks the model to introspect on its own confidence rather than apply a defined categorical rule, and confidence self-reports of this kind have already been shown not to move precision in this scenario.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Adding "be conservative" or "only report high-confidence findings" is the reliable way to reduce false positives in a review prompt.Why is that wrong?
The model may follow the instruction literally: it finds the bugs and then doesn't report the ones it judges below a bar you never defined. Recall falls, and nothing specifically targets the unwanted categories. The guidance recommends stating the bar concretely.
Covered in Why "be conservative" doesn't set a precision bar
2.A category with a high false positive rate should stay enabled while its prompt is tuned, because disabling it loses coverage.Why is that wrong?
A noisy category costs trust in every other category, so the accurate findings get ignored too. The design the exam expects is to put the category on the REVIEW.md "Do not report" list, improve its criteria against evals, and re-enable it once precision is acceptable.
Covered in Why "be conservative" doesn't set a precision bar
3.Telling the model not to nitpick is enough to keep style issues out of a review.Why is that wrong?
"Nitpick" is itself a judgement call. A concrete criterion names what gets omitted (pure style or naming preferences) and what always gets reported (anything that could cause incorrect behavior, a test failure, or a misleading result).
Covered in Writing report and skip lists
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practicesOfficial docs
“Claude responds well to clear, explicit instructions. Being specific about your desired output can help enhance results.”
↩︎ Why vague instructions produce vague findings“Think of Claude as a brilliant but new employee who lacks context on your norms and workflows.”
↩︎ Why vague instructions produce vague findings - 2.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-4-8Official docs
“It does not silently generalize an instruction from one item to another, and it does not infer requests you didn't make.”
↩︎ Why vague instructions produce vague findings“When a review prompt says things like "only report high-severity issues," "be conservative," or "don't nitpick," Claude Opus 4.8 may follow that instruction”
↩︎ Why vague instructions produce vague findings“Precision typically rises, but measured recall can fall even though the model's underlying bug-finding ability has improved.”
↩︎ Why "be conservative" doesn't set a precision bar“be concrete about where the bar is rather than using qualitative terms like "important"”
↩︎ Key concept“it may investigate the code just as thoroughly, identify the bugs, and then not report findings it judges to be below your stated bar.”
↩︎ Exam trap 1“report any bugs that could cause incorrect behavior, a test failure, or a misleading result; only omit nits like pure style or naming preferences.”
↩︎ Exam trap 3 - 3.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5Official docs
“ask it to report everything and filter in a separate pass instead.”
↩︎ Why "be conservative" doesn't set a precision bar - 4.
“a verification step checks results against actual code behavior to filter out false positives”
↩︎ Why "be conservative" doesn't set a precision bar - 5.
“Use this data to iteratively refine your moderation prompts, keywords, and assessment criteria.”
↩︎ Why "be conservative" doesn't set a precision bar - 6.https://code.claude.com/docs/en/code-reviewOfficial docs
“Use it to say what your team wants flagged, at what severity, and how findings are reported.”
↩︎ Why "be conservative" doesn't set a precision bar“Reserve Important for findings that would break behavior, leak data,”
↩︎ Why "be conservative" doesn't set a precision bar“Use it to say what your team wants flagged, at what severity, and how findings are reported.”
↩︎ Writing report and skip lists“Use it to say what your team wants flagged, at what severity, and how findings are reported.”
↩︎ Exam trap 2 - 7.
“Severity is assigned per finding based on exploitability in your codebase, not the category itself”
↩︎ Why "be conservative" doesn't set a precision bar - 8.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-sonnet-5Official docs
“be concrete about where the bar is rather than using qualitative terms like "important"”
↩︎ Writing report and skip lists