CertSafari
    CLAUDE-CERTIFIED-ARCHITECT-FOUNDATIONS-CCAR-F · Lessons

    Domain 4 · Lesson 20/30

    Few-Shot Prompting for Consistent Output and Judgment

    Apply few-shot prompting to improve output consistency and quality

    15 min read
    3.33% of exam
    10 sources
    Published 28 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Explain why a few worked examples produce consistent, actionable output when detailed instructions alone do not
    • Design 2-4 targeted examples for ambiguous cases, each showing why one action was chosen over a plausible alternative
    • Describe the role of few-shot examples in demonstrating ambiguous-case handling, such as tool selection for ambiguous requests and branch-level test coverage gaps
    • Explain how few-shot examples with reasoning enable the model to generalize judgment to novel patterns rather than matching only pre-specified cases
    • Use examples that pin down an output format such as location, issue, severity and suggested fix
    • Cut false positives by giving examples that separate acceptable code patterns from genuine issues
    • Reduce fabricated values and empty required fields in extraction by showing examples that span informal wording and varied document structures

    Key concept

    Few-shot examples as specification — A few worked input-to-output pairs in the prompt show the model the format and the judgment you want. Instructions describe what you want in the abstract. Examples show it, and that closes gaps that more instructions keep leaving open.

    1.When instructions stop working, show the output

    Picture a code-review prompt with a careful paragraph of instructions: report each problem, say where it is, rate how serious it is, suggest a fix. Run it across a hundred files and the output drifts. Some findings give a line number and some only name a function. Severity shows up as 'high', 'P1', 'critical' or not at all. Some fixes are concrete diffs, and others say 'consider refactoring'. Each instruction was followed in some reading, but no two responses read the same way, so a downstream tool or a busy reviewer can't use them reliably.

    This is the situation this subdomain is about: detailed instructions alone give inconsistent results. The usual reaction is to add more rules, but that seldom helps, because the missing information isn't another rule. It is a concrete picture of the finished output. Anthropic's consistency guidance says it directly: showing an example of the output you want works better than describing it in the abstract. The ticket-routing guide calls providing examples the most effective way to improve performance.

    Anthropic's consistency guide constrains output with an example format rather than a description. The model then fills this exact shape for every competitor it analyses.xml
    <competitor>
      <name>Rival Inc</name>
      <overview>A 50-word summary.</overview>
      <swot>
        <strengths>- Bullet points</strengths>
        <weaknesses>- Bullet points</weaknesses>
        <opportunities>- Bullet points</opportunities>
        <threats>- Bullet points</threats>
      </swot>
      <strategy>A 30-word strategic response.</strategy>
    </competitor>

    Applied to the code reviewer, this means writing two or three complete findings in the exact shape you want, with four fields every time: location (file and line), issue (one sentence on what is wrong), severity (one value from a fixed set), and suggested fix (a concrete change, not advice). Once the model has seen those examples, it copies their granularity, vocabulary and field order much more reliably than it follows a sentence that lists the same fields.

    There is a flip side. Current Claude models pay close attention to the details of the examples you give. If every example happens to report severity 'high', or every location points into a test file, the model may pick up that accident as a rule. So vary your examples in every way you don't care about, and keep them identical in the format you do care about.

    Sources123

    2.Examples for ambiguous cases: show the reasoning, not just the answer

    Format isn't the only problem examples solve. The harder problem is judgment on inputs where two actions both look reasonable. This is the role of few-shot examples in demonstrating ambiguous-case handling: they show the model what to do on exactly the inputs where the instructions run out.

    The first classic case is tool selection for ambiguous requests. Take an agent with two search tools, one over documentation and one over the codebase. A request like 'where is the retry limit configured' could go to either. With the default automatic tool choice, Claude calls a tool when the request matches that tool's described capability. When an ambiguous request matches both descriptions equally well, the model has no stable basis for choosing, so similar sessions go different ways. Anthropic's tool-choice cookbook notes that Claude can be over-eager to call tools and that a detailed prompt helps it decide when to call a tool and when not to. Few-shot examples are the most concrete form of that detail.

    The fix is a small set of targeted examples, typically 2-4, aimed squarely at that ambiguity. Each example shows the request, the action chosen, and the reason it beat the plausible alternative. For example: 'Request asks where a value is *configured*. Chose code search, because configuration lives in source and config files, and documentation may describe a default that has since changed. Documentation search would be right if the request asked what the setting *means*.' The contrast with the rejected option is what makes the example useful.

    The second classic case is branch-level test coverage gaps. A reviewer that looks for missing tests will see a function with a test file next to it and call it covered. An example that walks through the branches ('the happy path is tested; the branch that runs when the upstream call times out has no test, so this is a gap even though the function is exercised') demonstrates the ambiguous-case handling you want: the model learns to reason at the branch level instead of the file level, and it learns that 'a test exists' and 'every branch is tested' are different claims.

    The reasoning matters because it is how few-shot examples enable the model to generalize judgment to novel patterns rather than matching only pre-specified cases. An example that shows only an answer teaches the model to match that input, so every variant would have to be enumerated. An example that shows a rationale teaches a rule the model can apply to inputs you never wrote down: a new phrasing of the search request, a coverage gap in a branch shape your examples never mentioned. Anthropic's ticket-routing guide makes this point: including a classification rationale for nuanced cases helps Claude generalize the logic to other tickets, and its general prompting guidance says Claude is smart enough to generalize from an explanation.

    Newer models make this more important. Claude Opus 4.8, for instance, interprets prompts literally and does not silently generalize an instruction from one item to another. Left with bare answer-only examples, it will treat them as a list of pre-specified cases. Given the reasoning, it has a principle it can carry to the novel pattern in front of it.

    A pull-request review agent is meant to flag branches (conditional paths) that lack test coverage. In practice it inconsistently flags newly introduced branches that are already exercised indirectly by an existing integration test, producing noisy false positives that erode reviewer trust. Detailed instructions about 'coverage' have not resolved the inconsistency. What should the team add to the prompt?

    Sources24567

    3.Separating acceptable patterns from genuine issues

    False positives are a special case of ambiguity: code that matches a 'bad pattern' on the surface and is fine in context. A broad exception handler is a classic example. Inside a leaf function that swallows every error, it hides bugs. At a top-level boundary that logs and re-raises, it is deliberate and correct. A reviewer that flags both cases trains its human readers to ignore it, which costs more than the missed bugs it was meant to catch.

    The tempting fix is an instruction such as 'be conservative' or 'only report high-severity issues'. Anthropic's guidance for Claude Opus 5 warns that the model may follow such an instruction literally and report less. You lose real findings along with the noise, because the instruction lowers the reporting volume without making the model any better at telling the two cases apart.

    The better fix is a pair of contrasting examples. Show the same pattern twice: once marked as a genuine issue with the reason ('the handler swallows the exception and returns a default, so callers can't tell success from failure'), and once marked as acceptable with the reason ('this is the process entry point; the handler logs and re-raises, so nothing is hidden'). The model learns the property that separates the two cases, whether the error is hidden or passed on. It can then apply that to patterns your examples never mentioned, such as a retry wrapper or a background worker. False positives drop while real issues are still caught.

    A team's code-review assistant prompt already spells out an exhaustive, itemized rubric for flagging issues, but reviewers still receive inconsistently formatted findings across runs: some list severity before location, others omit the suggested fix entirely. The team wants the most effective fix for this output-consistency problem. What should they do?

    Sources82

    4.Extraction: fewer invented values, fewer empty fields

    Extraction pipelines fail in two opposite ways, and both have the same cause: the model hasn't seen what a correct answer looks like for this kind of input.

    Hallucinated precision. The source says 'a couple tablets' or 'about half a cup', and the schema has a numeric dosage field, so the model makes up a number. An example that shows the right handling fixes this better than a warning does. One example keeps the informal quantity as written and flags it as approximate. Another leaves the field empty with a note when the text states no amount at all. Anthropic's hallucination guidance recommends explicitly allowing the model to express uncertainty, and an example is the most concrete way to show what that looks like in your output format.

    Empty required fields. The information is in the document, but somewhere the model doesn't expect. A research-paper extractor finds citations when they are gathered in a bibliography and returns an empty list when they are inline, as in '(Smith, 2021)'. It finds the sample size under a 'Methodology' heading and returns null when the same detail sits in a sentence of the results section. The fix is examples drawn from the structures you actually receive: one inline-citation document and one bibliography document, one with a dedicated methodology section and one where the method is mentioned in passing. Each example shows the field filled correctly despite the different layout.

    Few-shot design by failure mode: what instructions alone tend to produce, and what the examples need to show instead
    Failure modeTypical symptom with instructions onlyWhat the examples should show
    Inconsistent formatFields present but worded, ordered and scaled differently each timeTwo or three complete outputs in the exact target shape
    Ambiguous action choiceSimilar requests handled differently across sessionsThe chosen action plus why it beat the plausible alternative
    False positivesEvery instance of a surface pattern flaggedA contrasting pair: genuine issue vs acceptable use, each with its reason
    Invented valuesPrecise numbers produced from vague source textInformal wording kept as written and flagged, or the field left empty with a note
    Empty required fieldsField is null when the data sits in an unexpected placeCorrect extraction from each document structure you actually receive

    Structured outputs don't remove the need for this. Constrained decoding guarantees that the response matches your schema: valid JSON, correct types, required keys present. A schema can't tell the model that 'about half a cup' shouldn't become 120, or that inline citations count as citations. The schema controls the shape of the output. The examples control whether the values are right.

    A team built a classification prompt with twenty exact input-output pairs, one for every edge case they had personally encountered in their historical data. The prompt performs well on those twenty inputs but degrades noticeably whenever a customer submits a new input that is similar to, but not identical to, one of the twenty. What change would best help the model generalize its judgment to these novel-but-similar inputs?

    Sources910

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.If output is inconsistent despite detailed instructions, the fix is to make the instructions longer or more emphatic.Why is that wrong?

      The missing piece is a concrete picture of the output, not another rule. Examples of the desired output work better than restated abstract instructions.

      Covered in When instructions stop working, show the output

    2. 2.Few-shot examples only teach the model to handle the exact cases you listed, so every variant has to be enumerated.Why is that wrong?

      Examples that include the reasoning behind each choice teach a rule the model can apply to new inputs. Examples that show only answers invite surface matching.

      Covered in Examples for ambiguous cases: show the reasoning, not just the answer

    3. 3.When an agent picks the wrong tool on ambiguous requests, the fix is to force a tool with tool_choice or make the tool descriptions longer.Why is that wrong?

      Forcing a tool removes the choice rather than teaching it, and longer descriptions still leave requests that match both tools equally. A few targeted examples with the reason one tool beat the other give the model a basis for deciding on requests you never listed.

      Covered in Examples for ambiguous cases: show the reasoning, not just the answer

    4. 4.Telling a code reviewer to 'be conservative' or 'only report high-severity issues' is the right way to cut false positives.Why is that wrong?

      The model may take that literally and report less overall, dropping real findings too. Contrasting examples of acceptable and genuine cases teach the difference instead of just lowering the volume.

      Covered in Separating acceptable patterns from genuine issues

    5. 5.Turning on structured outputs with a strict schema stops the model from fabricating values or leaving fields empty.Why is that wrong?

      Constrained decoding guarantees the output matches the schema's shape and types. Whether a value is faithful to the source, or found at all in an unusual document layout, depends on the prompt, which is where examples come in.

      Covered in Extraction: fewer invented values, fewer empty fields

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Provide examples of your desired output. This is more effective than abstract instructions.”
      ↩︎ When instructions stop working, show the output
      “Provide examples of your desired output. This is more effective than abstract instructions.”
      ↩︎ Key concept
      “Provide examples of your desired output. This is more effective than abstract instructions.”
      ↩︎ Exam trap 1
    2. 2.
      “Despite providing examples being the most effective way to improve performance”
      ↩︎ When instructions stop working, show the output
      “include a classification rationale for particularly nuanced ticket intents, so that Claude can better generalize the logic to other tickets”
      ↩︎ Examples for ambiguous cases: show the reasoning, not just the answer
      “consider providing explicit instructions or examples in the prompt of how Claude should handle the edge case”
      ↩︎ Examples for ambiguous cases: show the reasoning, not just the answer
      “consider providing explicit instructions or examples in the prompt of how Claude should handle the edge case”
      ↩︎ Separating acceptable patterns from genuine issues
      “include a classification rationale for particularly nuanced ticket intents, so that Claude can better generalize the logic to other tickets”
      ↩︎ Exam trap 2
    3. 3.
      “Claude 4.x and similar advanced models pay very close attention to details in examples.”
      ↩︎ When instructions stop working, show the output
    4. 4.
      “It calls a tool when the request maps to that tool's described capability and the answer isn't already in context.”
      ↩︎ Examples for ambiguous cases: show the reasoning, not just the answer
    5. 5.
      “Often, Claude can be over-eager to call tools. Writing a detailed prompt helps Claude determine when to call a tool and when not to.”
      ↩︎ Examples for ambiguous cases: show the reasoning, not just the answer
      “Writing a detailed prompt helps Claude determine when to call a tool and when not to.”
      ↩︎ Exam trap 3
    6. 6.
      “It does not silently generalize an instruction from one item to another, and it does not infer requests you didn't make.”
      ↩︎ Examples for ambiguous cases: show the reasoning, not just the answer
    7. 8.
      “the model may follow that instruction literally and report less”
      ↩︎ Separating acceptable patterns from genuine issues
      “the model may follow that instruction literally and report less”
      ↩︎ Exam trap 4
    8. 9.
      “Explicitly give Claude permission to admit uncertainty. This simple technique can drastically reduce false information.”
      ↩︎ Extraction: fewer invented values, fewer empty fields
    9. 10.
      “Structured outputs guarantee schema-compliant responses through constrained decoding”
      ↩︎ Extraction: fewer invented values, fewer empty fields
      “Structured outputs guarantee schema-compliant responses through constrained decoding”
      ↩︎ Exam trap 5