CertSafari
    CLAUDE-CERTIFIED-ARCHITECT-FOUNDATIONS-CCAR-F · Lessons

    Domain 4 · Lesson 22/30

    Schema Errors vs Semantic Errors in Claude Extraction

    Implement validation, retry, and feedback loops for extraction quality

    12 min read
    3.33% of exam
    6 sources
    Published 29 Sep 2026
    Docs as of 24 Sep 2026

    What you will be able to do

    • Tell apart the extraction errors that tool use and strict schemas prevent from the ones they cannot prevent
    • Explain why a schema-valid extraction can still be wrong, and what has to check it
    • Design paired fields, such as a calculated total next to a stated total, that make an extraction expose its own inconsistencies
    • Build a retry that feeds the specific validation errors back to the model, and recognise when a retry cannot succeed because the information is absent from the source
    • Record which code construct triggered each finding, so that dismissed findings can be analysed for false-positive patterns

    Key concept

    Schema-valid is not the same as correct — Constrained decoding guarantees that an extraction has the right shape: it parses, required fields are present and types match. It does not guarantee the values are right, so whether values reconcile and sit in the right fields has to be checked by your own validation.

    1.What tool use and strict schemas remove

    An extraction pipeline can fail in two very different ways, and each needs a different fix. The first kind is structural: the output does not parse, a required key is missing, or a value has the wrong type. Anthropic's structured outputs documentation lists these as the failures you can hit without schema enforcement, even with careful prompting: parsing errors from invalid JSON, missing required fields, inconsistent data types, and schema violations that need error handling and retries.

    Two features remove them. JSON outputs (output_config.format) make Claude's text response follow a JSON schema. Strict tool use (strict: true on a tool definition) makes the tool name and its input follow the tool's input_schema. Both use constrained decoding: the model's token sampling is limited to output the schema allows. So the guarantee is not a higher success rate. It is a property of every response. The documentation's booking example makes this concrete. With passengers typed as an integer, strict mode always returns passengers: 2, never "two" or "2".

    This matters for retry design. A retry loop built to catch malformed JSON, missing required keys or wrong types is solving a problem the platform already solves. The documentation says so directly: with structured outputs, schema violations need no retries.

    Sources12

    2.What a valid schema still lets through

    A schema describes shape: which fields exist, what types they have, and which are required. It cannot describe most of what makes an extraction correct. Take the invoice model from the structured outputs documentation:

    An invoice extraction model: the schema fixes the types, not whether the values agree with each otherpython
    class Invoice(BaseModel):
        invoice_number: str
        date: str
        total_amount: float
        line_items: list[dict]
        customer_name: str

    Constrained decoding guarantees that total_amount is a float and line_items is a list. It says nothing about whether the line items add up to the total, or whether a value that exists in the document ended up in the right field. These are semantic errors: the output is schema-valid and still wrong. The exam guide gives two examples: values that don't sum, and a value placed in the wrong field. A meter reading stored under an address object is schema-valid whenever both objects accept a string.

    The SDK itself treats some constraints as checks that run after generation. If a Pydantic field has minimum: 100, that constraint is not in the schema sent to the model. The SDK sends a plain integer, adds "Must be at least 100" to the field description, and then validates the response against the original constraint. Semantic validation in general works the same way: your code checks what the schema could not enforce. For Pydantic models, client.messages.parse() runs this validation and returns parsed_output.

    Where each kind of extraction error gets caught
    ErrorExampleWhat catches it
    Invalid JSON syntaxResponse fails to parseoutput_config.format or strict: true (constrained decoding)
    Wrong type or missing required fieldpassengers: "two" instead of 2strict: true
    Constraint not in the sent schemaminimum: 100 on a Pydantic fieldSDK validation against the original constraint
    Values that don't reconcileline_items not summing to total_amountYour own validation code

    Sources1

    3.Designing the schema to expose its own mistakes

    Once you accept that some errors are semantic, the next step is to make the schema carry the evidence your validator needs. The exam guide describes two such designs. The first extracts a calculated_total alongside the stated_total. The model reports the total printed on the document and, separately, the sum it computes from the line items. Your code compares the two, and a mismatch raises a flag. The second is a conflict_detected boolean. It lets the model report that the source data is inconsistent instead of silently picking one version.

    The documentation supplied for this lesson does not describe these specific fields, so their names and purpose come from the exam guide. The documentation does support the principle behind them. The Agent SDK guidance says to match the schema to the task. Anthropic's Skills best practices describe the loop that these fields feed: run a validator, fix errors, repeat. The guidance says this pattern greatly improves output quality.

    The two fields answer different questions. If calculated_total disagrees with stated_total because the model misread one line, the error is in the extraction, and it can be corrected. If conflict_detected is true because the document's printed total really does disagree with its own line items, then the extraction is accurate and the document is the problem. Re-extracting will not change the document, so that case goes to review, not to a retry.

    Sources34

    4.Retrying with the validation error in the prompt

    When your validator flags a semantic error that the extraction can fix, such as a calculated_total that disagrees with stated_total, the retry should not be a blind repeat of the first request. A repeat with the same prompt tends to reproduce the same mistake. Instead, the follow-up request carries three things: the original document, the failed extraction, and the specific validation errors your code found. The model then has something concrete to correct against. The exam guide calls this retry-with-error-feedback.

    Anthropic's tool-use documentation shows the same mechanism at the API level. When a tool call is invalid, you continue the conversation with a tool_result that states the error and sets is_error to true. Claude reads the error and tries the tool again with the missing information filled in. The documentation notes that Claude will retry two to three times with corrections before giving up. A semantic validation failure can be fed back in exactly this form: the error text names the field and what was wrong with it.

    A tool_result that feeds a specific error back so the next attempt can correct itjson
    {
      "role": "user",
      "content": [
        {
          "type": "tool_result",
          "tool_use_id": "toolu_01A09q90qw90lq917835lq9",
          "content": "Error: Missing required 'location' parameter",
          "is_error": true
        }
      ]
    }

    The specificity of the error is what makes the loop work. The Skills best practices describe the same discipline for validation scripts: review the error message carefully, fix the issues, run validation again, and only proceed when validation passes. Anthropic's evaluator-optimizer pattern names the precondition: the loop is a good fit when responses can be demonstrably improved once feedback is provided. "Validation failed" gives the model nothing to improve. "line_items sum to 480.00 but stated_total is 500.00" does.

    Retries have a limit, and it is not the retry count. A retry can only fix an error whose correct answer is in the material the model was given. Format mismatches, a value placed under the wrong parent object, a total that was misread: the document contains the right answer, so a retry with the error attached will usually succeed. But if the required information is simply absent from the source document, no number of retries will produce it. A purchase order number that exists only in a separate email the model was never shown cannot be extracted from the invoice, however precise the error feedback. The Agent SDK guidance points at the fix: if the task might not have all the information your schema requires, make those fields optional. The SDK's own structured-output loop ends with an error_max_structured_output_retries result when repeated attempts fail, which is the signal that you are retrying something a retry cannot fix.

    Deciding whether a retry will help
    Validation failureWill a retry with error feedback fix it?Why
    Wrong parameter types or missing required fields, no strict modeUsually, but strict: true removes the needStructural error; the schema can express the constraint
    A value misread from the document, so totals do not reconcileYesThe correct value is in the document; the error tells the model where to look
    A value placed under the wrong parent objectYesStructural placement error with the right answer present in the source
    A required field whose information is not in the documentNoMake the field optional or flag it; retries cannot invent missing information

    Sources5463

    5.Feedback loops across runs: recording what triggered each finding

    The retry loop corrects one extraction. A second kind of feedback loop improves the extractor itself over many runs, and it needs its own field in the schema. Consider a code-review pipeline in which Claude emits structured findings and developers accept or dismiss each one. A dismissed finding is a likely false positive, but on its own it tells you little. The exam guide's design is to add a detected_pattern field to every finding that names the code construct which triggered it: a bare except, a string-built SQL query, a nested ternary. When developers dismiss findings, you can group the dismissals by detected_pattern and see which constructs keep producing findings that nobody wants.

    As with the self-correction fields, the documentation supplied for this lesson does not name detected_pattern; the field and its purpose come from the exam guide. The documentation does describe the practice it enables. The Skills best practices review loop asks the reviewer to note each issue with a specific section reference rather than a general complaint, because a specific record is what makes the next revision targeted. The evaluator-optimizer pattern depends on clear evaluation criteria. A dismissal report grouped by pattern is how you find out which of your criteria are unclear, and then tighten the prompt for exactly those constructs instead of guessing.

    Without the field, the only feedback available is a dismissal rate, which cannot be acted on. With it, systematic analysis becomes possible: a pattern dismissed nine times out of ten is a rule to rewrite or drop, and a pattern rarely dismissed is one to keep. The field costs the model one short string per finding and turns free-text human judgments into a dataset.

    Sources46

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.With strict tool use or output_config.format turned on, an extraction needs no further validation.Why is that wrong?

      Constrained decoding guarantees shape only. Even the SDK validates some constraints after generation because they are not in the schema sent to the model. Whether values add up, and whether they sit in the right fields, still has to be checked in code.

      Covered in What a valid schema still lets through

    2. 2.Putting a value under the wrong parent object is a schema error, so strict mode prevents it.Why is that wrong?

      Strict mode guarantees that the input follows the input_schema. If the misplaced value is a valid type in the wrong location, the output still conforms, and only semantic validation will catch it.

      Covered in What tool use and strict schemas remove

    3. 3.If an extraction keeps failing validation, raising the retry count will eventually get the right value.Why is that wrong?

      Retries with error feedback fix format and structural errors because the correct answer is in the document. When the required information is absent from the source, no retry can supply it. The fix is to make the field optional or flag it, not to retry more.

      Covered in Retrying with the validation error in the prompt

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Reliable: No retries needed for schema violations”
      ↩︎ What tool use and strict schemas remove
      “Schema violations requiring error handling and retries”
      ↩︎ What tool use and strict schemas remove
      “the SDK updates the description to "Must be at least 100" and validates the response against the original constraint.”
      ↩︎ What a valid schema still lets through
      “The parse() method automatically transforms your Pydantic model, validates the response, and returns a parsed_output attribute.”
      ↩︎ What a valid schema still lets through
      “Structured outputs guarantee schema-compliant responses through constrained decoding”
      ↩︎ Key concept
      “the SDK updates the description to "Must be at least 100" and validates the response against the original constraint.”
      ↩︎ Exam trap 1
    2. 2.
      “Without strict mode, Claude might return incompatible types ("2" instead of 2) or omit required fields, breaking your functions and causing runtime errors.”
      ↩︎ What tool use and strict schemas remove
      “Tool input strictly follows the input_schema”
      ↩︎ Exam trap 2
    3. 3.
      “If the task might not have all the information your schema requires, make those fields optional.”
      ↩︎ Retrying with the validation error in the prompt
      “error_max_structured_output_retries”
      ↩︎ Retrying with the validation error in the prompt
      “If the task might not have all the information your schema requires, make those fields optional.”
      ↩︎ Exam trap 3
    4. 4.
      “Common pattern: Run validator → fix errors → repeat”
      ↩︎ Designing the schema to expose its own mistakes
      “This pattern greatly improves output quality.”
      ↩︎ Designing the schema to expose its own mistakes
      “Review the error message carefully”
      ↩︎ Retrying with the validation error in the prompt
      “Only proceed when validation passes”
      ↩︎ Retrying with the validation error in the prompt
      “Note each issue with specific section reference”
      ↩︎ Feedback loops across runs: recording what triggered each finding
    5. 5.
      “Claude will try to use the tool again with the missing information filled in”
      ↩︎ Retrying with the validation error in the prompt
      “If a tool request is invalid or missing parameters, Claude will retry 2-3 times with corrections before apologizing to the user.”
      ↩︎ Retrying with the validation error in the prompt