What you will be able to do
- Explain why a high aggregate accuracy figure can hide a failing document type or field, and break results down by segment before reducing human review
- Describe how field-level confidence scores become usable routing thresholds only after calibration against a labeled validation set
- Design a review queue that sends low-confidence or ambiguous or contradictory extractions to humans first, making the best use of limited reviewer capacity
- Use stratified random sampling of high-confidence extractions to keep measuring error rates and catch new error patterns after automation
Key concept
Segment-level trust — Treat an extraction pipeline as trustworthy one document type and one field at a time, never on the strength of one overall number. Human review drops only where accuracy has been measured and holds, confidence scores have been calibrated, and ongoing sampling keeps confirming it.
1.Why 97% overall can still mean a broken pipeline
Structured-extraction systems usually report one headline number, such as "97% of fields extracted correctly." That number is a weighted average, and the weights come from volume. Say typed invoices make up most of your traffic and come out almost perfect. Then a small share of handwritten forms, or one hard field such as a dosage, a policy number or a date inside a long table, can be wrong far more often and barely move the overall figure. The average is accurate arithmetic, but it answers the wrong question. The question behind any decision to cut human review is "where is this pipeline safe to trust?", and an overall number can't answer that.
Anthropic's own evaluation guidance keeps pushing toward breaking results down by segment. The support-ticket routing guide doesn't stop at one routing accuracy. It asks teams to check whether accuracy holds across segments, for example to measure accuracy "across different languages, aiming for no more than a 5–10% drop in accuracy for non-primary languages." Swap "languages" for "document types" or "fields" and you have the discipline this objective tests. A team building a clinical-abstraction system on Claude put it the same way: granular evaluation lets you find the cause of a failure "rather than staring at an aggregate score wondering what went wrong."
An architect asks Claude to output a confidence score from 0 to 1 for each extracted field in a structured JSON response. Before using these scores to decide which fields skip human review, what step should the architect take to make the scores trustworthy for routing decisions?
Correct answer: B — Calibrate review thresholds against a labeled validation set, checking whether fields the model scores highly are actually correct at the rate that score implies before trusting it for routing.
- A. Incorrect. Forcing high scores for successful-looking extractions removes the signal the score is meant to carry; a self-consistently high score doesn't indicate the score maps to real-world correctness rates.
- B. Correct. A raw confidence score is only useful for routing once it has been calibrated: checking against labeled ground truth whether, say, 0.9-scored fields are actually correct roughly 90% of the time, and adjusting the review threshold based on that empirical relationship.
- C. Incorrect. Averaging a field's score with unrelated neighboring fields conflates independent extraction decisions and provides no evidence that the resulting number reflects actual correctness likelihood.
- D. Incorrect. Converting to a categorical label doesn't address whether the underlying score is calibrated; an uncalibrated categorical label is just as unreliable for routing as an uncalibrated numeric one.
2.Validate every segment before you automate it
Validating before automating means setting a bar per segment, then showing that each segment you plan to automate clears it. In practice: take the labeled evaluation set you already have and group it by document type (auto, home and health claims, say). Within each type, group it by field. Compute accuracy for each cell of that grid. Every cell clears the bar or it doesn't, and you only reduce review for the cells that do. If a segment has too few labeled examples to give a stable number, that means you need more labeled data for it. It isn't a pass.
| Criterion in the guide | Stated target | What it teaches for extraction review |
|---|---|---|
| Accuracy across languages | no more than a 5–10% drop in accuracy for non-primary languages | A minority segment gets its own measured tolerance and is not absorbed into the average |
| Accuracy across customer groups | consistent routing accuracy (within 2–3%) across all customer groups | Accuracy must be consistent across every segment, not only high on average |
| Edge cases | at least 80% accuracy on these challenging inputs | Hard inputs get a dedicated test set and a separate number |
| Consistency over time | a consistency rate of 95% or higher | Re-test periodically with standardized inputs, not once at launch |
The edge-case row matters most here. The guide says to "create a test set of edge cases and measure the routing accuracy" as its own number. Handwritten forms, poor scans and unusual layouts are the edge cases of document extraction. If they are left mixed into the main set, their failures disappear into the average again.
A logistics company's bill-of-lading extraction pipeline shows 94% field accuracy in aggregate. A new architect discovers that the 'weight' field is correct only 70% of the time specifically when the source document's units are ambiguous (for example, a number with no unit label present). All other conditions for the weight field exceed 95%. What review policy should be applied to the weight field going forward?
Correct answer: A — Route the weight field to human review whenever the source document does not clearly specify units, while allowing high-confidence, unambiguous cases to bypass review.
- A. Correct. The failure is specifically tied to source ambiguity (missing units), so the appropriate policy routes exactly those ambiguous cases to human review while letting unambiguous, well-performing cases continue with reduced review, matching review effort to the identified risk.
- B. Incorrect. Citing the blended aggregate ignores the documented sub-condition where accuracy drops to 70%, which is well below an acceptable bar for unreviewed automation; leaving the policy unchanged would let that specific failure mode continue unflagged.
- C. Incorrect. Removing the field from automation entirely discards the fact that unambiguous cases already exceed 95% accuracy, which is an overcorrection that wastes reviewer capacity on cases that don't need it.
- D. Incorrect. Raising temperature increases output variability rather than addressing the root cause, which is that the source document itself lacks the information needed to resolve the unit ambiguity; this would not improve accuracy or reviewer usefulness.
Sources1
3.Field-level confidence, calibrated against labeled data
Once you know which segments are weak, you need a way to point reviewers at the specific extractions that are likely wrong. The standard approach has the model return a confidence score for each field, not one score per document. A document-level score can't tell a reviewer that the invoice total is solid but the due date is a guess. Anthropic's prompting guidance recommends the same shape for other structured outputs: ask for a confidence level on every item so a later stage can act on it. The exact wording is "include your confidence level and an estimated severity so a downstream filter can rank them."
The catch is that a model's self-reported confidence is only a raw signal. A 0.9 from the model doesn't mean the field is right nine times in ten. To make it meaningful, you calibrate it. Run the pipeline over a labeled validation set, where every field has a golden answer to check against. Group the extractions by stated confidence, and measure how often each group is actually correct, broken down by field and document type. Then set each routing threshold at the confidence level where the measured error rate for that field is low enough to accept. You set the number from measured data. You don't pick it because it sounds reasonable.
The threshold was never calibrated. Nobody checked the model's stated confidence against labeled ground truth, so high scores don't reliably mean correct fields, at least for some fields or document types. The fix is to measure accuracy at each confidence level on a labeled validation set, per field and segment, and reset thresholds from that data. Raising one global threshold doesn't fix it. Once reviewers stop trusting the score, the queue has lost its purpose.
A team building a resume-parsing pipeline wants Claude to output a confidence score for each extracted field so reviewers can prioritize their limited time. During prompt design, which approach best supports later calibration of these scores against a labeled validation set?
Correct answer: B — Instruct the model to output a per-field confidence score alongside each extracted value, using a consistent numeric scale across all fields and documents.
- A. Incorrect. A single document-level score cannot be calibrated against or used to route individual fields; a resume with one bad field and nine good ones would get one blended score, defeating the purpose of field-level review routing.
- B. Correct. Field-level scores on a consistent scale allow the team to compare each field's stated confidence against its actual observed correctness rate in a labeled validation set, which is exactly what's needed to calibrate thresholds and route review attention per field.
- C. Incorrect. Omitting scores for fields judged 'easy' removes exactly the data needed to verify that assumption; some of those omitted fields could still have meaningfully lower accuracy that only a labeled validation set would reveal.
- D. Incorrect. A free-text explanation without a numeric or ordinal score cannot be directly compared against a labeled validation set to compute a calibration curve or set a quantitative routing threshold.
4.Spending limited reviewer time where it counts
Reviewer capacity is always limited, so the review queue works as a triage system. Two kinds of extraction belong at the front. The first is fields whose calibrated confidence falls below their threshold. The second is extractions whose source is ambiguous or contradicts itself: two different totals on one invoice, a date that conflicts with another field, a document that fits none of the known types. The model may report high confidence and still be wrong in those cases, because the source doesn't support a single answer. Ambiguity and contradiction should therefore be separate routing signals, not something folded into one score.
Anthropic's prompting guidance suggests separating the two jobs. The model reports everything, with its confidence attached, and a separate stage decides what to act on, because "moving confidence filtering out of the finding step often helps." For extraction, that means the model doesn't quietly drop or quietly guess an uncertain field. It returns the field with its confidence and any conflict it noticed, and the routing layer, driven by calibrated thresholds, decides whether a human sees it. Then watch the queue itself. The content-moderation cookbook advises "Route needs_review to humans and watch its rate". If one field keeps showing up in review, look at its instructions or its segment.
5.Keep sampling what you have automated
Automating the high-confidence extractions doesn't end the measuring. If nobody ever looks at auto-accepted fields, you can't see their error rate, and you won't notice when a new vendor template, a new form version or a change in scan quality introduces an error pattern your validation set never contained. So keep sending a random sample of high-confidence, auto-accepted extractions to human review. The sample exists for measurement, not for catching individual errors. The monitoring guidance in Anthropic's writing on evals describes this post-launch job: "Production monitoring kicks in post-launch to detect distribution drift and unanticipated real-world failures."
The sample should be stratified, not a simple random draw from the whole stream. A simple random sample reproduces your traffic mix, so a document type that is 2% of volume gets 2% of the sample, too few examples to estimate its error rate. Stratified sampling draws a set number from each document type (and, where it matters, each field or confidence level), so every segment gets enough reviewed examples. Anthropic's own benchmark reporting uses the same idea to keep a small sample representative: one result was "measured on a 50-task subset stratified across all themes." The reviewed examples do two jobs. They give each segment a running error rate for its auto-accepted fields, which tells you whether a threshold still holds. They also surface novel errors, which you can add to the labeled validation set before recalibrating.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A high overall accuracy figure such as 97% is enough evidence to reduce human review across the whole pipeline.Why is that wrong?
An overall figure weighted by volume can hide a document type or field that performs far worse. Break accuracy down by document type and field, and reduce review only for segments that clear the bar themselves.
2.The model's self-reported confidence can be used directly as a review threshold, for example sending everything below 0.85 to review.Why is that wrong?
Raw confidence has to be checked against a labeled validation set, per field, before a threshold means anything. If it isn't, high-confidence fields can still be wrong often and reviewers stop trusting the score.
Covered in Field-level confidence, calibrated against labeled data
3.Once high-confidence extractions are auto-accepted, human review only needs to see the low-confidence ones.Why is that wrong?
Without a stratified random sample of auto-accepted, high-confidence extractions, their error rate can't be seen, and new error patterns from drift or new document formats go unnoticed.
Covered in Keep sampling what you have automated
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Measure the routing accuracy across different languages, aiming for no more than a 5–10% drop in accuracy for non-primary languages.”
↩︎ Why 97% overall can still mean a broken pipeline“Create a test set of edge cases and measure the routing accuracy, aiming for at least 80% accuracy on these challenging inputs.”
↩︎ Validate every segment before you automate it“aiming for consistent routing accuracy (within 2–3%) across all customer groups”
↩︎ Key concept - 2.https://claude.com/blog/carta-healthcare-clinical-abstractorSecondary source
“rather than staring at an aggregate score wondering what went wrong”
↩︎ Why 97% overall can still mean a broken pipeline“rather than staring at an aggregate score wondering what went wrong”
↩︎ Exam trap 1 - 3.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-4-8Official docs
“For each finding, include your confidence level and an estimated severity so a downstream filter can rank them.”
↩︎ Field-level confidence, calibrated against labeled data“moving confidence filtering out of the finding step often helps”
↩︎ Spending limited reviewer time where it counts - 4.https://platform.claude.com/cookbook/misc-building-evalsSecondary source
“A "golden answer" to which we compare the model output.”
↩︎ Field-level confidence, calibrated against labeled data“A "golden answer" to which we compare the model output.”
↩︎ Exam trap 2 - 5.
“Route needs_review to humans and watch its rate”
↩︎ Spending limited reviewer time where it counts - 6.https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligenceOfficial docs
“measured on a 50-task subset stratified across all themes”
↩︎ Keep sampling what you have automated - 7.
“Production monitoring kicks in post-launch to detect distribution drift and unanticipated real-world failures.”
↩︎ Keep sampling what you have automated“Production monitoring kicks in post-launch to detect distribution drift and unanticipated real-world failures.”
↩︎ Exam trap 3