What you will be able to do
- Explain how next-token generation produces both fluent answers and fabricated specifics
- Explain why the same prompt can return different outputs across runs
- List everything that counts toward Claude's context window and name which models have 1M or 200k windows
- Tell zero-shot, single-shot and multishot prompting apart, and write multishot examples that are relevant, diverse and structured
- Choose between extended thinking, adaptive thinking, effort levels and fast mode for a given latency, cost or reasoning need
Key concept
Next-token prediction — A language model builds its answer one piece at a time. At each step it picks a continuation based on patterns in its training, rather than looking up a stored answer. The same mechanism explains its fluency, its fabrications and why its output varies from run to run.
1.How an LLM writes: one token after another
Anthropic's Academy course on AI capabilities says generative AI is closer to a very sophisticated autocomplete than to a search engine. The model writes its answer word by word, and at each step it continues with whatever tends to follow what it has written so far. It does not retrieve a stored answer. It generates one.
That single mechanism explains both what the model is good at and where it fails. Tasks that look like patterns the model has seen many times fall in its capability zone: summarizing, reformatting, explaining common concepts. Novel or sparse territory is its limitation zone. So is any task that depends on telling what is true apart from what merely sounds true. The course puts it plainly: the same property gives you both fluency and hallucination.
Fabrications don't turn up evenly across a response. They cluster in specific details: names, dates, statistics, citations, URLs and quotes. The more precise a claim is, the more it needs checking. Product features such as citations, uncertainty signalling, constrained generation and generator-verifier loops exist to push this limit further out.
Sources1
2.Sampling: why the same prompt gives different answers
Probably not. The next token is sampled, not retrieved, so two runs of the same request can take different paths. The Academy exercise makes this point directly. It has you run the same specific-facts request in a fresh conversation and compare what stayed the same with what changed. It calls that variation sampling at work.
In practice, this non-determinism means one good run tells you little about the next. When you evaluate a prompt or a setting, compare outputs across several runs rather than trusting a single one. It also matters for the specific details from the previous section: a citation that shows up in one run and not in another deserves suspicion.
Sources1
3.The context window: the model's working memory
The context window is all the text the model can refer to while it generates a response, and that includes the response itself. It is not the training data. Anthropic calls it the model's "working memory". Its size is measured in tokens, which is the unit the API counts, limits and bills. The sources here don't explain how text is split into tokens. What the exam tests is what gets counted.
Everything in a request counts toward the window: the system prompt, every message (including tool results, images and documents), and your tool definitions. The output Claude generates for the turn counts too. In a conversation, each response becomes input for the next turn, so usage grows with every turn. Every response reports its consumption in the usage field, and the token counting API estimates a request before you send it.
With prompt caching, the input count is split across input_tokens, cache_read_input_tokens and cache_creation_input_tokens, and all three count toward the window. Caching changes what you pay for a cached prefix. It does not free any space.
| Models | Context window | Notes |
|---|---|---|
| Claude Fable 5.1, Claude Mythos 5.1, Claude Opus 5.5, Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Sonnet 5, Claude Sonnet 4.6 | 1M tokens | The default, with no beta header needed. Up to 128k output tokens (max_tokens) per request. |
| Claude Sonnet 4.5, Claude Haiku 4.5, and other Claude models | 200k tokens | Up to 100 images or PDF pages per request, compared with 600 on 1M-token models. |
A bigger window is not automatically better. As the token count grows, accuracy and recall get worse, which Anthropic calls context rot. That makes choosing what goes into the context as important as how much room there is. For long conversations and agentic workflows, the primary strategy is server-side compaction, which summarizes earlier parts of the conversation on the server so the conversation can continue past the limit.
A team is building an autonomous coding agent expected to run for well over 30 minutes, making repeated tool calls and exploring a large codebase, with a token budget in the millions. Which effort level should they start with?
Correct answer: D — xhigh
- A. Low effort is the most conservative setting and is intended for simple, short-scoped tasks, not extended multi-tool agentic exploration.
- B. Medium effort offers moderate token savings for balanced workloads but is not tailored to long-horizon agentic work spanning many tool calls.
- C. High is the API default and suits complex reasoning and coding, but it is not the level specifically designed for long-running agentic work exceeding 30 minutes.
- D. Xhigh is described as extended capability for long-horizon agentic and coding tasks over 30 minutes with token budgets in the millions, matching this scenario exactly.
Sources2
4.Zero-shot, single-shot and multishot prompting
The "shot" in these names is the number of worked examples you include in the prompt. Anthropic's prompting guide covers the multi-example form, which it calls few-shot or multishot prompting, and describes examples as one of the most reliable ways to steer Claude's output format, tone and structure. The guide supplied here doesn't separately define zero-shot (instructions only, no examples) or single-shot (one example). Read those terms by the same counting rule: zero examples or one example.
Examples work best when they meet three conditions. They should be relevant, meaning they closely mirror your real use case. They should be diverse: cover edge cases and vary enough that Claude doesn't pick up patterns you never intended, such as copying one example's length or opening. And they should be structured: wrap each one in <example> tags, and a set in <examples> tags, so Claude can tell the examples apart from the instructions.
It breaks diversity. Claude may treat shortness or politeness as part of the pattern and struggle with long or angry emails. Add examples that cover those edge cases.
Sources3
5.Model options: thinking, effort and fast mode
Prompting is one lever. Claude also offers model-side options that trade capability, latency and cost. Multi-shot examples steer the format, tone and structure of an answer. The options below change how much the model reasons and how quickly it produces output.
Extended thinking, in manual mode, gives you direct control over how much Claude thinks. You set a token budget with thinking: {type: "enabled", budget_tokens: N}, and Claude thinks against that budget before its final answer. The budget has a minimum of 1,024 tokens and must be less than max_tokens. It is a target rather than a strict cap: Claude may stop reasoning early, and max_tokens remains the hard ceiling on total output.
Adaptive thinking removes the manual budget. With thinking: {type: "adaptive"}, Claude decides when and how deeply to think based on the request, so thinking token usage varies from request to request. On Claude Opus 5.5, adaptive thinking is always on and can't be turned off. Either way, thinking helps most on complex tasks such as math, coding and analysis, and it has a cost: thinking tokens are billed as output tokens and count toward max_tokens.
Effort levels are set with output_config.effort and control how many tokens Claude spends responding. The levels run from low (most efficient, some capability reduction) through medium and high (the default on most models) to xhigh and max. Effort affects all tokens in the response, including text, tool calls and thinking, so it works whether or not thinking is enabled. Claude Opus 5.5 defaults to medium rather than high.
Fast mode runs the same model with a faster inference configuration, giving up to 2.5x higher output tokens per second at premium pricing. You opt in with speed: "fast" and the fast-mode-2026-02-01 beta header. It is not a different or weaker model, and its gain is in output tokens per second, not time to first token. Switching between fast and standard speed invalidates the prompt cache, and fast mode is not available with the Batch API.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Prompt caching frees context-window space, so you can cache a large document and fit more into the window.Why is that wrong?
Cached tokens still occupy the context window. Caching changes only what you pay for them.
2.A larger context window always improves results, so you should fill it with as much material as possible.Why is that wrong?
Accuracy and recall degrade as the token count grows (context rot), so choosing what goes into context matters as much as its size.
3.The same prompt sent twice will return the same answer.Why is that wrong?
Generation samples each next token, so repeated runs of the same request can differ.
Covered in Sampling: why the same prompt gives different answers
4.Fast mode gives you a smarter or different model, so it improves answer quality.Why is that wrong?
Fast mode runs the same model with a faster inference configuration. It changes output speed, not intelligence or capabilities.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://academy.claude.com/courses/ai-capabilities-and-limitations/next-token-predictionOfficial docs
“It writes answers word by word based on what tends to follow what.”
↩︎ How an LLM writes: one token after another“That single property gives you both the fluency and the hallucination.”
↩︎ How an LLM writes: one token after another“Fabrication concentrates in specificity: names, dates, statistics, citations, URLs, quotes.”
↩︎ How an LLM writes: one token after another“The variation you see is Next Token Prediction's sampling at work.”
↩︎ Sampling: why the same prompt gives different answers“It writes answers word by word based on what tends to follow what.”
↩︎ Key concept“The variation you see is Next Token Prediction's sampling at work.”
↩︎ Exam trap 3 - 2.
“The "context window" refers to all the text a language model can reference when generating a response, including the response itself.”
↩︎ The context window: the model's working memory“This is different from the large corpus of data the language model was trained on”
↩︎ The context window: the model's working memory“the system prompt, every message in messages (including tool results, images, and documents), and your tool definitions”
↩︎ The context window: the model's working memory“Generates a text response that becomes part of the input for the next turn”
↩︎ The context window: the model's working memory“Other Claude models, including Claude Sonnet 4.5, have a 200k-token context window.”
↩︎ The context window: the model's working memory“A single request to any of them can generate up to 128k output tokens (max_tokens).”
↩︎ The context window: the model's working memory“As token count grows, accuracy and recall degrade, a phenomenon known as context rot.”
↩︎ The context window: the model's working memory“Cached prompt prefixes still occupy the context window: prompt caching changes what you pay for those tokens, not whether they count.”
↩︎ Exam trap 1“As token count grows, accuracy and recall degrade, a phenomenon known as context rot.”
↩︎ Exam trap 2 - 3.https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practicesOfficial docs
“A few well-crafted examples (known as few-shot or multishot prompting) improve accuracy and consistency.”
↩︎ Zero-shot, single-shot and multishot prompting“Examples are one of the most reliable ways to steer Claude's output format, tone, and structure.”
↩︎ Zero-shot, single-shot and multishot prompting“Diverse: Cover edge cases and vary enough that Claude doesn't pick up unintended patterns.”
↩︎ Zero-shot, single-shot and multishot prompting“Wrap examples in <example> tags (multiple examples in <examples> tags) so Claude can distinguish them from instructions.”
↩︎ Zero-shot, single-shot and multishot prompting“A few well-crafted examples (known as few-shot or multishot prompting) improve accuracy and consistency.”
↩︎ Model options: thinking, effort and fast mode - 4.
“The budget is a target rather than a strict cap.”
↩︎ Model options: thinking, effort and fast mode - 5.
“which lets Claude decide when and how deeply to think based on the request”
↩︎ Model options: thinking, effort and fast mode“the tokens Claude spends reasoning are billed as output tokens”
↩︎ Model options: thinking, effort and fast mode - 6.
“The effort parameter lets you control how many tokens Claude spends when responding to requests.”
↩︎ Model options: thinking, effort and fast mode“Because effort applies to every output token, it works whether or not thinking is enabled.”
↩︎ Model options: thinking, effort and fast mode - 7.
“Fast mode runs the same model with a faster inference configuration. There is no change to intelligence or capabilities.”
↩︎ Model options: thinking, effort and fast mode“Switching between fast and standard speed invalidates the prompt cache.”
↩︎ Model options: thinking, effort and fast mode“Fast mode runs the same model with a faster inference configuration. There is no change to intelligence or capabilities.”
↩︎ Exam trap 4