What you will be able to do
- Explain the two costs of loading every tool definition and instruction into context up front
- Describe how Agent Skills use progressive disclosure across three loading levels
- Identify which content costs tokens at startup, on trigger, and only when accessed
Key concept
Progressive disclosure (progressive discovery) — A context strategy where the agent sees only a lightweight index of what it can do and loads the full instructions, tool definitions or files when a task needs them. The opposite is a monolithic context, where everything is placed in the context window before any work starts.
1.What a monolithic context costs
The simplest way to give an agent capabilities is the monolithic approach: put every tool definition, and every piece of guidance, into the context window at the start. It works when the agent has only a few tools. As the library grows, it causes two separate problems.
The first is context bloat. The documentation gives a concrete example: a typical multiserver setup of GitHub, Slack, Sentry, Grafana and Splunk can use about 55k tokens on definitions alone, before Claude does any work. The Agent SDK documentation gives a smaller rule of thumb: 50 tools can use 10–20K tokens, which leaves less room for the actual task.
The second is selection accuracy, and it matters more. The documentation states that Claude's ability to pick the right tool degrades once you exceed 30–50 available tools. So a monolithic toolset does not just cost more tokens. It makes the agent choose tools less reliably, because it has too many options to pick from.
Anthropic's context-engineering guidance explains why this happens. As a context window fills up, the model's ability to recall information from it goes down, so context has to be treated as a finite resource that gives less back the more you add. The guidance names bloated tool sets as one of the most common failure modes it sees: sets that cover too much functionality or leave the agent unsure which tool to use. The same thinking applies to prompts. Rather than stuffing a prompt with a laundry list of edge cases, aim for the smallest set of high-signal tokens that does the job.
A team builds an internal agent that calls exactly 7 tools, nearly all of which are needed on almost every request, and the combined tool definitions total roughly 1,200 tokens. Which context strategy should they use for tool loading?
Correct answer: A — Load all seven tool definitions into context on every turn, since the toolset is small and every tool is used regularly.
- A. Correct. With fewer than roughly 10 tools whose definitions fit comfortably in context, loading everything upfront (monolithic loading) is typically faster than paying the extra search round-trip that tool search introduces.
- B. Tool search exists to solve context bloat and selection accuracy problems that appear at larger tool counts (dozens to thousands); for 7 small, frequently-used tools it adds an unnecessary search step.
- C. Deferring every tool, including in a setup this small, adds a search round-trip on tools that are needed almost every request, and at least one tool must stay non-deferred for search to even function.
- D. Splitting a 7-tool, 1,200-token catalog across servers and enabling search on each adds architectural complexity and search overhead with no context-bloat problem to solve.
2.Progressive disclosure in Agent Skills: three loading levels
Agent Skills are Anthropic's clearest example of the progressive alternative. A Skill is a directory on the filesystem of Claude's virtual machine. It holds instructions, code and reference material, and each kind of content is loaded at a different time. Nothing beyond a short index is loaded until a task needs it.
| Level | When loaded | Token cost | Content |
|---|---|---|---|
| Level 1: Metadata | Always (at startup) | ~100 tokens per Skill | name and description from YAML frontmatter |
| Level 2: Instructions | When Skill is triggered | Under 5k tokens | SKILL.md body with instructions and guidance |
| Level 3+: Resources | As needed | None until accessed | Bundled files; scripts run through bash and only their output enters context |
Level 1 is the index. Claude loads each Skill's YAML frontmatter at startup and includes it in the system prompt. Because only the name and description are loaded, you can install many Skills without filling up the context.
---
name: pdf-processing
description: Extract text and tables from PDF files, fill forms, merge documents. Use when working with PDF files or when the user mentions PDFs, forms, or document extraction.
---The description has a design consequence. Claude compares each request against it to decide whether to trigger the Skill, so it has to say both what the Skill does and when to use it. In a progressive design, the index is how Claude finds a capability. If the description is vague, Claude cannot find the Skill, even though its instructions are there.
Sources4
3.What never enters context unless it is used
Level 3 is where progressive disclosure differs most from a monolithic prompt. When a Skill is triggered, Claude uses bash to read SKILL.md into context. If those instructions point to other files, Claude reads those only when the task needs them. A Skill can bundle dozens of reference files, and the ones that are not read stay on the filesystem and cost zero tokens.
Scripts save even more. When Claude runs a bundled script, the script's code never enters the context window. Only its output does, for example "Validation passed" or a specific error message. So a Skill can include full API documentation, large datasets or many examples without a context penalty for the parts that go unused.
The frontmatter was already in the system prompt from startup. Claude then runs cat pdf-processing/SKILL.md, which loads the instructions. It decides form filling is not needed, so FORMS.md is never read. REFERENCE.md and fill_form.py also stay on the filesystem, so only the metadata and the SKILL.md body use tokens.
Sources4
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Installing many Skills bloats the context window, because each Skill's instructions are loaded at startup.Why is that wrong?
Only the Level 1 metadata (name and description, about 100 tokens per Skill) is loaded at startup. The SKILL.md body is loaded only when the Skill is triggered.
Covered in Progressive disclosure in Agent Skills: three loading levels
2.If the model's context window is large enough to hold every tool definition, loading them all up front has no downside.Why is that wrong?
Fitting in the window is not the only concern. Tool selection accuracy drops once the agent has more than 30–50 tools, however much room is left.
Covered in What a monolithic context costs
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“A typical multiserver setup (GitHub, Slack, Sentry, Grafana, and Splunk) can consume ~55k tokens in definitions before Claude does any work.”
↩︎ What a monolithic context costs“Claude's ability to pick the right tool degrades once you exceed 30–50 available tools.”
↩︎ What a monolithic context costs“Claude's ability to pick the right tool degrades once you exceed 30–50 available tools.”
↩︎ Exam trap 2 - 2.
“Tool definitions can consume large portions of the context window (50 tools can use 10-20K tokens), leaving less room for actual work.”
↩︎ What a monolithic context costs - 3.
“Context, therefore, must be treated as a finite resource with diminishing marginal returns.”
↩︎ What a monolithic context costs“bloated tool sets that cover too much functionality or lead to ambiguous decision points about which tool to use”
↩︎ What a monolithic context costs“good context engineering means finding the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome.”
↩︎ What a monolithic context costs - 4.
“until a Skill is triggered, only its name and description occupy context.”
↩︎ Progressive disclosure in Agent Skills: three loading levels“The description is what Claude matches your request against when determining whether to trigger the Skill”
↩︎ Progressive disclosure in Agent Skills: three loading levels“The rest stay on the filesystem and cost zero tokens.”
↩︎ What never enters context unless it is used“When Claude runs validate_form.py, the script's code never loads into the context window.”
↩︎ What never enters context unless it is used“Progressive disclosure ensures only relevant content occupies the context window at any given time.”
↩︎ What never enters context unless it is used“This filesystem-based architecture enables progressive disclosure: Claude loads information in stages as needed, rather than consuming context upfront.”
↩︎ Key concept“until a Skill is triggered, only its name and description occupy context.”
↩︎ Exam trap 1