What you will be able to do
- Explain what the context window is and how it differs from what Claude learned in training
- Identify everything that uses up space in a conversation, including instructions, pasted material and Claude's own replies
- Recognise how long conversations fail: context rot, lost-in-the-middle, and silent truncation at the edge of the window
- Explain why a larger context window does not remove the need to manage what goes into it
- Decide whether to restart a conversation, let it be summarised, or persist standing context into a Project or saved instructions
Key concept
Context window (working memory) — The fixed amount of material Claude can pay attention to while it writes a response: your messages, its replies, and anything you have pasted or attached. Anything outside the window has no effect on the answer, however important it was earlier in the conversation.
1.Working memory, not what Claude learned
When you chat with Claude, two kinds of knowledge are in play. The first is what the model learned in training: broad, general, and the same for every user. The second is the conversation in front of it: your request, the documents you pasted, the instructions you set, and everything said so far. Only the second kind is the context window. Anthropic's documentation describes it as working memory, and draws the line clearly: the window is separate from the training data.
This distinction explains a lot of everyday behaviour. Claude can discuss a well-known topic without being told about it, because that comes from training. It cannot know your team's agreed deadlines, your style guide or a correction you made yesterday unless that material is in the current window. The Academy course puts the rule bluntly: the model can attend to what is inside the window and cannot attend to anything outside it.
It will not. The rules existed only in last week's window. A new chat starts with a new, empty working memory. Everything else in this lesson, including when to restart, when to summarise and when to persist, comes from that one fact.
That one fact leaves three ways to handle a conversation that is outgrowing its window, and the objective names them: restart, summarize, or persist. Restarting means opening a fresh chat and supplying only the context the next task needs. The Academy course says that knowing how the window works tells you how to structure context, when to front-load, and when to start fresh. Summarising means condensing the older turns into a short summary that stands in for them, so the conversation can carry on. Anthropic's platform documentation calls this compaction: it replaces the older turns of a conversation with a summary that Claude writes, and it keeps the active context small because response quality degrades as a conversation grows. Persisting means moving durable knowledge out of any single chat and into standing context, such as a Project, saved instructions or uploaded reference documents, so every new conversation starts with it already in place.
Summarising trades detail for room. The summary is written to carry the state, next steps and learnings needed to continue, and the raw history is replaced by it, so anything the summary leaves out is no longer in working memory. Choose it when the thread's accumulated decisions still matter to the next step. Restart when they do not: a clean window beats a summary of material the next task never needed. Persist when you keep re-explaining the same background. The Academy course asks you to annotate which tasks need standing context set up (a project, saved instructions, uploaded reference docs) and which work fine cold. If a piece of context is needed in most conversations, it belongs in a Project, not in each chat.
Let the chat be summarised (or ask Claude for a summary and restart from it): the early decisions still matter, so a condensed summary that carries the state and next steps is better than a cold restart or pushing on as quality degrades. The brand guidelines are standing context needed in every conversation, so persist them into a Project or saved instructions rather than pasting them each time.
2.What uses up the space
People tend to picture the window as holding their own messages. It holds far more than that. The platform documentation lists what counts: the instructions set for the conversation, every message (including documents, images and tool results), and the output Claude produces, including any extended thinking. Every turn is added to what the next turn has to read, so a conversation's size grows with each reply as well as each question.
| Item | Counts toward the window? | Practical consequence |
|---|---|---|
| Standing instructions for the conversation | Yes, on every turn | Long instructions leave less room for the material you actually want analysed |
| Your messages, pasted documents and images | Yes | Pasting a large reference set uses space before you ask anything |
| Claude's replies | Yes, and each one becomes input for the next turn | Long answers speed up how fast the window fills |
| Extended thinking produced for a turn | Yes | Deep-reasoning turns cost space as well as time |
Taken together, the table shows that context is a budget you spend on purpose. The context-editing documentation calls it a finite resource with diminishing returns, and warns that irrelevant content degrades the model's focus. Material you add 'just in case' is not harmless: it takes space, and it competes for attention with the material that matters.
Cut the standing instructions down to what each task actually needs. They are re-read on every turn, so they use space in every exchange and dilute attention on the pasted documents. Shorter, sharper instructions give the working material more room.
3.How long conversations fail
A long conversation does not fail in one way. There are three failure modes, and the first begins well before the window is full. As the token count grows, accuracy and recall decline. The documentation calls this context rot. A conversation can be well within its limit and still give worse answers than it did at the start, simply because there is more for the model to sift through.
The second is position. The Academy course names burying critical information in the middle of long input as a limitation zone, and its exercises ask you to test it: hide one important instruction in the middle of a long document, ask a question that depends on it, then move the instruction to the top and compare. Where you put something affects whether it is used.
The third is the edge. The course says the property has a cliff rather than a gradient: things work until they don't. The failure is silent truncation, and you won't always be warned. The platform documentation adds that chat interfaces such as claude.ai can manage the window on a rolling, first-in-first-out basis. In practice, the oldest material is what is most likely to stop influencing answers. That is often exactly where the ground rules were set.
A teacher pastes a whole term of student feedback into a single chat and keeps asking Claude to summarise themes as the conversation grows very long. What is the greatest risk to the quality of her summaries?
Correct answer: B — Detail from early in the chat may be condensed away, so her theme counts can quietly omit some feedback.
- A. Incorrect. The model in use does not change because a conversation is long. Approaching the limit affects available working memory, not model quality.
- B. Correct. As a conversation approaches the context limit, earlier messages are condensed to make room. The summary still reads confidently, but it may no longer rest on every comment she pasted in.
- C. Incorrect. Reaching a length limit does not force a plan upgrade; she can continue in a new chat or move the material into a Project.
- D. Incorrect. Condensing earlier messages does not erase them. The full chat history stays readable, so she can still go back and check.
4.Why a bigger window is not the whole answer
Larger windows help. The documentation says a larger context window lets the model handle more complex and lengthy prompts, so a model with more room suits a genuinely large body of material that has to be considered together. Size does not cure context rot, though. The same page says more context isn't automatically better, and that curating what is in the window matters as much as how much room there is. The Academy course lists larger windows alongside memory features, compaction and projects as ways to push the cliff further out. It does not say they remove it.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The biggest context window is always the best choice, because more context always gives better answers.Why is that wrong?
A larger window fits more material, but recall and accuracy decline as the token count grows. Choosing what goes into the window matters as much as how big it is.
Covered in Why a bigger window is not the whole answer
2.Claude will clearly warn you before earlier parts of a long conversation stop affecting its answers.Why is that wrong?
Hitting the limit is a cliff, not a gradual slope, and the typical failure is silent truncation. Early material can stop influencing answers without any notice.
Covered in How long conversations fail
3.Summarising a long conversation keeps everything that was said, so nothing is lost and it is always better than restarting.Why is that wrong?
A summary replaces the raw history. Whatever it does not capture is no longer in working memory. If the earlier decisions do not matter to the next task, a clean restart is the better choice.
Covered in Working memory, not what Claude learned
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“This is different from the large corpus of data the language model was trained on”
↩︎ Working memory, not what Claude learned“Compaction automatically summarizes earlier parts of the conversation on the server, so the conversation can continue past the context window limit.”
↩︎ Working memory, not what Claude learned“Everything in the request counts toward the context window”
↩︎ What uses up the space“As token count grows, accuracy and recall degrade, a phenomenon known as context rot.”
↩︎ How long conversations fail“Chat interfaces such as claude.ai can also manage the context window on a rolling”
↩︎ How long conversations fail“A larger context window allows the model to handle more complex and lengthy prompts”
↩︎ Why a bigger window is not the whole answer“As token count grows, accuracy and recall degrade, a phenomenon known as context rot.”
↩︎ Exam trap 1 - 2.
“Everything the AI is paying attention to lives inside a fixed-size workspace called the context window.”
↩︎ Working memory, not what Claude learned“Knowing how the window works tells you how to structure context, when to front-load, and when to start fresh.”
↩︎ Working memory, not what Claude learned“which tasks need standing context set up (a project, saved instructions, uploaded reference docs) to be worth running, and which work fine cold”
↩︎ Working memory, not what Claude learned“burying critical info in the middle of long input”
↩︎ How long conversations fail“Silent truncation is the failure mode”
↩︎ How long conversations fail“Memory features, compaction, projects, larger windows, and multi-agent workflows all exist to push this cliff further out.”
↩︎ Why a bigger window is not the whole answer“Everything the AI is paying attention to lives inside a fixed-size workspace called the context window.”
↩︎ Key concept“This property has a cliff rather than a gradient.”
↩︎ Exam trap 2 - 3.
“Compaction replaces the older turns of a conversation with a summary that Claude writes on the server”
↩︎ Working memory, not what Claude learned“it keeps the active context small, because response quality degrades as a conversation grows”
↩︎ Working memory, not what Claude learned - 4.
“Write down anything that would be helpful, including the state, next steps, learnings etc.”
↩︎ Working memory, not what Claude learned“where the raw history above may not be accessible and will be replaced with this summary”
↩︎ Exam trap 3 - 5.
“context is a finite resource with diminishing returns, and irrelevant content degrades model focus”
↩︎ What uses up the space