CertSafari
    CLAUDE-CERTIFIED-ARCHITECT-FOUNDATIONS-CCAR-F · Lessons

    Domain 5 · Lesson 25/30

    Case Facts and Full History in Long Claude Conversations

    Manage conversation context to preserve critical information across long interactions

    15 min read
    2.5% of exam
    8 sources
    Published 29 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Explain why every API request must carry the complete conversation history, and what server-side context editing does and does not change about that
    • Identify which details a running summary or compaction is likely to blur, and why it matters that compaction is lossy
    • Design a persistent case-facts block, and a per-issue structured layer for multi-issue sessions, that sits outside the summarized history
    • Structure aggregated inputs so key findings sit at the beginning and detailed results sit under explicit section headers, to counter the lost-in-the-middle effect
    • Trim verbose tool outputs to the relevant fields before they accumulate in context
    • Specify subagent outputs as structured facts carrying dates, source locations and methodological context, instead of verbose content and reasoning chains

    Key concept

    Case facts layer — A structured block of exact transactional facts: amounts, dates, order numbers and statuses. The application extracts them and sends them with every request, outside the summarized history, so no summary can blur them.

    1.Each request carries the whole conversation

    A multi-turn conversation stays coherent because of what your application sends. The context-windows documentation describes each turn in two phases. The input phase holds all previous conversation history plus the current user message. The output phase is the model's reply, which becomes part of the input for the next turn. As turns go by, each user message and assistant response stays in the window, and earlier turns are kept whole. In practice your code keeps a message list and sends all of it every time.

    The basic multi-turn loop: append the user turn, send the full history, then append the assistant reply so the next request includes itpython
    history: list[BetaMessageParam] = []
    for turn, question in enumerate(QUESTIONS, start=1):
        history.append({"role": "user", "content": question})
        response = client.beta.messages.create(
            model="claude-opus-5-5",
            max_tokens=8192,
            system=SYSTEM,
            betas=["compact-2026-09-04"],
            messages=history,
        )
        history.append({"role": "assistant", "content": response.content})

    If you leave out the assistant turns, or send only the newest question, the model answers without the conversation that came before it. Its references to "the order" or "that date" then have nothing to point back to.

    That is the importance of passing the complete conversation history in every subsequent API request: it is what maintains conversational coherence. Each output phase generates a response that becomes part of the input for the next turn, so the chain only holds if every later request carries the earlier turns again. A request that sends a partial history leaves the model with a partial conversation.

    Sources12

    2.Why running summaries lose the details customers care about

    Sending everything every turn has a cost that goes beyond tokens. The context-windows documentation warns that more context isn't automatically better: as token count grows, accuracy and recall degrade. This effect is called context rot. Long sessions therefore get condensed at some point. Compaction summarizes a conversation that is nearing its limit and starts again from that summary. The client-side SDK variant generates a summary and replaces the full conversation history. After that, anything the summary left out is gone from what the model sees.

    The cookbook is frank about the trade-off. Compaction is lossy by design, and overly aggressive compaction can lose subtle but critical context whose importance only becomes apparent later. The exam guide names what is typically lost when summarization is applied again and again: numerical values, percentages, dates and customer-stated expectations, each flattened into vague wording. "$142.50 refund promised by Friday" becomes "a refund was discussed". The sources describe the mechanism (a lossy, whole-transcript operation) rather than listing those specific losses, but the mechanism explains why exact figures are the first to go.

    Three context-management primitives and what each does to earlier content
    PrimitiveScopeWhat survivesRisk to exact facts
    CompactionWhole transcript: user messages, assistant messages, tool calls, tool results and earlier compaction blocks are merged into one summaryArchitectural decisions, unresolved questions and key facts that the summary chooses to keepHigh. It is lossy by design, and details can be dropped before their importance is clear
    Tool-result clearingSub-transcript: replaces old tool_result blocks onlyUser messages, assistant reasoning and the tool_use recordLow for the conversation. Cleared payloads can be fetched again by calling the tool
    MemoryNotes persisted outside the context windowWhatever the agent writes and your storage backend keepsDepends on the guidance you give about what gets written

    The cookbook's first remedy is to replace the default compaction prompt so it keeps what your agent needs. That helps, but a summary is still a rewrite, and a rewrite can still get a number wrong. For facts that must stay exact, the more robust fix is to keep them out of the summary altogether.

    Sources123

    3.Keep transactional facts in a layer the summary cannot touch

    The pattern is to extract transactional facts (amounts, dates, order numbers, statuses) as soon as they appear. Store them as structured data and include that block in every prompt, separate from the history that gets summarized. This is the memory idea from the cookbook applied to a single session: the agent, or your code, writes notes that live outside the context window and brings them back in later. After a reset the notes are still there, word for word.

    The compaction documentation already rebuilds history from separate parts. In the keep-recent-turns variant, only the older messages are summarized. The new history is the returned compaction block followed by the recent turns, which are kept verbatim:

    Compacting only the older turns, then rebuilding history as the summary block plus the untouched recent turnspython
            split = -2 * KEEP_TURNS
            older, recent = history[:split], history[split:]
            summary = client.beta.messages.create(
                model="claude-opus-5-5",
                max_tokens=4096,
                system=SYSTEM,
                betas=["compact-2026-09-04"],
                messages=older,
                compaction={"type": "summarize"},
            )
            if summary.stop_reason == "compaction":
                history = [{"role": "assistant", "content": summary.content}, *recent]

    Your application chooses what sits next to the summary. A case-facts block assembled from extracted fields can be placed in every request, for example in the system prompt. It never passes through the summarization call, so compaction has nothing to paraphrase.

    Sessions with several open issues need one more step. A single narrative that mixes a delivery complaint with a warranty question is where ambiguity starts: a later mention of "the amount" could belong to either. Keep a separate structured entry for each issue, keyed by its own identifier, with its own amount and status. The model can then resolve a follow-up to the right record instead of guessing from blended prose.

    A billing support agent has been summarizing a lengthy chat every few turns to keep the prompt short. After the third summarization pass, the running summary reads "the customer wants a refund soon and mentioned an order from last month," even though the original messages contained an exact order number, a refund amount of $214.88, and a customer-stated deadline of the 15th. Which practice best addresses this failure mode going forward?

    A research assistant aggregates outputs from five subagents (market sizing, competitor pricing, regulatory risk, customer sentiment, distribution channels) into one long combined document that is then passed to a synthesis step. The synthesis step's final memo omits the regulatory risk finding, which appeared in the third of five sections in the middle of the document. What is the best way to structure the aggregated input to prevent this in future runs?

    Sources43

    4.Put key findings where the model reliably reads them

    Where a fact sits in a long input matters as much as whether it is there at all. This is the "lost in the middle" effect: models reliably process information at the beginning and end of long inputs but may omit findings from middle sections. Anthropic's Academy course links it to the serial position effect in human memory. In a 2023 Stanford test, accuracy was highest when a key fact appeared at the very beginning or very end of a long context window, and it dropped by more than 30% when the fact was buried in the middle. The course calls this structural rather than a quirk: transformer attention patterns weight the edges of the context window more heavily.

    Aggregated inputs are where this hurts most. When a coordinator pastes a dozen subagent reports into one prompt, the findings in the middle reports are the ones at risk. The mitigation has two parts. First, place a key findings summary at the beginning of the aggregated input, so the most important results sit in the position the model processes reliably. Second, organize the detailed results under explicit section headers, one per report or topic, so each finding is labeled and easy to locate instead of dissolving into one long block. The Academy advice is to put the most important instructions at the beginning and end of the context, and not to rely on the model giving equal weight to everything in between. The context engineering post recommends XML tags or Markdown headers to delineate sections.

    Sources56

    5.Trim tool outputs before they accumulate

    Tool results accumulate in context. Once a tool_result is appended, it becomes part of the conversation history and counts against the context budget on every subsequent turn. The tool-use documentation names accumulated tool_result blocks, alongside tool definitions, as what consumes your context window in long-running agents. That consumption is disproportionate to relevance: the exam guide's example is an order lookup that returns 40+ fields when only 5 are relevant to the customer's return. The other 35-odd fields consume tokens on every later request, and the context editing documentation notes that irrelevant content degrades model focus.

    The fix is to trim verbose tool outputs to only the relevant fields before they accumulate in context. For a return request, your code keeps the return-relevant fields from the order lookup (such as the order number, amount, purchase date and status) and drops the rest before the result is appended to history. The context engineering post states the same principle at the tool boundary: tools should promote efficiency by returning information that is token efficient.

    Trimming is not the same as the API's context-management primitives, and they compose. Tool-result clearing replaces old tool_result blocks with a placeholder later, once they are stale. Programmatic tool calling keeps intermediate results out of the conversation history entirely. Trimming acts the moment a result arrives, so the irrelevant fields never take up space in the first place, and the relevant ones can also feed the case-facts block.

    Sources7326

    6.Make upstream agents return compact, attributed facts

    Multi-agent systems move the same problem between agents. Anthropic reports that its multi-agent research system uses about 15 times more tokens than chats. In that system the subagents act as intelligent filters: they search, then return a list of findings to the lead agent so it can compile a final answer. The filtering is the point. When downstream agents have limited context budgets, modify the upstream agents to return structured data (key facts, citations, relevance scores) instead of verbose content and reasoning chains. The downstream synthesizer needs the facts, not the path taken to find them, and more context isn't automatically better.

    Structure alone does not guarantee accurate downstream synthesis. Require subagents to include metadata in their structured outputs: the date of each finding, its source location, and methodological context such as how a figure was measured. Without it, the synthesizer cannot tell an old estimate from a recent one, cannot reconcile two figures measured in different ways, and cannot attribute a claim. Anthropic's system shows how much work attribution takes once it is missing: a separate CitationAgent processes the documents and the research report to identify specific locations for citations, so that all claims are properly attributed to their sources.

    Sources81

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Because context editing trims history on the server, the client can store and resend only the trimmed conversation.Why is that wrong?

      Context editing is applied server-side before the prompt reaches Claude. The client keeps and resends the full, unmodified history and does not sync to the edited version.

      Covered in Each request carries the whole conversation

    2. 2.A high-fidelity compaction summary can be trusted to carry exact amounts, dates and promised deadlines forward.Why is that wrong?

      Compaction is lossy by design and can drop details whose importance only shows up later. Exact facts belong in a persistent block outside the summarized history.

      Covered in Why running summaries lose the details customers care about

    3. 3.If a finding in the middle of a long aggregated input is missed, adding more surrounding context will help the model catch it.Why is that wrong?

      More context pushes other material further into the middle, where recall is weakest. Put a key findings summary at the beginning and organize the details under explicit section headers.

      Covered in Put key findings where the model reliably reads them

    4. 4.Appending full, untrimmed tool responses is harmless because each one is only used on the turn it arrives.Why is that wrong?

      Once appended, a tool result stays in the history and is paid for on every later turn. Trim it to the relevant fields before it accumulates.

      Covered in Trim tool outputs before they accumulate

    5. 5.When a synthesis agent runs short of context, the best fix is to have it condense the subagents' verbose reports itself.Why is that wrong?

      By then the verbose reports have already used up its budget. The subagents should act as filters and return structured facts with citations and metadata.

      Covered in Make upstream agents return compact, attributed facts

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Input phase: Contains all previous conversation history plus the current user message”
      ↩︎ Each request carries the whole conversation
      “each user message and assistant response accumulates within the context window, and previous turns are preserved completely.”
      ↩︎ Each request carries the whole conversation
      “Output phase: Generates a text response that becomes part of the input for the next turn”
      ↩︎ Each request carries the whole conversation
      “As token count grows, accuracy and recall degrade, a phenomenon known as context rot.”
      ↩︎ Why running summaries lose the details customers care about
      “more context isn't automatically better”
      ↩︎ Make upstream agents return compact, attributed facts
    2. 2.
      “Your client application maintains the full, unmodified conversation history.”
      ↩︎ Each request carries the whole conversation
      “Generates a summary and replaces full conversation history.”
      ↩︎ Why running summaries lose the details customers care about
      “irrelevant content degrades model focus”
      ↩︎ Trim tool outputs before they accumulate
      “Your client application maintains the full, unmodified conversation history.”
      ↩︎ Exam trap 1
    3. 3.
      “it's lossy by design, but handles all context growth, not just tool results”
      ↩︎ Why running summaries lose the details customers care about
      “replacing the default compaction prompt to preserve what your agent needs”
      ↩︎ Why running summaries lose the details customers care about
      “the agent regularly writes notes persisted outside the context window, then pulls them back in at later times”
      ↩︎ Keep transactional facts in a layer the summary cannot touch
      “the results become part of the conversation history and count against the context budget on every subsequent turn”
      ↩︎ Trim tool outputs before they accumulate
      “maintaining critical context that would otherwise be lost across dozens of tool calls or across context resets”
      ↩︎ Key concept
      “overly aggressive compaction can lose subtle but critical context whose importance only becomes apparent later”
      ↩︎ Exam trap 2
      “the results become part of the conversation history and count against the context budget on every subsequent turn”
      ↩︎ Exam trap 4
    4. 5.
      “dropped by more than 30% when it was buried in the middle”
      ↩︎ Put key findings where the model reliably reads them
      “put your most important instructions at the beginning and end of the context”
      ↩︎ Put key findings where the model reliably reads them
      “Transformer attention patterns naturally weight the edges of the context window more heavily.”
      ↩︎ Put key findings where the model reliably reads them
      “Every piece of context you add pushes other pieces further into the middle”
      ↩︎ Exam trap 3
    5. 6.
      “using techniques like XML tagging or Markdown headers to delineate these sections”
      ↩︎ Put key findings where the model reliably reads them
      “returning information that is token efficient”
      ↩︎ Trim tool outputs before they accumulate
    6. 7.
      “Tool definitions and accumulated tool_result blocks consume your context window.”
      ↩︎ Trim tool outputs before they accumulate
      “The intermediate results never enter the conversation history.”
      ↩︎ Trim tool outputs before they accumulate
    7. 8.
      “multi-agent systems use about 15× more tokens than chats”
      ↩︎ Make upstream agents return compact, attributed facts
      “the subagents act as intelligent filters by iteratively using search tools to gather information”
      ↩︎ Make upstream agents return compact, attributed facts
      “identify specific locations for citations. This ensures all claims are properly attributed to their sources.”
      ↩︎ Make upstream agents return compact, attributed facts
      “the subagents act as intelligent filters by iteratively using search tools to gather information”
      ↩︎ Exam trap 5