Conceptual illustration of many conversation notes condensed into a handover sheet, with recent notes kept separate.AI-generated conceptual illustration: a concise handover preserves key decisions while recent notes remain separate.

Anthropic’s 14 September 2026 API update gives developers more control over when a long conversation is condensed. For a business building an assistant that works across many steps, that matters because the useful history is rarely just the latest message. Earlier decisions, corrections and unresolved questions can still determine what the assistant should do next.

The new option is on-demand compaction, announced in the Claude Platform release notes. It is an API beta, not a newly announced button for every Claude chat user. Developers can request a summary separately, continue working while it is prepared and retain recent conversation turns after that summary.

For a business owner, the practical question is straightforward: can the assistant shorten its working history without losing the details that control the next decision? That is the question to put into a pilot, rather than assuming a longer-running conversation is automatically a more reliable one.

What is new, and what already existed?

Claude already had threshold-based compaction, where the API manages summarisation as a conversation reaches a configured size. The September update adds application-controlled timing through the top-level compaction parameter and the compact-2026-09-04 beta header.

The distinction changes who decides when to summarise. A developer can choose a natural handover point in a workflow, rather than relying solely on an automatic size trigger. The release notes also describe keeping recent turns word for word after the summary. That can be useful when the newest correction must remain explicit.

This is a feature for teams integrating the Claude API. It does not establish a new storage policy for a business, replace its project records or prove that every summarised detail will remain available later.

Availability and implementation boundaries

The on-demand section of Anthropic’s compaction documentation identifies this option as available on the Claude API, but not Amazon Bedrock or Google Cloud. Check that specific section rather than assuming the broader compaction compatibility list applies to every variant.

The response is a signed compaction block. Subsequent requests place it first, replacing the messages it summarises. The documentation warns that leaving summarised messages before that block produces an error. Recent turns that should remain verbatim need to be excluded from the summary request and retained after the returned block.

There are additional conditions for preserving thinking in retained turns, including unchanged relevant system instructions and tool definitions. An integration owner should check those conditions for the chosen model and workflow. Treat that review as an implementation task, not evidence that your deployment is compatible.

A useful pilot: an assistant preparing a client handover

Consider a hypothetical agency assistant that helps prepare a client handover. Its conversation contains an initial brief, several draft plans, the client’s corrections and a final review. Most early brainstorming is no longer important. A rejected promise, an agreed deadline and the person responsible for approval still matter.

For example, let the invented client replace a Friday delivery date with Monday, reject a promise of round-the-clock support and leave final sign-off with Maya. A successful handover must retain Monday, omit the rejected support promise and ask Maya for approval. These are test inputs, not reported customer results.

Build a small set of invented test conversations before using real client work. Put a few details in each conversation that should survive every handover:

  • A decision that replaced an earlier decision.
  • A constraint that rules out an otherwise attractive suggestion.
  • An unresolved question that must not be presented as settled.
  • A source document that the next reviewer needs to consult.

Then ask the assistant to continue the same task with the original history and with the proposed shortened history. Compare the resulting actions, not just how fluent the answers sound. If the shorter version confidently revives a rejected promise, the pilot has found a meaningful problem even if its response is faster.

Use the same reviewer and questions for both versions. Keep the comparison small enough that a person can identify why an answer changed. The aim is to discover what the workflow needs to preserve, not to manufacture an impressive score.

Define a handover contract before choosing a trigger

A handover contract is a short statement of what must be true before the assistant moves to its next stage. For the agency example, it might require the approved scope, the latest deadline, the approval owner and any unresolved dependencies to remain available.

Here is a compact review prompt to adapt for your own test. It is an evaluation prompt, not an API configuration:

Compare this proposed handover with the approved project record. Identify changed decisions, missing constraints, unresolved questions presented as settled, and unsupported commitments. For each issue, point to the relevant record. Do not approve or send the handover.

Keep the authoritative project record outside the assistant’s summary. That gives the reviewer something stable to check. It also makes recovery practical if a conversation becomes confused: the team can return to the approved record rather than treating a generated summary as the sole account of what happened.

A summary is most useful when everyone knows its job. Asking it to carry every draft, attachment and decision indefinitely makes its success difficult to judge. Define the next task, preserve the information needed for that task and keep important source material accessible through the normal business process.

Measure the work around the model

Do not assess the change using response speed alone. Record how much time the reviewer spends finding omissions, how often the assistant repeats completed work and whether a correction stays respected later in the conversation.

For example, an assistant might finish drafting sooner but cause the reviewer to reopen several source documents to reconstruct a missing condition. That may still be acceptable in one workflow and unacceptable in another. The decision depends on the consequence of the omission and the overall work required.

Include failure handling in the pilot. Decide what the application should do if no usable summary arrives, who can review a questionable handover and how the team returns to its last approved state. Avoid connecting an unproven continuation directly to an irreversible business action.

Our guide to using AI in a small business without losing control explains how to put review and approval around a pilot. If the workflow produces marketing material, the small-business content system provides a useful place to define the editorial approval stage.

For persistent notes between coding sessions, compare this with our Grok Build memory review guide. Its three-check pilot separates relevant recall, changed instructions and project boundaries.

Reader Q&A

Is on-demand compaction a new Claude chat setting?

The announcement concerns the Claude API. It should not be described as a new setting available to every person using the Claude chat app.

How does it differ from threshold compaction?

Threshold compaction manages summarisation when a configured size is reached. On-demand compaction lets the application choose when to request a summary.

Can a business use it without a developer?

A business using a third-party assistant should ask its supplier whether the feature is supported and how continuity is tested. The announcement itself describes an API integration, not a no-code rollout.

Does a shorter history guarantee a better answer?

No. Test whether the continued workflow respects approved decisions, constraints and unresolved questions. A fluent answer can still omit something important.

What should remain outside the summary?

Keep authoritative records, approvals and important source material in the normal business systems. A generated summary should not become the only copy of information needed to check or recover the work.

What is a sensible first step?

Choose one bounded workflow, create a few synthetic conversations with deliberate corrections and constraints, and compare continuation with and without the proposed shortened history before expanding the pilot.

What to do next

If a team already maintains a long-running Claude API integration, this update is worth an implementation review. Ask the developer to explain the intended handover point, the information that must survive it and the recovery route if it fails. Those answers make the feature assessable in business terms before it becomes part of everyday work.

author avatar
Garry Knight
I'm Garry Knight, the person behind Prodify Digital. I write about email list building, email marketing, SEO, AI search and the tools that connect them. My aim is to make online marketing easier to understand, so creators and small business owners can make informed decisions about building an audience and keeping people engaged. Here you'll find straightforward guides and product reviews that explain what something does, where it fits and which limitations matter. The focus is on clear explanations and useful next steps—not hype, shortcuts or promises of easy earnings.
One thought on “Claude Adds On-Demand Compaction: What Changes for Business AI Workflows”

Leave a Reply

Discover more from Prodify Digital

Subscribe now to keep reading and get access to the full archive.

Continue reading