Start typing to search the documentation.

Docs navigation

Compaction

Compaction makes room in a long session by replacing older conversation with a summary, so the session can keep going past the model’s context limit.

How it works

When a session approaches the model’s context limit, OpenCode summarizes everything except the most recent conversation, about 15,000 tokens by default. Once the summary is ready, it is placed in front of that recent conversation and the session continues from there.

before   [ system prompt ][ older conversation ............ ][ recent 15k ][ pending work ]
after    [ system prompt ][ summary ][ recent 15k ][ pending work ]

The model sees the summary as past conversation. Later compactions update the same summary rather than starting over.

Start

Automatic compaction is on by default:

opencode.jsonc
{
  "$schema": "https://opencode.ai/config.json",
  "compaction": {
    "auto": true,
    "keep": { "tokens": 15000 },
  },
}

Compaction is lossy. Increase keep.tokens when exact recent details matter, or request compaction yourself before the automatic one.

Manual

Request compaction at any time, even for a short history:

curl -X POST http://localhost:4096/api/session/ses_example/compact \
  -H 'content-type: application/json' \
  -d '{}'

The response confirms the request was accepted; it does not wait for the summary. Follow session.compaction.* events to learn when it finishes. See the API reference for the full operation.

RuleBehavior
TimingRuns at the next safe point, before prompts that are still waiting.
RepeatsRequests made while one is pending merge into it.
AutomationWorks when compaction.auto is false.

Pass an id to make an exact retry idempotent. Reusing that ID for a different request returns a conflict.

curl -X POST http://localhost:4096/api/session/ses_example/compact \
  -H 'content-type: application/json' \
  -d '{"id":"msg_compact_once"}'

Settings

Add compaction to any OpenCode configuration file:

opencode.jsonc
{
  "$schema": "https://opencode.ai/config.json",
  "compaction": {
    "auto": false,
    "keep": { "tokens": 24000 },
    "buffer": 16000,
  },
}
FieldDefaultBehavior
autotrueCompact automatically near the context limit, and recover once when a provider rejects a request as too long. Manual compaction always works.
keep.tokens15000Approximate recent conversation kept beside the summary. Larger values preserve more detail but leave less room for new work.
buffer10% of the limitTokens to keep free below the model’s limit. Larger values start automatic compaction earlier.

keep.tokens and buffer accept non-negative integers.

Providers

By default OpenCode writes the summary itself, using the session’s model. Some providers can compact the conversation on their side instead. Enable it per provider or model with settings.compaction:

opencode.jsonc
{
  "$schema": "https://opencode.ai/config.json",
  "providers": {
    "openai": {
      "settings": { "compaction": { "type": "native" } },
      "models": {
        "gpt-4.1": { "settings": { "compaction": { "type": "summary" } } },
      },
    },
  },
}

A model setting overrides the provider setting. native uses the provider’s compaction; summary uses OpenCode’s own.

With native compaction, OpenCode sends the conversation to the provider’s compaction endpoint. The provider returns an opaque, encrypted item that stands in for the older conversation. Your recent messages and attachments, up to keep.tokens, are kept alongside it:

before   [ older conversation ........................ ][ recent user messages ]
after    [ encrypted checkpoint + recent user messages ]
TopicBehavior
SupportOpenAI Responses models. Support varies by deployment and model.
ThresholdSame auto and buffer settings as summary compaction. Manual requests work too.
PortabilityAn encrypted checkpoint only works with the same provider, model, and endpoint. If you switch models, the session continues from the original conversation instead.
Too longIf the provider rejects a compaction as too long, OpenCode retries it with a smaller request. It does not switch to a summary.

Summaries

The summary is structured so another agent could pick the work up from it:

## Objective
Finish the authentication migration.

## Next Move
- Update the callback handler.
- Verify the login flow.

It covers the objective and requirements, decisions, completed and active work, blockers and next moves, and relevant files. In the recent conversation kept beside it, long tool output is shortened and attachment data is replaced by a short description.

After compaction, the current versions of your instruction files become the session’s baseline; see Instructions.

Limits

LimitResult
ModelCompaction uses the session’s model. There is no separate compaction model.
HistoryCompaction needs older conversation to replace. It cannot create room when a request is mostly fixed instructions and tool schemas.
SizeIf the conversation to compact is itself too long, OpenCode sends a shortened text version and may leave out the oldest exchanges.
RecoveryA request rejected as too long is compacted and retried once. A second rejection is returned as an error.
StorageEarlier messages remain stored even when they are no longer sent to the model.

For example, compaction cannot help here:

128k context = 120k fixed instructions and tools + 8k conversation

Migration

V1 used tail-turn and pruning behavior. V2 uses summaries and compaction.keep.tokens:

{
  "compaction": {
    "keep": { "tokens": 15000 }
  }
}

The settings and behavior on this page apply to V2.