> ## Documentation Index
> Fetch the complete documentation index at: https://docs.budecosystem.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Context Compaction

> Keep long agent conversations inside the model's context window

## Overview

An agent replays its whole conversation to the model on every turn. A long conversation eventually
outgrows the model's context window, and the model refuses it with a
`context_length_exceeded` error.

With **context compaction** on, Bud summarises the older turns of the conversation into a
**checkpoint** once the next request would pass a threshold. From then on the agent sends the
checkpoint plus the turns after it. When the conversation grows past the threshold again, the next
checkpoint absorbs the previous one.

* The summary is written by the **agent's own model**, with the agent's own credential, and is billed
  to the agent's project like any other call.
* The stored conversation is **never changed**: `GET /v1/responses/{id}` and `input_items` keep
  returning the original turns.
* Turns below the threshold are sent exactly as before.

```mermaid theme={null}
graph LR
    A[Next request] --> B{Size vs threshold}
    B -- below --> C[Send as is]
    B -- past threshold --> D[Summarise in the background, send as is]
    B -- over the model's budget --> E[Summarise first, then send]
    D --> F[Checkpoint stored on the thread]
    E --> F
    F --> G[Later turns start from the checkpoint]
```

## Turn it on for an agent

1. Open the agent and select **System Prompt** settings.
2. Check **Compact long conversations**.
3. Optionally set **Compact at (% of the context budget)** — between 50 and 95, default 85 — and the
   **Summary style**.
4. Click **Update**.

| Setting | Meaning |
| - | - |
| Compact at | When a request reaches this share of the model's input budget, the older turns are summarised in the background and the next turn uses the summary. |
| Summary style | **Briefing** (recommended) rewrites one running brief each time. **Gist** keeps a chain of structured notes. **Extractive** quotes the most relevant passages and makes no model call. |

The input budget is the deployment's context window, minus room for the reply (the agent's maximum
output tokens, or its reasoning budget if larger) and a small safety margin. For a deployment on a
Bud cluster the window is the length the model was actually served with, which is often smaller
than the model's advertised maximum.

## Turn it on per request

A `/v1/responses` request can turn compaction on for that call with OpenAI's `context_management`
field, whether or not the agent has it on:

```bash theme={null}
curl https://<gateway>/v1/responses \
  -H "Authorization: Bearer $BUD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": {"id": "my-agent"},
    "previous_response_id": "resp_...",
    "input": "And what about the refund?",
    "context_management": [{"type": "compaction", "compact_threshold": 24000}]
  }'
```

`compact_threshold` is in tokens and is clamped to 50–95% of the input budget. A bare object
(`{"type": "compaction"}`) is accepted too.

## What you see

* The run view shows an **Earlier turns summarised** step with the checkpoint's generation and the
  token counts before and after.
* The summary call appears in the trace as a child call tagged `bud.call.purpose = compaction`.
* If a conversation still does not fit (for example compaction is off), the request fails with
  `400 context_length_exceeded` instead of a server error.

## Limits

* Compaction applies to stored conversations (`store: true`).
* A deployment whose context window Bud cannot determine is never compacted.
* The first turn to pass the threshold is sent unchanged; the summary is used from the next turn.
  A turn that would not fit at all is summarised before it is sent.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.