ReqLLM.Compaction (ReqLLM v1.26.0)

View Source

Context compaction for the OpenAI Responses API (OpenAI and Azure OpenAI).

Compaction asks the service to fold a long conversation into one or more opaque compaction items that carry the essential prior state in far fewer tokens. Those items come back as :provider_block content parts on the assistant message of the returned response and are replayed verbatim by the Responses encoder on the next request.

Usage

{:ok, first} = ReqLLM.generate_text(model, "Draft a landing page.")

{:ok, compacted} = ReqLLM.compact_context(model, first.context)

next = ReqLLM.Context.append(compacted.context, ReqLLM.Context.user("Add a booking form."))
{:ok, follow_up} = ReqLLM.generate_text(model, next)

compacted.context holds the complete returned window in one message. Its metadata.responses_replay preserves all API items, including retained user messages, assistant messages, and tool items, in their original order. Do not prune this window: the next request must replay it as returned by the service. Append the next user message to continue. The service can also compact a stored response instead of a replayed context:

{:ok, compacted} = ReqLLM.compact_context(model, nil, previous_response_id: first.id)

Server-side compaction on ordinary requests is enabled with the context_management provider option; the resulting compaction items land on the response message the same way and replay automatically.

The compacted message keeps the compaction response id under metadata.compaction_response_id rather than metadata.response_id, so the next turn replays the compaction items instead of chaining through previous_response_id.

Summary

Functions

Compacts a conversation through POST /responses/compact.

Prepares a compaction response message for replay without pruning its items.

Whether a message carries at least one compaction item.

Whether a content part is a replayable compaction item.

Returns the option schema used for tuple-model defaults.

Drops every message before the most recent one carrying a compaction item.

Functions

compact_context(model_spec, messages, opts \\ [])

@spec compact_context(
  ReqLLM.model_input(),
  ReqLLM.Context.t() | ReqLLM.Message.t() | [term()] | String.t() | nil,
  keyword()
) :: {:ok, ReqLLM.Response.t()} | {:error, term()}

Compacts a conversation through POST /responses/compact.

messages is anything ReqLLM.Context.normalize/2 accepts, or nil when previous_response_id: names the stored response to compact; passing both is an error. Returns a ReqLLM.Response whose context contains only the returned window (see compacted_message/1).

compact_context!(model_spec, messages, opts \\ [])

@spec compact_context!(
  ReqLLM.model_input(),
  ReqLLM.Context.t() | ReqLLM.Message.t() | [term()] | String.t() | nil,
  keyword()
) :: ReqLLM.Response.t() | no_return()

Same as compact_context/3 but raises on error.

compacted_message(message)

@spec compacted_message(ReqLLM.Message.t()) :: ReqLLM.Message.t()

Prepares a compaction response message for replay without pruning its items.

The complete output window in metadata.responses_replay is authoritative. Retained messages and tool items must stay in their original order. The compaction response ID cannot be used as a normal previous_response_id.

compaction_message?(message)

@spec compaction_message?(ReqLLM.Message.t()) :: boolean()

Whether a message carries at least one compaction item.

compaction_part?(arg1)

@spec compaction_part?(ReqLLM.Message.ContentPart.t()) :: boolean()

Whether a content part is a replayable compaction item.

schema()

@spec schema() :: NimbleOptions.t()

Returns the option schema used for tuple-model defaults.

trim(context)

@spec trim(ReqLLM.Context.t()) :: ReqLLM.Context.t()

Drops every message before the most recent one carrying a compaction item.

The compaction item carries the context needed to continue, so earlier history only adds request size. Returns the context unchanged when it holds no compaction item.