ReqLLM. Compaction
(ReqLLM v1.26.0)
View Source
Context compaction for the OpenAI Responses API (OpenAI and Azure OpenAI).
Compaction asks the service to fold a long conversation into one or more
opaque compaction items that carry the essential prior state in far fewer
tokens. Those items come back as :provider_block content parts on the
assistant message of the returned response and are replayed verbatim by the
Responses encoder on the next request.
Usage
{:ok, first} = ReqLLM.generate_text(model, "Draft a landing page.")
{:ok, compacted} = ReqLLM.compact_context(model, first.context)
next = ReqLLM.Context.append(compacted.context, ReqLLM.Context.user("Add a booking form."))
{:ok, follow_up} = ReqLLM.generate_text(model, next)compacted.context holds the complete returned window in one message. Its
metadata.responses_replay preserves all API items, including retained user
messages, assistant messages, and tool items, in their original order. Do not
prune this window: the next request must replay it as returned by the service.
Append the next user message to continue. The service can also compact a
stored response instead of a replayed context:
{:ok, compacted} = ReqLLM.compact_context(model, nil, previous_response_id: first.id)Server-side compaction on ordinary requests is enabled with the
context_management provider option; the resulting compaction items land
on the response message the same way and replay automatically.
The compacted message keeps the compaction response id under
metadata.compaction_response_id rather than metadata.response_id, so the
next turn replays the compaction items instead of chaining through
previous_response_id.
Summary
Functions
Compacts a conversation through POST /responses/compact.
Same as compact_context/3 but raises on error.
Prepares a compaction response message for replay without pruning its items.
Whether a message carries at least one compaction item.
Whether a content part is a replayable compaction item.
Returns the option schema used for tuple-model defaults.
Drops every message before the most recent one carrying a compaction item.
Functions
@spec compact_context( ReqLLM.model_input(), ReqLLM.Context.t() | ReqLLM.Message.t() | [term()] | String.t() | nil, keyword() ) :: {:ok, ReqLLM.Response.t()} | {:error, term()}
Compacts a conversation through POST /responses/compact.
messages is anything ReqLLM.Context.normalize/2 accepts, or nil when
previous_response_id: names the stored response to compact; passing both
is an error. Returns a ReqLLM.Response whose context contains only the
returned window (see compacted_message/1).
@spec compact_context!( ReqLLM.model_input(), ReqLLM.Context.t() | ReqLLM.Message.t() | [term()] | String.t() | nil, keyword() ) :: ReqLLM.Response.t() | no_return()
Same as compact_context/3 but raises on error.
@spec compacted_message(ReqLLM.Message.t()) :: ReqLLM.Message.t()
Prepares a compaction response message for replay without pruning its items.
The complete output window in metadata.responses_replay is authoritative.
Retained messages and tool items must stay in their original order. The
compaction response ID cannot be used as a normal previous_response_id.
@spec compaction_message?(ReqLLM.Message.t()) :: boolean()
Whether a message carries at least one compaction item.
@spec compaction_part?(ReqLLM.Message.ContentPart.t()) :: boolean()
Whether a content part is a replayable compaction item.
@spec schema() :: NimbleOptions.t()
Returns the option schema used for tuple-model defaults.
@spec trim(ReqLLM.Context.t()) :: ReqLLM.Context.t()
Drops every message before the most recent one carrying a compaction item.
The compaction item carries the context needed to continue, so earlier history only adds request size. Returns the context unchanged when it holds no compaction item.