front: configurable compaction model + live progress to the client
Auto-compaction can now summarize with a DIFFERENT model than the one
serving the request, with a global default (config.json `compaction`) and
a per-model override (models.json `auto_compact_model`). Empty = the
request's own model, as before.
- config: new CompactionConfig (enabled/pct/strategy/model) + round-trip
- text.py: resolve effective settings (per-model over global), resolve the
summarizer LAZILY (only when actually over threshold, so a separate model
isn't loaded on every request); map-reduce the dropped history into chunks
sized to the CHOSEN summarizer's own context, reducing iteratively until
it fits; stream status + live per-chunk progress to the client as content
deltas (queue-bridged from the summarizer's callback)
- admin: global compaction card (settings) + per-model summarizer dropdown
(models, shown only for the summarize strategy)
Raw two-pass path is skipped (prompt is built from system + last user turn).
Co-Authored-By:
Claude Opus 4.8 <noreply@anthropic.com>
Showing
This diff is collapsed.
Please
register
or
sign in
to comment