Skip to content

smart-compaction

Verified

smart-compaction Β· v0.2.3 Β· MIT

dsh plugin: a compact_now tool letting the model trigger compaction itself at a safe point, plus a context_status tool so it can check real token usage against its context window, instead of only the automatic token-threshold trigger.

Install

dsh plugin add smart-compaction

Confirm the layer applied with dsh --profile default --dump-config β€” see the install guide.

Source

Tags

Creators

Readme

smart-compaction

A DeepSeek Harness (dsh) plugin that gives the model two tools: compact_now, so it can trigger compaction itself at a point it knows is safe β€” right after finishing a step, never mid-edit β€” instead of only ever being interrupted by dsh's automatic token-threshold trigger; and context_status, so it can check real, current token usage against its actual context window instead of guessing whether a conversation "feels long."

Why

dsh's built-in compaction is purely reactive: an automatic trigger fires between agent steps whenever token pressure crosses a fixed threshold, with no idea whether that moment is a good one to interrupt. It can land between two dependent steps β€” e.g. right after reading a file and right before editing it based on what was just read β€” and the resulting summary loses exactly the context the very next step needed.

This plugin doesn't change how compaction works, only who decides when. It hands the model a tool it can call proactively, on its own judgment, instead of leaving that decision entirely to a blind token counter.

What it does

Two tools, both no-argument.

compact_now

  • Selects the largest currently-compactable span of conversation history β€” the same tool-call-pairing-safe boundary logic dsh's own compaction already guarantees, reimplemented here against dsh's public APIs only (no dsh core changes, no dependency on compaction-basic's internals).
  • Calls dsh's own ctx.compaction.compactRegion() to actually do the compaction β€” the same summarizer, the same durable compaction/start/compaction/end log events, the same guarantees as any other compaction in the session.
  • No-ops harmlessly ("Not enough compactable history yet.") if there isn't enough history yet β€” safe for the model to call speculatively.

context_status

  • Read-only β€” never modifies the session, safe to call anytime.
  • Reports current token usage and the model's actual context-window size for this session, e.g. "~42,300 / 77,824 tokens used (54.3%)."
  • Reads dsh's own contextPressure session projection (registered by @deepseek-ai/dsh-token-meter, mounted wherever compaction-basic is) via the public ctx.sessionProjections.stateOf() API β€” the same numbers the web UI's own context meter reads. Nothing is hardcoded: the context-window figure comes from whatever model this session is actually routed to, so it's correct unchanged across different models, profiles, and context-window sizes.
  • Exists because, without it, the model has zero visibility into its own context usage β€” the only prior signal was a vague "if the conversation feels long" in compact_now's own description.

How it works

Two entry points exist on dsh's compaction service: compactNow() (what the human /compact command uses β€” requires an idle agent, throws busy otherwise) and compactRegion() (what dsh's own automatic between-step trigger uses β€” works fine mid-turn). A tool's execute() always runs mid-turn β€” that's what calling a tool means β€” so compact_now is built on compactRegion(), not compactNow().

Install

From your dsh web profile (web below is the profile name; use whichever profile backs your session):

dsh plugin --profile web add smart-compaction

This installs the package and adds it to dsh.profile.bundles for you. (No local dsh binary? Run the equivalent by hand from the profile directory, e.g. ~/.dsh/profiles/web/: pnpm add smart-compaction, then add "smart-compaction" to that package.json's dsh.profile.bundles array yourself.)

No build step, no config. Restart your dsh service after adding it β€” new bundles are only picked up on boot.

compact_now requires a compaction service on your profile. Most profile templates ship one, but not all do (e.g. @deepseek-ai/dsh-web-app-based profiles don't by default). If yours doesn't, compact_now still installs cleanly (it won't break your profile's boot) but returns an error every time it's called: "no compaction service is configured on this profile". Add a compaction-basic bundle to get one.

context_status requires @deepseek-ai/dsh-token-meter mounted (it registers the contextPressure projection this tool reads). It's normally pulled in wherever compaction-basic is, so if compact_now works, context_status should too. If it isn't mounted, the tool still installs cleanly and just reports available: false instead of erroring.

Tell the model when to use it

This is baked into both tools' own descriptions, so it works out of the box with no setup: each tool's description tells the model to check context_status right after finishing a self-contained step (e.g. right after todo_write marks an item completed) and to follow up with compact_now once usage climbs past roughly 70-90%. Since both descriptions are sent to the model on every request automatically, no AGENTS.md edit is required for this behavior β€” unlike an early version of this plugin, which relied entirely on a hand-written AGENTS.md rule and (measured directly against real session logs) got essentially no organic use as a result: the description alone wasn't a strong enough signal.

That said, standing instructions carry more weight than a tool description competing against everything else in a long tool catalog. If you want to reinforce it further, copy agents-snippet.md into your AGENTS.md (or whatever your profile injects as standing instructions).

Tying this to todo-completion matters: it's a real, already-tracked signal for "I just finished a self-contained unit of work," instead of asking the model to estimate its own remaining work or guess whether a conversation "feels long," both of which it's generally bad at. context_status replaces that guess with the real number.

The snippet deliberately does not say "compact once usage crosses 70-80%" as the primary rule β€” a threshold check alone is still reactive: a step that turns out bigger than expected (a large file, a long diff, a subagent dispatch) can blow straight past a comfortable-looking percentage during that step, which is exactly the mid-step interruption this plugin exists to avoid. The percentage is kept only as a hard backstop; the primary check is comparing remaining headroom against the size of the step about to start, and compacting early β€” before that step, not during or after it β€” whenever the fit looks tight.

Building from source

Requires Node 22+. Plain JS, no build step β€” git clone, pnpm install, done.

git clone https://github.com/JoblessJoe/smart-compaction.git
cd smart-compaction
pnpm install
node test.js

test.js is a pure unit-test check of the range-selection logic (select-range.js) and the context-usage summary logic (context-status.js) β€” no live dsh/Ollama session required.

Configuration

None. The tool takes no arguments and needs no setup beyond installing it.

Status

compact_now verified end-to-end against a real dsh session, including:

  • Basic call: the tool loads, the model calls it, it selects a valid boundary-safe range, and it drives compactRegion() mid-turn without ever hitting the busy failure this design exists to avoid.
  • A real tool-call/tool-result pair (todo_write, marked in_progress then completed) sitting in the compacted range β€” stays correctly paired, nothing split.
  • Three compact_now calls in a row in one session (one accidental, caught by dsh's own duplicate-call guard mid-task) β€” each completed cleanly, no crash, no busy, no corrupted state from operating on a surface that already contains an earlier compaction's checkpoint message.

Troubleshooting: if the tool returns an error like summarization produced no text summary content, that's not this plugin β€” compact_now selected a valid range and handed it to dsh's own summarizer, which returned nothing. Observed specifically when the compacted range contains only injected boilerplate (e.g. standing instructions) with no real assistant-generated text yet β€” succeeded reliably once genuine conversation content was in the range. If it happens on ranges with real content too, check whether your model's "thinking" level is an actually-enforced token budget or just an instruction (e.g. Ollama's think: low is unenforced β€” a model can still spend its entire output budget on hidden reasoning and return no visible summary text). Either way, this is a model/summarizer-side issue, not something compact_now itself controls.

License

MIT