smart-compaction
Verifiedsmart-compaction Β· v0.2.3 Β· MIT
dsh plugin: a compact_now tool letting the model trigger compaction itself at a safe point, plus a context_status tool so it can check real token usage against its context window, instead of only the automatic token-threshold trigger.
Install
dsh plugin add smart-compaction Confirm the layer applied with dsh --profile default --dump-config β see the install guide.
Source
Tags
Creators
Readme
smart-compaction
A DeepSeek Harness (dsh) plugin that gives the
model two tools: compact_now, so it can trigger compaction itself at a point it knows is safe
β right after finishing a step, never mid-edit β instead of only ever being interrupted by dsh's
automatic token-threshold trigger; and context_status, so it can check real, current token
usage against its actual context window instead of guessing whether a conversation "feels long."
Why
dsh's built-in compaction is purely reactive: an automatic trigger fires between agent steps whenever token pressure crosses a fixed threshold, with no idea whether that moment is a good one to interrupt. It can land between two dependent steps β e.g. right after reading a file and right before editing it based on what was just read β and the resulting summary loses exactly the context the very next step needed.
This plugin doesn't change how compaction works, only who decides when. It hands the model a tool it can call proactively, on its own judgment, instead of leaving that decision entirely to a blind token counter.
What it does
Two tools, both no-argument.
compact_now
- Selects the largest currently-compactable span of conversation history β the same
tool-call-pairing-safe boundary logic dsh's own compaction already guarantees, reimplemented
here against dsh's public APIs only (no dsh core changes, no dependency on
compaction-basic's internals). - Calls dsh's own
ctx.compaction.compactRegion()to actually do the compaction β the same summarizer, the same durablecompaction/start/compaction/endlog events, the same guarantees as any other compaction in the session. - No-ops harmlessly (
"Not enough compactable history yet.") if there isn't enough history yet β safe for the model to call speculatively.
context_status
- Read-only β never modifies the session, safe to call anytime.
- Reports current token usage and the model's actual context-window size for this session,
e.g.
"~42,300 / 77,824 tokens used (54.3%)." - Reads dsh's own
contextPressuresession projection (registered by@deepseek-ai/dsh-token-meter, mounted wherevercompaction-basicis) via the publicctx.sessionProjections.stateOf()API β the same numbers the web UI's own context meter reads. Nothing is hardcoded: the context-window figure comes from whatever model this session is actually routed to, so it's correct unchanged across different models, profiles, and context-window sizes. - Exists because, without it, the model has zero visibility into its own context usage β the
only prior signal was a vague "if the conversation feels long" in
compact_now's own description.
How it works
Two entry points exist on dsh's compaction service: compactNow() (what the human /compact
command uses β requires an idle agent, throws busy otherwise) and compactRegion() (what
dsh's own automatic between-step trigger uses β works fine mid-turn). A tool's execute()
always runs mid-turn β that's what calling a tool means β so compact_now is built on
compactRegion(), not compactNow().
Install
From your dsh web profile (web below is the profile name; use whichever profile backs your
session):
dsh plugin --profile web add smart-compaction
This installs the package and adds it to dsh.profile.bundles for you. (No local dsh binary?
Run the equivalent by hand from the profile directory, e.g. ~/.dsh/profiles/web/: pnpm add smart-compaction, then add "smart-compaction" to that package.json's dsh.profile.bundles
array yourself.)
No build step, no config. Restart your dsh service after adding it β new bundles are only picked up on boot.
compact_now requires a compaction service on your profile. Most profile templates ship
one, but not all do (e.g. @deepseek-ai/dsh-web-app-based profiles don't by default). If yours
doesn't, compact_now still installs cleanly (it won't break your profile's boot) but returns an
error every time it's called: "no compaction service is configured on this profile". Add a
compaction-basic bundle to get one.
context_status requires @deepseek-ai/dsh-token-meter mounted (it registers the
contextPressure projection this tool reads). It's normally pulled in wherever compaction-basic
is, so if compact_now works, context_status should too. If it isn't mounted, the tool still
installs cleanly and just reports available: false instead of erroring.
Tell the model when to use it
This is baked into both tools' own descriptions, so it works out of the box with no setup: each
tool's description tells the model to check context_status right after finishing a
self-contained step (e.g. right after todo_write marks an item completed) and to follow up
with compact_now once usage climbs past roughly 70-90%. Since both descriptions are sent to the
model on every request automatically, no AGENTS.md edit is required for this behavior β unlike
an early version of this plugin, which relied entirely on a hand-written AGENTS.md rule and
(measured directly against real session logs) got essentially no organic use as a result: the
description alone wasn't a strong enough signal.
That said, standing instructions carry more weight than a tool description competing against
everything else in a long tool catalog. If you want to reinforce it further, copy
agents-snippet.md into your AGENTS.md (or whatever your profile injects
as standing instructions).
Tying this to todo-completion matters: it's a real, already-tracked signal for "I just finished a
self-contained unit of work," instead of asking the model to estimate its own remaining work or
guess whether a conversation "feels long," both of which it's generally bad at. context_status
replaces that guess with the real number.
The snippet deliberately does not say "compact once usage crosses 70-80%" as the primary rule β a threshold check alone is still reactive: a step that turns out bigger than expected (a large file, a long diff, a subagent dispatch) can blow straight past a comfortable-looking percentage during that step, which is exactly the mid-step interruption this plugin exists to avoid. The percentage is kept only as a hard backstop; the primary check is comparing remaining headroom against the size of the step about to start, and compacting early β before that step, not during or after it β whenever the fit looks tight.
Building from source
Requires Node 22+. Plain JS, no build step β git clone, pnpm install, done.
git clone https://github.com/JoblessJoe/smart-compaction.git
cd smart-compaction
pnpm install
node test.js
test.js is a pure unit-test check of the range-selection logic (select-range.js) and the
context-usage summary logic (context-status.js) β no live dsh/Ollama session required.
Configuration
None. The tool takes no arguments and needs no setup beyond installing it.
Status
compact_now verified end-to-end against a real dsh session, including:
- Basic call: the tool loads, the model calls it, it selects a valid boundary-safe range, and it
drives
compactRegion()mid-turn without ever hitting thebusyfailure this design exists to avoid. - A real tool-call/tool-result pair (
todo_write, markedin_progressthencompleted) sitting in the compacted range β stays correctly paired, nothing split. - Three
compact_nowcalls in a row in one session (one accidental, caught by dsh's own duplicate-call guard mid-task) β each completed cleanly, no crash, nobusy, no corrupted state from operating on a surface that already contains an earlier compaction's checkpoint message.
Troubleshooting: if the tool returns an error like summarization produced no text summary content, that's not this plugin β compact_now selected a valid range and handed it to dsh's own
summarizer, which returned nothing. Observed specifically when the compacted range contains only
injected boilerplate (e.g. standing instructions) with no real assistant-generated text yet β
succeeded reliably once genuine conversation content was in the range. If it happens on ranges with
real content too, check whether your model's "thinking" level is an actually-enforced token budget
or just an instruction (e.g. Ollama's think: low is unenforced β a model can still spend its
entire output budget on hidden reasoning and return no visible summary text). Either way, this is a
model/summarizer-side issue, not something compact_now itself controls.
License
MIT