Skip to content

dsh-memory

Verified

@try-works/dsh-memory Β· v0.1.16 Β· MIT Β· Web UI

Git-backed, workspace-scoped memory for the DeepSeek Harness (global layer + per-workspace layers)

Install

dsh plugin add @try-works/dsh-memory

Confirm the layer applied with dsh --profile default --dump-config β€” see the install guide.

Source

Published to npm without a public repository. Inspect the package contents before installing.

Tags

Readme

dsh-memory

Git-backed, workspace-scoped memory for the DeepSeek Harness.

Where things stand: STATUS.md is the tracker and planner β€” what is verified working, what is written but not yet loaded, what is unverified, and what is left. Updated against measurements rather than intentions, so if you want to know whether a capability is real, read that rather than this.

Features

  • Global layer (global/) + per-workspace layers (workspaces/<key>/).
  • Workspaces auto-associate with local folders via DSH ctx.workspaceRegistry.
  • The whole memory folder is a git repo: durable markdown is committed, derived state is gitignored. GitHub sync (pull + push) runs hourly (+ boot + explicit).
  • Web-UI settings (storage root, GitHub origin, token, cadences) via ctx.settings.
  • Deep DSH hooks: session/event capture, agent/pre-step injection + cadence, in-process ctx.llm consolidation, ctx.subprocess git, ctx.sessionQuery capture, ctx.goals.

Self-describing: SYSTEM.md

The store explains itself. On every boot the plugin renders SYSTEM.md at the store root from src/SYSTEM.template.md, filling in the live configuration and a table of what the store currently holds. It is committed, so the same description travels with the memory to every machine β€” and regenerated every boot, so it can never describe a version of the plugin that no longer exists.

It documents the intent, not just the mechanics: why the files are the truth and everything else is derived, why a model decides what to store but never how it is written, why each layer exists and what belongs in it, what the reply protocol is and what the earlier JSON contract got wrong, and how the whole thing is meant to be used. An agent that reads it can place a fact in the right layer and know what it is safe to trust.

The document version is stamped into the file and exposed as systemDocStale(root), so a stale copy is detectable rather than silently wrong.

Store layout

<memoryRoot>/
β”œβ”€β”€ .gitignore            # ignores index/ state/ digests/ logs/
β”œβ”€β”€ memory.json           # schemaVersion + git.remote/branch + workspace registry
β”œβ”€β”€ global/               # GLOBAL layer (identity, prefs, standing rules)
β”‚   β”œβ”€β”€ INDEX.md          # hand-written topic map over the whole scope
β”‚   β”œβ”€β”€ semantic.md       # INDEX of semantic topics
β”‚   β”œβ”€β”€ semantic/         # one file per topic
β”‚   β”‚   └── <topic>.md
β”‚   β”œβ”€β”€ episodic.md  + episodic/<topic>.md
β”‚   β”œβ”€β”€ procedural.md + procedural/<topic>.md
β”‚   β”œβ”€β”€ learnings.md  + learnings/<topic>.md
β”‚   β”œβ”€β”€ working.md        # current-state head (single file, never indexed)
β”‚   └── junk.md           # archaeology (single file, never indexed)
└── workspaces/<key>/     # WORKSPACE layer (one per associated local folder/repo)
    β”œβ”€β”€ meta.json
    β”œβ”€β”€ codebase.md  + codebase/<topic>.md   # repository map (WORKSPACE ONLY)
    └── (same shape as global/)

Layered layers

semantic, episodic, procedural, learnings and codebase are multi-layered: the layer file is an index, and the entries live in one file per topic under a sibling directory. A single append-only file per layer grows without bound and becomes hard to read, diff and grep; one file per topic keeps every file small and searchable.

# global/semantic.md  β€” the index

Topic index. Entries live in one file per topic under semantic/.

- [Store root](semantic/store-root.md) β€” 1 entry, updated 2026-09-30
- [Plugin config rows](semantic/plugin-config-rows.md) β€” 2 entries, updated 2026-10-01

Properties this buys:

  • The index is regenerated from the topic files after every write, so it can never drift.
  • A write reindexes only the files it touched (updateIndex). A full rebuild re-reads and re-tokenizes the whole store and costs seconds once the store holds a few hundred entries; buildIndex remains the repair path.
  • A topic file is ordinary Format-2 markdown, so parseEntries, the ledger, search and the context card work on it unchanged β€” the layer is resolved from the path, not the basename.
  • Id allocation spans the whole layer (NNNN is unique per scope+layer, not per file).
  • Adding a topic file by hand is fine: the next rebuild picks it up.
  • Existing stores migrate on load: entries still in a single layer file are moved into topic files and the index is rebuilt. The migration is idempotent.

Install

The package is a self-contained DSH plugin: plain ESM, no build step, and a bundle patch that mounts it by name. Install it from a path or a registry:

# from a local checkout
dsh plugin --profile web add D:/DEV/dsh-memory

# or from a registry, once published
dsh plugin --profile web add dsh-memory

Adding it composes cordis.patch.yml into the profile bundles, which mounts one entry:

- insert:
    - id: dsh-memory
      name: dsh-memory

The entry carries no config on purpose. The store root, memory model, git origin and branch belong to the user, so they are written by Settings -> Memory as a separate top-level config: row for the same id in your profile patch. Mounting by hand works the same way:

- insert:
    - id: dsh-memory
      name: D:/DEV/dsh-memory/src/index.js
- id: dsh-memory          # written by the settings page
  config:
    memoryRoot: E:\dsh-memory

The entry id is the settings namespace, so either memory or dsh-memory works as a row id; the client half resolves whichever the host actually serves.

Package invariants

npm run check (or node scripts/check-package.mjs) verifies what an install-by-name needs: the entry exports name/inject/apply/Config and no default export, the bundle patch mounts the package by name, the client half and SYSTEM.template.md ship, and every imported peer is declared. test/package.test.js asserts the same invariants so packaging cannot drift.

Configure (web UI β†’ Settings β†’ Memory)

The plugin ships a browser client half (src/client/index.js, declared via dsh.client in package.json), so Settings shows a Memory page with editable fields. Saving writes to this profile's layer and applies live.

The client half is plain JavaScript that requests only react from the browser module table. It deliberately imports no harness client package: those change without notice, plain JS gets no type check against them, and a throwing component blanks the slot entry. Its styles are local, class-prefixed, and built only from theme tokens that exist on the shipped settings pages (--dsw-alias-*, --dsw-radius-*, --ds-font-family-code). A token that does not exist is an invalid declaration, the browser drops it, and the page degrades into unstyled prose β€” which is exactly what happened before the current styling was added.

Settings -> Memory

Field Purpose
memoryRoot local folder/drive for the store (default ~/.dsh-memory)
summaryModel model for memory writes (provider/model); empty = the session's own model
gitRemote GitHub origin URL (empty = local-only)
gitBranch default main

Config file only (advanced tuning)

Set these in the profile patch's config: block. They have sane defaults and are deliberately kept out of the settings page.

Field Purpose
gitToken optional PAT (secret; empty = OS git credential helper)
gitName / gitEmail optional commit identity
autoCaptureIntervalMs per-session capture throttle (default 300000)
commitDebounceMs local commit debounce (default 10000)
autoSyncIntervalMs GitHub push+pull cadence (default 3600000 = 1h)
autoConsolidateIntervalMs / autoConsolidateMinPending / autoConsolidateBatchSize consolidation cadence
autoRefreshIntervalMs freshness scan cadence (default 86400000 = 1d)
injectionMaxChars / injectionRefreshMs injection budget
digestMaxUserTurns / digestMaxAssistantGists / digestMaxMessageChars what the summarizer may see per pass
digestMaxBytes hard ceiling on summarizer input per pass
summaryMaxBullets durable bullets kept from one summary
consolidateMaxPromptBytes / consolidateMaxDigestBytes consolidation prompt + per-digest ceilings

Agent tools

  • memory_search(query, scope?, workspace?, limit?) β€” ranked facts with file:line citations.
  • memory_remember(text, layer, scope?, workspace?, topic?, source?, importance?, until?) β€” deterministic write.
  • memory_workspaces() / memory_associate(key, title?, path?) β€” workspace association.
  • memory_sync() β€” commit local changes and sync (pull + push) with the GitHub origin.
  • memory_forget(id, undo?) β€” forget a memory entry (supersede + suppress re-adding), or release it.
  • memory_consolidate(limit?) β€” run the LLM consolidation pass now (digests β†’ entries) instead of waiting for the cadence.
  • memory_codebase(topic, text, importance?) β€” record what a module or subsystem of the current repository does (workspace codebase map).
  • memory_decision(decision, rationale, context?, alternatives?, consequences?, scope?, workspace?, importance?) β€” record a settled choice with why it won. Available in both scopes.
  • memory_state(section, content) β€” record or refresh one section of the repository state overview.
  • memory_working(section?, content?, activity?, clearActivity?) β€” record what you are working on: a current section, or activity notes.
  • memory_status() β€” where the store lives, how many digests await consolidation, whether the last sync conflicted.
  • memory_refresh(limit?) β€” run the freshness pass now instead of waiting for the daily cadence.
  • memory_status() β€” now also reports model calls, event counts and saved calls (thin/trivial).
  • memory_history(limit?) β€” ingest repository commit history into the workspace history layer.

ctx.memory service: search, remember, workspaces, associate, resolveKey, status, workspaceOf, pending, consolidate.

Capture pipeline

Capture turns sessions into memory without ever storing raw transcripts:

  1. Relevance filter (digest.js) β€” deterministic and bounded. Tool calls reduce to a count (tool mechanics are not memory); user turns and assistant gists are clipped; bare acknowledgements are dropped; the window is trimmed oldest-first to digestMaxBytes.
  2. Summarization (summarize.js) β€” the session's own model turns that material into 5-12 durable bullets under explicit instructions: bullets, **label:**, date-led when time-bound, keep decisions/constraints/preferences/environment facts, drop tool chatter, code and pleasantries. A model that finds nothing durable replies NO_DURABLE_CONTENT and nothing is stored.
  3. Incremental storage (capture.js) β€” a lastSeq watermark in state/sessions.json means each pass summarizes only events that arrived since the previous one, so capture is O(new events) rather than O(session). Bullets are appended as a dated section.

A digest file therefore holds a bounded metadata header plus dated bullet sections β€” never message bodies. Without a model, capture degrades to deterministic clipped notes instead of failing; a null consolidatedAt marks the digest as work for the consolidation pass.

Working memory and sync safety

  • working.md is per repo/workspace, never global: it is the current-state head of one repository, so mixing repositories would be meaningless. The global layer has no working file, and a template-only global/working.md left by an older version is retired on boot.
  • A failed git pull/push is never swallowed. The failure is logged, recorded in state/sync.json, surfaced by status(), reported by the memory_sync tool, and leaves a state/CONFLICT.md marker that the next clean sync clears.
  • Consolidation attributes every entry to the digest it came from. A digest the model never attributes stays pending and is retried; after maxAttempts it is marked needsReview instead of being silently dropped.
  • A model reply is parsed with escalating recovery: exact JSON, then repaired JSON (trailing commas, placeholder values), then whatever complete objects the text contains β€” so a truncated or sloppy reply degrades instead of failing the pass. If nothing is recoverable, the raw prompt and reply are written to logs/consolidate-failed-<stamp>.txt and the tool reports parseFailed rather than discarding the evidence.
  • Replies are assembled from text blocks only. A reasoning model streams its thinking as reasoning blocks; folding those into the reply prepends prose to the answer, and any "parse the JSON" step then matches the wrong braces. This was the cause of a real failure: the reasoning preamble plus the echoed prompt example were handed to the parser.
  • The reply protocol is line-oriented, not JSON: one - [<digest> | global|<workspace:key> | <layer> | <importance>] <topic> :: <text> line per fact, plus a DONE <digest> line per digest processed. Nested JSON proved fragile in practice β€” one invalid placeholder, one truncated brace or one reasoning preamble lost the entire pass β€” while a line protocol parses per line, so one bad line cannot kill the batch and prose is simply ignored. A JSON object is still accepted if a model volunteers one, with escalating recovery (exact, repaired, then salvaged objects).
  • Nothing recoverable is discarded: the raw prompt and reply are written to logs/consolidate-failed-<stamp>.txt and the tool reports parseFailed.
  • Each digest heading carries the workspace key resolved from that session's own working directory, and the model is told to use it. It was previously left to derive a key, which it did from the filesystem path β€” and a path key created a workspace directory that the write then could not use. remember() now sanitizes the key once, up front, so a key and the directory it targets can never disagree.
  • The last consolidation exchange is always kept at logs/consolidate-last.txt, and the tool reports writeFailed with the reasons, so a pass that writes nothing is visible.
  • One malformed entry cannot fail a whole pass: it is recorded, and a workspace request with no usable key falls back to the global layer instead of throwing.

Codebase map

codebase is the repository map: one entry per module, file or subsystem, written as the agent discovers it. It is deliberately workspace-only β€” a global codebase map would mix unrelated repositories together, exactly like working memory.

Two things feed it:

  • memory_codebase(topic, text, importance?) β€” the agent records what a module does. The topic is the path or subsystem name, so re-discovery updates the same topic file. It resolves the workspace from the session, so no key is needed.
  • Traversal evidence. Capture records the distinct file paths a session's tool calls point at (arguments only β€” results are full of paths mentioned in file contents, which say nothing about intent), bounded and deduped. Those paths ride the digest, so summarization and consolidation can turn exploration into map entries without the agent writing anything.

The injected context lists the topics already mapped for this workspace, so the agent can see what it can rely on and what is still unexplored.

Commit history

history records how the repository changed, complementing the codebase map (which says what each module is). Also workspace-only and multi-layered, but its topics are calendar days β€” one file per day, so a long-lived repo yields many small readable files:

# workspaces/role-model/history.md
- [2026-09-27](history/2026-09-27.md) β€” 11 entries
- [2026-09-28](history/2026-09-28.md) β€” 16 entries
- [2026-10-01](history/2026-10-01.md) β€” 2 entries

Each entry is an ordinary Format-2 entry (subject + short hash as the heading, the full commit body, the changed files, and **Source:** git:<hash>), so the ledger, ranking and search work on it unchanged.

Ingest is incremental in two layers:

  1. The log range starts at the last recorded commit (<head>..HEAD), so a normal pass reads only new commits. When the range is unreachable (a rebase or force-push rewrote history) the pass falls back to a bounded walk and reports fellBack.
  2. A hash watermark in state/history.json skips what is already recorded β€” an optimisation, not the authority.
  3. The layer itself is the authority. Every entry carries a Commit: <full hash> line and ingest reads them back, so a lost watermark, a fresh clone or a rewound head cannot produce a duplicate. Ingest is idempotent even if state/ is deleted entirely.

Runners report status inconsistently β€” git.js documents code, while the host runner and ctx.subprocess return exitCode β€” so the check reads either. Reading only one made every successful ranged query look like a failure and silently fell back to a full re-read.

Recording happens on a cadence (autoHistoryIntervalMs, default hourly, autoHistoryLimit commits per pass) and on demand via memory_history. A repository that is missing, not a repo, or has no new commits is skipped silently β€” it never interrupts a turn.

Decisions

A decision is neither a fact nor a lesson: it is a settled choice and why it won β€” a dependency, a convention, an architecture call. It exists so a later session does not re-litigate a closed question or silently reverse it.

Unlike codebase and history, it is available in both scopes: a convention can be a global preference, while a dependency choice belongs to the repository that made it.

The shape is the point. The decision statement is the entry heading, so the layer index reads as a list of decisions, and the reasoning lives in fixed sections:

## Use plain ESM with no build step for dsh-memory

**Context:**
This plugin is plain JavaScript loaded directly by path...

**Rationale:**
Plain ESM keeps the shipped source equal to the code under test...

**Alternatives considered:**
- Compile TypeScript or bundle the host half
- Wrap the entry in a default export

**Consequences:**
The plugin entry must keep named exports and must not add a default export...

Fixed sections make decisions comparable and greppable, and parseDecisionBody reads them back so a stored decision can be inspected without re-deriving it from prose. The write path renders this shape β€” a model deciding what is durable never has to produce it.

Decisions in force are surfaced in the injected context next to learnings, because they are the constraints a new session has to respect.

Freshness metadata

Every entry carries its own decay metadata, so staleness is determinable from the file alone β€” no external clock, no global assumption, and it survives a clone because it travels with the text.

Field Meaning
**Volatility:** How fast this decays: stable (years β€” a language, a license, a convention), normal (months β€” how a subsystem works), volatile (days β€” a port, a version in flight).
**Verified:** When the fact was last confirmed true. This is the age basis.
**Valid-until:** A hard expiry for facts that genuinely end.

Freshness is therefore a property of the entry, not a constant applied to a write date. A plugin using plain ESM stays true for years; the port it listens on can change tomorrow β€” and only the writer knows which is which:

same age (40 days), different verdicts
  dsh-memory-sem-0029  Volatility: volatile  ->  stale  (window 3d)
  dsh-memory-sem-0028  Volatility: stable    ->  fresh  (window 365d)

**Verified:** matters because writing and confirming are different acts: an entry dated 400 days ago but re-confirmed yesterday is fresh, and status quo previously could not express that. Where volatility is absent (entries written before the field existed) time-bounded language still hints at it, then the normal window applies. memory_remember accepts volatility, and the consolidation protocol carries it as an optional fifth field.

Workspace state

state.md is a short, section-based overview of what the repository is right now β€” layout, what is running, what is pending, health. It is the one layer whose entire value is being current, so it is treated differently from every other:

  • Replaced per section, never appended. State describes now, so a newer statement about a section supersedes the old one instead of piling up beside it. Other sections keep their own timestamps β€” editing one area must not make an unrelated section look freshly verified.
  • Every section carries **Verified:**, and freshness is graded rather than binary: fresh / aging / stale / unverified.
  • Staleness is always shown, never hidden. A stale section appears in the injected context with its age ([STALE 5d]), so it reads as a hypothesis to check rather than a fact to trust. An unverified section is marked [UNVERIFIED].
  • Workspace-only and a single file: a repository state file in the global layer would describe no repository, and an overview that needs an index is not an overview.

What capture keeps (and what it refuses)

A session is projected to candidate material before any model sees it. Two rules were added after observing real research sessions silently produce nothing:

  • Injected context is not user intent. The harness prepends runtime snapshots, skill catalogs, policy notices and the memory plugin's own context block as user turns. They are dropped, because in one measured session they were 82% of the material. Capturing the memory block specifically is a feedback loop: the system treating its own output as new input.
  • A research answer is its findings, not its status line. Analysis sessions open with a one-line announcement ("Analysis complete.") and put everything in the body. Capture now skips a bare status opener and carries the substance, and takes the assistant text block rather than its private reasoning.

Both were invisible in unit tests and obvious the moment one real session was traced.

Memory belongs to the main agent

A subagent does not write memory. It reads freely - it can search and consult the store - but recording is the main agent's job.

The reason is ownership, not permission. A subagent is a delegate working inside a task the parent already owns, and the parent's transcript carries the delegation and its result. Letting the child write as well produces two partial accounts of one task, and the child's - written without the parent's context - competes with the parent's in search.

Enforced in three places:

  • Capture skips a subagent session entirely, so it never becomes a digest.
  • Every writing tool refuses with an explanation naming the tool; reads are unrestricted.
  • The injected context tells a subagent to report its findings to the caller instead of instructing it to write.

A session counts as a subagent when delegationDepth > 0 or origin === "subagent" - both are checked because they are set independently, and an absent depth means top-level.

Which model does memory work

Memory work needs a model that returns text. Reasoning-heavy models sometimes stream only reasoning blocks and no answer at all - measured here as 7 of 11 consolidation passes writing nothing, on the same prompts that another model turned into 35 entries.

So the model is not configured as a constant. Each call builds a ladder from the live registry:

  1. the configured summaryModel, when set (a preference, not a requirement)
  2. the session's own route
  3. every other model this deployment can route to right now

The first model that returns text wins. A model that returns nothing, or throws, is skipped and the same prompt goes to the next rung. The winner is remembered, so the normal case still costs one call; a model that returned nothing is benched for 15 minutes and sinks to the end of the ladder - never dropped, so a deployment where everything is benched still tries.

The ladder is capped at 4 models, and if every one returns nothing the call raises an error rather than reporting success, so the session is retried instead of silently settled.

Failures retry; conclusions do not

A summary that fails (empty reply, unparseable prose, or a thrown call) is not the same as a model concluding there is nothing durable. Failure keeps its material and is retried up to summaryMaxAttempts (default 2) before being laid to rest as failed. A deliberate NO_DURABLE_CONTENT settles immediately. Treating a failure as a conclusion is how whole sessions were marked handled and lost for good.

Memory may correctly write nothing

A worker is not required to add a memory every capture interval or every turn. In fact, writing something merely to fill a gap is a failure: it pollutes search, creates false context, and makes later decisions harder to trust.

This is enforced at multiple levels:

  • The injected footer says explicitly that writing NOTHING is a correct and expected outcome.
  • The capture relevance filter drops tool mechanics, acknowledgements and low-signal material.
  • summaryMinSignalChars (default 400) skips the model on a thin new delta entirely.
  • The summarizer may return NO_DURABLE_CONTENT without creating pending work.
  • Consolidation may return only DONE <digest> lines: a handled digest with zero facts is successful, writes zero entries, and is not retried. It never invents a fact for output.
  • memory_remember itself says not to call it to fill a gap.

The desired ratio is not "one turn in, one memory out". It is high signal, low noise, with a cheap no-op path for the common case where nothing durable changed.

Working memory belongs to the agent

Two files describe "now", and they are not the same thing:

File Owner Answers
state.md the repository What is this repo right now? β€” layout, what is running, pending, health
working.md the agent What am I in the middle of? β€” task, recent activity, next steps

A repository has one current state, but several agents may work in it with entirely different intents β€” and one intent may span several repositories. Merging them would lose both meanings.

working.md states its own ownership in the file, on every write, so an agent reading it directly (not just through the injected card) knows it is theirs:

THIS FILE BELONGS TO THE AGENT. It is not repository state (that is state.md) and not
durable memory (those are the layered entries). It is your own head for this workspace.

You own it. Read it, edit it, and keep it honest. Write it with memory_working, or edit
this file directly. Nothing else writes here, so a stale or empty file means the intent
was never recorded.

Two section behaviours, because the information has two shapes:

  • Task / Next steps / Plan / Blocked by β€” replaced on write: there is one current task.
  • Recent activity β€” a rolling log, newest last, trimmed oldest-first by entries and characters, with identical consecutive notes not repeated.

It decays in hours, not days, so it is graded on its own scale (fresh / aging / stale / not started), and an unrecorded file is called out in the injected card rather than omitted β€” silence would read as "no working memory" when it actually means "nothing written yet".

Operational observability

The derived logs/ and state/ surfaces are intentionally observable even though they are gitignored:

  • logs/memory.log is newline-delimited JSON for capture, consolidation, refresh, history, pruning, sync errors and model calls. It rotates at 2 MB.
  • logs/consolidate-last.txt preserves the latest prompt/reply exchange for diagnosis, and consolidate-failed-*.txt preserves parse failures rather than discarding the evidence.
  • state/metrics.json counts cadence events, model calls by purpose, and calls avoided by cost control (thin, trivial). memory_status exposes the useful subset.
  • state/activity.jsonl is a bounded human-readable operational feed.

Cost control

Every stage can decide not to spend a model call:

  • A trivial session (a ping) never reaches a model at all.
  • A content session whose new delta is below summaryMinSignalChars (default 400) is skipped as thin. The check counts the filtered candidates, not raw events, so a burst of tool noise cannot masquerade as content β€” a model call on a two-line delta costs more than it captures.
  • Consolidation batches and caps its prompt (consolidateMaxPromptBytes, consolidateMaxDigestBytes).

Captured digests are pruned on the idle cadence down to the newest digestKeepCount (default 50) consolidated ones. A pending digest is never removed: it is the only copy of work not yet in the store. Ordering is by when the session was captured, not file mtime, so digests written in the same instant cannot be pruned in the wrong order.

Hooks

Direction Hook Effect
read agent/pre-step injects the ## DSH memory context card at step 1 (re-injected per injectionRefreshMs)
write session/event turn/end incremental capture: new events β†’ relevance filter β†’ session-model summary β†’ digests/ (throttled by autoCaptureIntervalMs)
write agent/pre-step cadence LLM consolidation: pending digests β†’ entries
write agent/pre-step cadence freshness scan + LLM refresh (keep/update/supersede)
life session/event turn/start workspace auto-association via ctx.workspaceRegistry
sync agent/pre-step cadence debounced git commit + hourly pull/push
all idle ticker (60 s) runs the cadences even when nobody is chatting; needs no agent when summaryModel is set

On-demand equivalents: memory_consolidate (or ctx.memory.consolidate(agent)), memory_sync (or ctx.memory sync).

How it works

observe (session/event) β†’ capture (sessionQuery β†’ deterministic digest) β†’ consolidate (ctx.llm β†’ entries) β†’ index (FTS5) β†’ query (in-process search) β†’ maintain (freshness/forget) β†’ sync (git).

Development

node --test            # 72 tests (unit + behavior fixtures + eval)
  • src/fixture.js β€” versioned behavior-fixture suite (ranking/freshness/forgetting/budget/workspace-key).
  • src/eval.js β€” sufficiency eval: fact-recall probes, a judge layer, and deterministic anchor checks.

Local validation against an installed DSH (needs the profile node_modules on the resolve path):

node --input-type=module -e "import('file:///D:/DEV/dsh-memory/src/index.js').then(m => console.log(m.name, m.inject))"

Search

Normal queries use the unicode61 word index. CJK and other no-space scripts route to a trigram substring index; two-character queries (which are shorter than a trigram) use an escaped LIKE fallback. Partial English words also retry through substring matching. The parallel index is regenerated or updated alongside the word index, and scope filtering still applies to every result.

Optional rerank

memory_search can ask the session model to reorder hits by usefulness. It is off by default (rerankSearch), costs one call, and is strictly advisory: rerankHits returns the original deterministic order whenever the model is missing, throws, or replies with something unusable β€” and it refuses a partial answer that merely agrees with the existing order. Search never depends on a model.

Settings safety

The Memory page exposes only store root, summary model, GitHub origin and branch. The optional git token is a password field that is never pre-filled from the redacted host snapshot: typing a new token replaces it; leaving the field empty leaves the current token untouched.

Status

Implemented greenfield, no legacy code. Seven layers (semantic, episodic, procedural, learning, plus workspace-only codebase, history and state), all multi-layered except state.md and working.md. Eleven tools, a ctx.memory service for other plugins, deep DSH hooks (capture, consolidation, refresh, history ingest, git sync, idle ticker) and a self-describing SYSTEM.md. 170+ unit tests, a behaviour-fixture suite over five axes, and a wiring test that mounts the real host and executes every tool. See dsh-memory-v2-plan.md for the plan and SYSTEM.md in the store for the as-built description.