dsh-model-safety-gate
Đã xác minh@yadsh/dsh-model-safety-gate · v0.3.1 · MIT · Giao diện web
Independent two-layer safety gate around the DeepSeek Harness agent loop: deterministic and model-classifier verdicts for prompts, streamed output, tools, and tool results
Cài đặt
dsh plugin add @yadsh/dsh-model-safety-gate Xác nhận layer đã áp bằng dsh --profile default --dump-config — xem hướng dẫn cài plugin.
Mã nguồn
Phát hành lên npm mà không có repository công khai. Hãy kiểm tra nội dung package trước khi cài.
Thẻ
Readme
@yadsh/dsh-model-safety-gate
Independent defense-in-depth safety gate for DeepSeek Harness: a deterministic scanner plus an isolated small-model classifier check user prompts, streamed model output (text and reasoning), tool calls, and tool results before they reach the model, the user, or the execution layer.
The gate is an additional decision layer. It does not replace the DSH sandbox, the permission system, or approval gates, and it never patches DSH core.
Features
- Input guard (
agent/pre-step): prompts are normalized and scanned before the main-model request; matched prompts can be warned about or blocked outright, with a sanitized reason published to the session. - Output stream guard (
llm/stream):text-deltaandreasoning-deltachannels are checked with rolling windows. In the defaultbufferedmode chunks are quarantined until their window passes, so blocked content never reaches the UI or session; blocking aborts the upstream provider request. - Tool gate (
tools/pre-execute): fully assembled tool calls are checked before execution — allow, defer to the native DSH approval flow, or deny. - Indirect-injection guard (
tools/post-execute): tool results from untrusted sources are scanned; hits raise the turn risk state, which tightens decisions for subsequent sensitive tool calls. - Two classification layers: L0 is a fast local scanner (injection patterns, jailbreak markers, secrets/credential formats, Unicode obfuscation, zero-width characters, suspicious encoding, repeated-payload floods, user patterns). L1 is an isolated safety classifier on a small model — a DSH provider/model pair, an OpenAI-compatible endpoint, or off.
- Monotonic safety merge: deterministic red lines cannot be weakened by the classifier; verdicts only escalate.
- Failure modes: classifier timeout or malformed output follows the
configured failure mode —
closed,open,rules-only(default),ask. - Safety ≠ usefulness: quality verdicts (unclear, spam, low-information) warn by default and never block unless explicitly opted in.
- Sanitized audit: plugin logs and counters record decisions with content hashes, never raw blocked content or secret values (raw logging is opt-in). Audit records include the session id but never enter the Harness session journal.
- Classifier isolation: classifier calls run under a process-local bypass marker, so moderating a generation never recursively moderates the moderator; the classifier has no tools.
Install
Install the published npm package by name:
dsh plugin --profile web add @yadsh/dsh-model-safety-gate
Configuration
All options are optional; defaults are shown.
enabled: true # master switch: false silences every surface at once
mode:
warn # off | audit | warn | enforce — default decision profile
# off scans nothing; audit records findings but never enforces, including
# turn-risk escalation
classifier:
backend: none # none | dsh | openai-compatible
provider: "" # backend: dsh — DSH provider id for the classifier model
model: "" # backend: dsh — model id
baseURL: "" # backend: openai-compatible — endpoint base URL
apiKey: "" # backend: openai-compatible — API key (kept out of logs)
timeoutMs: 3000 # classifier request timeout
maxTokens: 128 # bounded structured response
temperature: 0
failureMode: rules-only # closed | open | rules-only | ask — on timeout/error/malformed
requireLocal: false # true forbids remote (openai-compatible) endpoints
input:
enabled: true # gate user prompts on agent/pre-step
safetyAction: block # allow | warn | block for safety verdicts
qualityAction: warn # allow | warn | block for quality-only verdicts (block = opt-in)
output:
enabled: true # gate main-model streaming output
mode: buffered # observe | interrupt | buffered
text: true # check the visible-answer channel
reasoning: true # check the reasoning channel when the provider streams it
checkEveryChars: 512 # new quarantined chars between classifier snapshots
windowChars: 1536 # snapshot window size sent to the classifier
lookbehindChars: 768 # preceding context included with each window
minCheckIntervalMs: 250
maxBufferedChars: 8192 # overflow fails closed in buffered mode
tools:
enabled: true # gate tool calls on tools/pre-execute
semanticClassifier: true
unanswerableAsk: deny # deny | ask — an escalation the session's approval policy refuses before asking anyone
toolResults:
enabled: true # scan tool results on tools/post-execute
classifyUntrustedSources: true
audit:
enabled: true
includeRawContent: false # opt-in raw content logging (default: hashes only)
ui:
enabled: true
showWarnings: true
allowSessionOverride: true # false forbids per-session downgrade of the global mode
Switching the gate off
enabled: false and mode: off are the same decision spelled twice, and both
mean the gate does nothing at all: no scan, no classifier call, no audit
record, no blocked decision — on the input, streaming-output, tool-call and
tool-result surfaces alike. Switching either one at runtime reaches the very
next check; there is nothing to restart.
Use mode: off when the profile is what changes between environments and
enabled: false when the plugin itself should be inert; audit is the middle
setting for a deployment that wants the findings without the enforcement.
Escalations the deployment cannot answer
A tool call the gate escalates is resolved by DSH, not by the gate: the call
becomes an ask decision, the tool runtime passes it to the approval
service, and the outcome carries no reason — the runtime turns it into its own
sentence. So the refusal a model sees is the user rejected tool "X", even
when nobody was asked, and the rule that actually fired is lost.
That is exactly what a session whose effective approval policy is never
produces: the service returns the refusal before dispatching to any
answerer, deterministically. A QA lockdown pins that policy, so every
escalation there reached the model as a human "no".
tools.unanswerableAsk: deny (default) refuses such an escalation with the
gate's own verdict instead, categories included:
Blocked by dsh-model-safety-gate (unsafe_tool_intent): this call needs
confirmation, but the session's approval policy is "never", so the request
could only ever be refused without asking anyone
The gate reads the policy the same way the approval service does — the
session's logged override first, else the deployment default — and only
refuses when that read completes: a host composing no approval service, or a
policy the gate cannot read, keeps the native ask, which the runtime still
resolves through the real seam. Set ask for a deployment whose own gate
answers asks ahead of that policy (for example a QA surface with
interaction.approvals: interactive, where the operator decides).
Deployment presets
| Profile | Input | Output | Tools | Failure mode |
|---|---|---|---|---|
| Personal | warn | interrupt | ask | rules-only |
| Balanced | block | buffered | ask | rules-only |
| Strict | block | buffered (text + reasoning) | block | closed, session override disabled |
Privacy
If the classifier backend is openai-compatible, prompts, streamed output,
and reasoning content are sent to that endpoint. The settings card states this
in place, next to the endpoint it would use; set classifier.requireLocal: true
to forbid remote endpoints entirely.
classifier.apiKey is declared a secret slot: configuration surfaces receive
only whether a key is configured, and the literal is never returned to a
browser. The safetyGate Remote projection redacts it as well.
Settings card
The plugin's row in the Plugins panel carries the card: opening the row's
configure control edits this plugin's configuration through the live fields of
its own profile entry — the settings namespace of a plugin is its entry id,
dsh-model-safety-gate, and it is also the row id the bundle's patch declares,
which is why moving the card off the old Settings tab changed nothing about
where a stored value lives. Every field the card writes is declared
.volatile() in the config schema, so a committed change re-resolves the running
gate on the spot: turning the gate off stops the next check, and switching the
mode to enforce blocks the next matching prompt.
The Host validates a write against the schema and nothing more, so a
combination the schema cannot express (a dsh backend without a provider, an
uncompilable customBlockPatterns entry) is stored, then refused when the gate
re-resolves it. The gate keeps running its last workable configuration and the
card says so — the refusal is reported by the safetyGate Remote as
configRejected, never swallowed. An operator who wants the change applied has
to correct the field.
The card also reports what the running gate is doing through the safetyGate
Typert Remote — effective mode, whether the classifier is actually wired, the
process-lifetime counters, and the last 50 sanitized verdicts. That view never
carries the classifier key, and it carries content only when raw logging is
switched on.
Chat moderation banners and a per-session shield control are not built yet
(design SPEC Phase 6); ui.* and allowSessionOverride are accepted keys with
no effect today.
Harness API
ctx.safetyGate.inspect(); // effective config, metrics, recent verdicts
The plugin registers one Cordis service (safetyGate) and publishes four
stable plugin-log categories (safety-gate/check|block|warn|classifier-error).
Compatibility
- DeepSeek Harness
>=0.1.7-rc.2 <0.2.0(channelnext), extension points:agent/pre-step,llm/stream,tools/pre-execute,tools/post-execute. - Node.js
^22.19.0 || >=24.0.0. - See compatibility.json for the machine-readable manifest.
Development
pnpm nx run dsh-model-safety-gate:lint
pnpm nx run dsh-model-safety-gate:typecheck
pnpm nx run dsh-model-safety-gate:test
pnpm nx run dsh-model-safety-gate:build
pnpm nx run dsh-model-safety-gate:verify
License
MIT — see LICENSE. Architecture credits are listed in NOTICE.md.