Skip to content

dsh-vision

Verified

@gitawego/dsh-vision Β· v0.2.1 Β· MIT Β· Web UI

Capability-aware vision + paste extension for DeepSeek Harness: describe_image tool, cache/retry/fallback delegation, paste markers, audit, local-only, /vision config, Web client (data-driven settings + tool cards).

Install

dsh plugin add @gitawego/dsh-vision

Confirm the layer applied with dsh --profile default --dump-config β€” see the install guide.

Source

Published to npm without a public repository. Inspect the package contents before installing.

Creators

Readme

dsh-vision

Capability-aware vision + paste extension for DeepSeek Harness, ported 1:1 from @gitawego/pi-vision. Published on npm as @gitawego/dsh-vision.

  • describe_image tool (single + batch image_paths) β€” hidden on multimodal primaries (native pass-through), visible on text-only primaries (delegation with cache/retry/fallback).
  • Paste UX: [Image-#N] markers, hint/auto/off modes.
  • Cache, audit log, local-only mode, batch concurrency, auto-detect, /vision command, Web Settings (data-driven, no hardcoded provider/model ids).

Installation

Install the plugin into a profile (the harness composes plugins per profile; the web profile backs the Web GUI). Pick one:

1. npm registry (recommended)

The package is published; add it from the registry and let the harness reconcile the bundle list:

dsh plugin --profile <name> add @gitawego/dsh-vision
dsh plugin --profile <name> install

Verify it composed, then restart the harness so the plugin (and its Web client) load:

dsh --profile <name> --dump-config   # expect a row for @gitawego/dsh-vision
dsh --profile <name>                 # restart, e.g. dsh web

2. Direct install (local path / no registry)

Useful for development or offline hosts: depend on a local checkout with a file: spec (exactly what pnpm would write for pnpm add file:...) and list the package in the profile's bundles, then run the official install command:

  1. Edit ~/.dsh/profiles/<name>/package.json:
{
  "dependencies": {
    "@gitawego/dsh-vision": "file:/absolute/path/to/dsh-vision"
  },
  "dsh": {
    "profile": {
      "bundles": [
        "@deepseek-ai/dsh-base",
        "@deepseek-ai/dsh-web-app",
        "@gitawego/dsh-vision"
      ]
    }
  }
}
  1. Install + verify + restart:
dsh plugin --profile <name> install
dsh --profile <name> --dump-config   # expect a row for @gitawego/dsh-vision
dsh --profile <name>                 # restart

Note: after rebuilding a local checkout, re-run dsh plugin --profile <name> install (pnpm copies file: dependencies into its store β€” not a live link).

Configuration

Configure the vision provider/model and behavior in the Settings β†’ Vision page, or with the /vision command (/vision show, /vision config provider <id>, ...).

  • provider / model β€” an image-capable model from the live harness catalog (data-driven; auto-detect offers a default). Used by the native sub-agent delegation.
  • delegation β€” auto (native sub-agent with the vision model; falls back to the direct http endpoint when the harness cannot deliver images natively), native (sub-agent only), http (plugin-owned direct endpoint call).
  • http block β€” baseUrl, credential (a DSH credential-ref name), model, protocol (openai | anthropic).

Example http configuration (any OpenAI-compatible vision endpoint):

/vision config delegation http
/vision config http.baseUrl https://api.example.com/v1
/vision config http.model vision-model
/vision config http.credential VISION_API_KEY   # a DSH credential ref

Android/Termux note: the harness attachment store cannot write under /data/data (its durability walk hits EACCES β€” a raw Android permission), so the native sub-agent path cannot deliver images there. Use delegation=http with the http block on Termux; auto falls back to http automatically when the store is unavailable. The same rule drives image routing: on Termux even a multimodal primary (e.g. MiniMax-M3) is treated as effectively text-only β€” it cannot receive native ImageBlocks, so describe_image becomes visible and your existing textOnlyPasteMode governs (hint = on-demand delegation via http, auto = automatic, off = markers only). The hint names native delivery as the reason instead of blaming the model, and pasted images are never silently dropped as a bare [Image-#N] marker. On hosts with a working attachment store, multimodal models keep native passthrough.

Routing & model-switch behavior

The plugin routes images by the current session model's capability (detected from the live harness catalog, never hardcoded):

Session primary model Image behavior
Multimodal (e.g. a vision "luna") Native pass-through: pasted path tokens become [Image-#N] markers + ImageBlock attachments; the model sees the raw image. describe_image is hidden (delegation would be wasteful).
Text-only (e.g. glm-5.2) Convert to text: path tokens and GUI image blocks are delegated to the configured vision model (cache/retry/fallback) or surfaced as hint markers for describe_image. The tool is visible.

Mid-session model switches are detected at the earliest point of the next step (system-prompt/assemble, which fires before the paste hook and before tool visibility is frozen), so the first step after a switch already routes correctly β€” paste auto-converts (or attaches natively), and describe_image is shown/hidden accordingly. /vision session-status reports the tracked model, the logged request header, and flags a switch that is still pending detection.

Harness (GUI) constraints on rc.6 β€” the dsh-host-apiproxy wire refuses two operations before this plugin ever sees the message, independent of the plugin:

  • sending a prompt with images while the session model is text-only β†’ attachment-error "Model X does not support image input";
  • switching to a text-only model while the session already contains images β†’ model-unavailable "…but this session already contains images".

So via the GUI, "upload an image while on glm-5.2" and "switch to glm-5.2 in an image session" are blocked by design. The plugin's supported route for those flows is image paths + textOnlyPasteMode: auto (paste the path; the vision model describes it; the text-only model reads the description), or switching models before images enter the session. Image blocks that DO reach the paste hook on a text-only primary (tui/headless/direct API, or future wires) are materialized to hash-named temp files under <DSH_HOME>/tmp/dsh-vision/ (Termux-safe: the OS tmpdir may be unwritable) and converted through the same pipeline β€” a raw image block never reaches a text-only model's request boundary. Auto-converted temp files are removed after delegation; hint-mode files are retained (content-hash-deduped) so the model can name them with describe_image.

Platform support

The harness loads client bundles only for platform: "web" (dsh-client-modules skips every other platform), so the plugin's Web client is available only in the web profile:

Surface web tui / headless
describe_image tool βœ“ βœ“ (host tool)
/vision command βœ“ βœ“ (host command)
Paste markers / auto βœ“ βœ“ (agent/pre-step hook)
Delegation (subagent/http) βœ“ βœ“
Settings page (Vision) βœ“ β€” (web-only slot)
Tool call card βœ“ β€” (web-only slot)

For tui / headless profiles, configure everything with the /vision command (/vision show, /vision config ...) or by editing vision: in ~/.dsh/settings.yaml β€” the settings page is a Web-only convenience, not a requirement.

Development

npm ci
npm run typecheck
npm test
npm run build      # server + client bundles (lib/ incl. lib/client.js)

Publishing is automated via GitHub Actions: tag v* to run typecheck/tests/build and publish @gitawego/dsh-vision to npm with provenance (OIDC trusted publishing), plus a GitHub Release.

Docs

  • SPEC.md β€” full implementation spec (feature parity, architecture, KV-cache requirements).
  • AGENTS.md β€” project architecture and design rules.
  • LESSONS.md β€” session history, debugging deep-dives, tooling pitfalls.