dsh-local-vision
已验证dsh-local-vision · v0.2.1 · MIT
DeepSeek Harness plugin: a model-callable local_vision tool that describes images with a local vision model through Ollama. No cloud, the image never leaves the machine. Two tiers: fast (text/colors/summary) and detailed (full description).
安装
dsh plugin add dsh-local-vision 用 dsh --profile default --dump-config 确认 layer 已生效 —— 参见安装指南。
源码
标签
作者
说明文档
dsh-local-vision
Give a text-only DeepSeek Harness agent local "eyes". This plugin registers a
model-callable local_vision tool that describes an image using a local
vision model through Ollama. The image never leaves your
machine — no cloud, no uploads, no API keys.
Requirements
Ollama running (default
http://localhost:11434).At least one vision-capable model pulled, e.g.:
ollama pull qwen2.5vl:3b # small, fast — text / colors / summary ollama pull gemma4:12b-it-qat # larger — thorough description
Install
# from npm (recommended)
npm install -g @deepseek-ai/dsh
dsh plugin --profile web add dsh-local-vision
# or from this repository (use `file:` so @deepseek-ai/dsh-tools resolves)
dsh plugin --profile web add file:/path/to/deepseek-harness-plugins/dsh-local-vision
Restart the DSH session so the new bundle is composed. The package then appears
in Settings → Plugins → Plugin list and the tool local_vision is available.
Usage
local_vision(image: "/abs/path/screenshot.png", mode: "fast")
local_vision(image: "/abs/path/mockup.png", mode: "detailed", prompt: "Describe the navigation and buttons")
local_vision(image: "screen", mode: "detailed") # capture the current display (macOS screencapture)
Modes
mode |
intent | default model |
|---|---|---|
fast |
transcribe visible text, name dominant colors, one-line summary | qwen2.5vl:3b |
detailed |
full factual description: layout, UI elements, all text, colors, state | gemma4:12b-it-qat |
Parameters
| Param | Type | Notes |
|---|---|---|
image |
string (required) | Absolute path to a PNG/JPEG/WebP/GIF, or the literal screen. |
mode |
"fast" | "detailed" |
Defaults to fast. |
prompt |
string | Optional focused question. Defaults per mode. |
model |
string | Optional override of the Ollama model id for this call. |
Configuration
Optional, at $DSH_HOME/local-vision.json (default ~/.dsh/local-vision.json):
{
"host": "http://localhost:11434",
"models": { "fast": "qwen2.5vl:3b", "detailed": "gemma4:12b-it-qat" },
"temperature": 0.1,
"timeoutMs": 300000
}
All keys are optional; defaults are shown above. The plugin sends the image to
Ollama's OpenAI-compatible /v1/chat/completions endpoint as a base64 data URI.
Notes & limitations
- The first call after the model is cold loads it into memory (a 7 GB model takes a while); subsequent calls are fast.
image: "screen"requires macOS and Screen Recording permission for the terminal.- The tool calls Ollama directly (not through the DSH sandbox executor).