dsh-vllm-effort-sync
已验证dsh-vllm-effort-sync · v0.1.2 · MIT
DSH plugin: read reasoning_efforts from OpenAI-compatible (vLLM) providers' /models listing and sync them into the llm-pi-ai settings section so the Web composer shows a per-model reasoning-effort picker.
安装
dsh plugin add dsh-vllm-effort-sync 用 dsh --profile default --dump-config 确认 layer 已生效 —— 参见安装指南。
源码
发布到 npm 但没有公开仓库。安装前请检查包内容。
标签
作者
说明文档
dsh-vllm-effort-sync
A DeepSeek Harness (DSH) plugin that makes the reasoning‑effort picker show up in the Web composer for OpenAI‑compatible (e.g. self‑hosted vLLM) models — by reading the reasoning_efforts each provider advertises in its GET /models listing and syncing it into the llm-pi-ai settings section.
Why it exists
Your settings.yaml connects a provider like this (via the llm-pi-ai adapter):
llm-pi-ai:
providers:
my:
apiKeyEnv: MY_API_KEY
api: openai-completions
baseURL: http://8.159.147.156:30011/v1
models:
- id: deepseek-v4-flash-vllm
contextWindow: 1048576
A hand‑declared model like deepseek-v4-flash-vllm has no reasoning capability by default. DSH itself never looks at reasoning_efforts in the /models reply (its built‑in discovery only reads id/name/contextWindow/maxTokens), so:
- the Web model selector shows no "推理等级 / reasoning effort" menu, and
- requests don't carry
reasoning_effort.
…even though your vLLM genuinely advertises it:
"reasoning_efforts": { "off": null, "low": "low", "medium": "medium", "high": "high", "max": "max" }
This plugin closes that gap automatically.
What it does
On load (and reactively whenever the llm-pi-ai settings section changes), it:
- Reads the
llm-pi-aisettings namespace. - For every provider with a
baseURLon the OpenAI‑compatible chat‑completions protocol, it resolves the API key (through yourapiKeyEnvcredential) and callsGET <baseURL>/models. - Extracts each model's
reasoning_effortsmap (pi‑ai levels → wire spelling). - Writes it back as a per‑model
reasoningEffortsdeclaration plus thecompatswitches pi‑ai needs to dispatchreasoning_effort(thinkingFormat: deepseek,supportsReasoningEffort: true).
It is additive and idempotent:
- Models the listing reports no efforts for are left untouched.
- A
reasoningEffortsyou already declared manually is never overwritten (manual wins). - Nothing is written when the produced section equals the current one (no settings churn, no loops).
- A provider that's unreachable, 401s, or isn't
openai-completionsis skipped and logged — it never takes the route down.
Because pi-ai re-reads its profiles on every request, the change takes effect without a restart; just refresh the Web page.
Install
The plugin is a DSH bundle, installed into your existing web profile with the dsh CLI (a dsh.cmd ships with your install and is on PATH):
cd /d D:\
dsh plugin --profile web add ./dsh-vllm-effort-sync
This initializes/uses the web profile, links the package into its node_modules, and appends dsh-vllm-effort-sync to the profile's bundle list (because package.json declares dsh.bundle). Verify the layer:
dsh --profile web --dump-config # look for the "# == dsh-vllm-effort-sync" layer
Then (re)start your web app:
node --import tsx/esm apps/cli/src/bin.ts web
Installing from npm (like any third-party bundle)
The package is a publish-ready npm bundle: main points at the compiled lib/index.js, types at lib/index.d.ts, and the @deepseek-ai/* imports are peer dependencies (resolved from the DSH runtime that loads it, exactly how @linxin666/dsh-web-ui-all / dsh-dafeiyu work). To publish:
# give it your own name first (npm already has "dsh-vllm-effort-sync")
npm login
npm publish # runs prepack -> tsc build -> tarball containing lib/ + patch
Then anyone (including you, on any machine) installs by package name:
dsh plugin --profile web add <your-scope>/dsh-vllm-effort-sync
Or pack a local tarball to distribute without a registry:
pnpm pack # -> dsh-vllm-effort-sync-0.1.0.tgz
dsh plugin --profile web add ./dsh-vllm-effort-sync-0.1.0.tgz
Note:
prepack/preparecompilesrc/index.tstolib/withtsc. A full build needs the DSH type sources (this checkout) available at publish time; the shipped tarball already contains the builtlib/, so consumers installing the tarball/npm package never need them.
Optional per‑plugin config lives in the profile's cordis.patch.yml:
- id: vllm-effort-sync
name: dsh-vllm-effort-sync
config:
providers: [my] # only sync this route (default: all with a baseURL)
applyCompat: true # add pi-ai compat switches (default true)
syncOnStart: true # sync at load (default true)
watch: true # re-sync on section changes (default true)
To remove:
dsh plugin --profile web remove dsh-vllm-effort-sync
What you should see
After the first sync pass, $DSH_HOME/settings.yaml gains, under your provider's model:
llm-pi-ai:
providers:
my:
apiKeyEnv: MY_API_KEY
api: openai-completions
baseURL: http://8.159.147.156:30011/v1
compat:
thinkingFormat: deepseek
supportsReasoningEffort: true
models:
- id: deepseek-v4-flash-vllm
contextWindow: 1048576
reasoningEfforts:
off:
low: low
medium: medium
high: high
max: max
Then in the Web UI: reopen the conversation composer, use the model selector (the button that reads "选择模型 / model"), and deepseek-v4-flash-vllm now offers a 推理等级 / reasoning effort submenu with Off / Low / Medium / High / Max, and picking one records reasoningEffort next to the model and sends reasoning_effort on the wire.
Files
src/index.ts— the plugin source (imports only@deepseek-ai/cordis,@deepseek-ai/dsh-credentials,@deepseek-ai/dsh-settings, all peer deps).lib/index.js+lib/index.d.ts— compiled build output (what package/entry/dsh.bundlepoints at).tsconfig.build.json— tsc config (compilessrc/→lib/).cordis.patch.yml— the bundle patch that inserts thevllm-effort-syncrow.package.json— the bundle manifest (dsh.bundle) with build/prepack scripts and peer deps.README.md— this file.
Notes & limitations
- It only inspects routes on the
openai-completionsprotocol (the one vLLM/OpenAI‑compatible gateways speak). Other protocols are left alone. - Wire spellings come straight from the listing, so if your vLLM uses non‑standard effort names, what it advertises is what's written.
- Resolving a key uses
ctx.credentials.resolve(apiKeyEnv)then the process environment, mirroring the pi-ai adapter. - This is an additive settings sync; it does not modify DSH source and needs no Web rebuild.