dsh-vllm-effort-sync
Đã xác minhdsh-vllm-effort-sync · v0.1.2 · MIT
DSH plugin: read reasoning_efforts from OpenAI-compatible (vLLM) providers' /models listing and sync them into the llm-pi-ai settings section so the Web composer shows a per-model reasoning-effort picker.
Cài đặt
dsh plugin add dsh-vllm-effort-sync Xác nhận layer đã áp bằng dsh --profile default --dump-config — xem hướng dẫn cài plugin.
Mã nguồn
Phát hành lên npm mà không có repository công khai. Hãy kiểm tra nội dung package trước khi cài.
Thẻ
Tác giả
Readme
dsh-vllm-effort-sync
A DeepSeek Harness (DSH) plugin that makes the reasoning‑effort picker show up in the Web composer for OpenAI‑compatible (e.g. self‑hosted vLLM) models — by reading the reasoning_efforts each provider advertises in its GET /models listing and syncing it into the llm-pi-ai settings section.
Why it exists
Your settings.yaml connects a provider like this (via the llm-pi-ai adapter):
llm-pi-ai:
providers:
my:
apiKeyEnv: MY_API_KEY
api: openai-completions
baseURL: http://8.159.147.156:30011/v1
models:
- id: deepseek-v4-flash-vllm
contextWindow: 1048576
A hand‑declared model like deepseek-v4-flash-vllm has no reasoning capability by default. DSH itself never looks at reasoning_efforts in the /models reply (its built‑in discovery only reads id/name/contextWindow/maxTokens), so:
- the Web model selector shows no "推理等级 / reasoning effort" menu, and
- requests don't carry
reasoning_effort.
…even though your vLLM genuinely advertises it:
"reasoning_efforts": { "off": null, "low": "low", "medium": "medium", "high": "high", "max": "max" }
This plugin closes that gap automatically.
What it does
On load (and reactively whenever the llm-pi-ai settings section changes), it:
- Reads the
llm-pi-aisettings namespace. - For every provider with a
baseURLon the OpenAI‑compatible chat‑completions protocol, it resolves the API key (through yourapiKeyEnvcredential) and callsGET <baseURL>/models. - Extracts each model's
reasoning_effortsmap (pi‑ai levels → wire spelling). - Writes it back as a per‑model
reasoningEffortsdeclaration plus thecompatswitches pi‑ai needs to dispatchreasoning_effort(thinkingFormat: deepseek,supportsReasoningEffort: true).
It is additive and idempotent:
- Models the listing reports no efforts for are left untouched.
- A
reasoningEffortsyou already declared manually is never overwritten (manual wins). - Nothing is written when the produced section equals the current one (no settings churn, no loops).
- A provider that's unreachable, 401s, or isn't
openai-completionsis skipped and logged — it never takes the route down.
Because pi-ai re-reads its profiles on every request, the change takes effect without a restart; just refresh the Web page.
Install
The plugin is a DSH bundle, installed into your existing web profile with the dsh CLI (a dsh.cmd ships with your install and is on PATH):
cd /d D:\
dsh plugin --profile web add ./dsh-vllm-effort-sync
This initializes/uses the web profile, links the package into its node_modules, and appends dsh-vllm-effort-sync to the profile's bundle list (because package.json declares dsh.bundle). Verify the layer:
dsh --profile web --dump-config # look for the "# == dsh-vllm-effort-sync" layer
Then (re)start your web app:
node --import tsx/esm apps/cli/src/bin.ts web
Installing from npm (like any third-party bundle)
The package is a publish-ready npm bundle: main points at the compiled lib/index.js, types at lib/index.d.ts, and the @deepseek-ai/* imports are peer dependencies (resolved from the DSH runtime that loads it, exactly how @linxin666/dsh-web-ui-all / dsh-dafeiyu work). To publish:
# give it your own name first (npm already has "dsh-vllm-effort-sync")
npm login
npm publish # runs prepack -> tsc build -> tarball containing lib/ + patch
Then anyone (including you, on any machine) installs by package name:
dsh plugin --profile web add <your-scope>/dsh-vllm-effort-sync
Or pack a local tarball to distribute without a registry:
pnpm pack # -> dsh-vllm-effort-sync-0.1.0.tgz
dsh plugin --profile web add ./dsh-vllm-effort-sync-0.1.0.tgz
Note:
prepack/preparecompilesrc/index.tstolib/withtsc. A full build needs the DSH type sources (this checkout) available at publish time; the shipped tarball already contains the builtlib/, so consumers installing the tarball/npm package never need them.
Optional per‑plugin config lives in the profile's cordis.patch.yml:
- id: vllm-effort-sync
name: dsh-vllm-effort-sync
config:
providers: [my] # only sync this route (default: all with a baseURL)
applyCompat: true # add pi-ai compat switches (default true)
syncOnStart: true # sync at load (default true)
watch: true # re-sync on section changes (default true)
To remove:
dsh plugin --profile web remove dsh-vllm-effort-sync
What you should see
After the first sync pass, $DSH_HOME/settings.yaml gains, under your provider's model:
llm-pi-ai:
providers:
my:
apiKeyEnv: MY_API_KEY
api: openai-completions
baseURL: http://8.159.147.156:30011/v1
compat:
thinkingFormat: deepseek
supportsReasoningEffort: true
models:
- id: deepseek-v4-flash-vllm
contextWindow: 1048576
reasoningEfforts:
off:
low: low
medium: medium
high: high
max: max
Then in the Web UI: reopen the conversation composer, use the model selector (the button that reads "选择模型 / model"), and deepseek-v4-flash-vllm now offers a 推理等级 / reasoning effort submenu with Off / Low / Medium / High / Max, and picking one records reasoningEffort next to the model and sends reasoning_effort on the wire.
Files
src/index.ts— the plugin source (imports only@deepseek-ai/cordis,@deepseek-ai/dsh-credentials,@deepseek-ai/dsh-settings, all peer deps).lib/index.js+lib/index.d.ts— compiled build output (what package/entry/dsh.bundlepoints at).tsconfig.build.json— tsc config (compilessrc/→lib/).cordis.patch.yml— the bundle patch that inserts thevllm-effort-syncrow.package.json— the bundle manifest (dsh.bundle) with build/prepack scripts and peer deps.README.md— this file.
Notes & limitations
- It only inspects routes on the
openai-completionsprotocol (the one vLLM/OpenAI‑compatible gateways speak). Other protocols are left alone. - Wire spellings come straight from the listing, so if your vLLM uses non‑standard effort names, what it advertises is what's written.
- Resolving a key uses
ctx.credentials.resolve(apiKeyEnv)then the process environment, mirroring the pi-ai adapter. - This is an additive settings sync; it does not modify DSH source and needs no Web rebuild.