跳到主要内容

dsh-vllm-effort-sync

已验证

dsh-vllm-effort-sync · v0.1.2 · MIT

DSH plugin: read reasoning_efforts from OpenAI-compatible (vLLM) providers' /models listing and sync them into the llm-pi-ai settings section so the Web composer shows a per-model reasoning-effort picker.

安装

dsh plugin add dsh-vllm-effort-sync

用 dsh --profile default --dump-config 确认 layer 已生效 —— 参见安装指南。

源码

发布到 npm 但没有公开仓库。安装前请检查包内容。

标签

作者

说明文档

dsh-vllm-effort-sync

A DeepSeek Harness (DSH) plugin that makes the reasoning‑effort picker show up in the Web composer for OpenAI‑compatible (e.g. self‑hosted vLLM) models — by reading the reasoning_efforts each provider advertises in its GET /models listing and syncing it into the llm-pi-ai settings section.

Why it exists

Your settings.yaml connects a provider like this (via the llm-pi-ai adapter):

llm-pi-ai:
  providers:
    my:
      apiKeyEnv: MY_API_KEY
      api: openai-completions
      baseURL: http://8.159.147.156:30011/v1
      models:
        - id: deepseek-v4-flash-vllm
          contextWindow: 1048576

A hand‑declared model like deepseek-v4-flash-vllm has no reasoning capability by default. DSH itself never looks at reasoning_efforts in the /models reply (its built‑in discovery only reads id/name/contextWindow/maxTokens), so:

  • the Web model selector shows no "推理等级 / reasoning effort" menu, and
  • requests don't carry reasoning_effort.

…even though your vLLM genuinely advertises it:

"reasoning_efforts": { "off": null, "low": "low", "medium": "medium", "high": "high", "max": "max" }

This plugin closes that gap automatically.

What it does

On load (and reactively whenever the llm-pi-ai settings section changes), it:

  1. Reads the llm-pi-ai settings namespace.
  2. For every provider with a baseURL on the OpenAI‑compatible chat‑completions protocol, it resolves the API key (through your apiKeyEnv credential) and calls GET <baseURL>/models.
  3. Extracts each model's reasoning_efforts map (pi‑ai levels → wire spelling).
  4. Writes it back as a per‑model reasoningEfforts declaration plus the compat switches pi‑ai needs to dispatch reasoning_effort (thinkingFormat: deepseek, supportsReasoningEffort: true).

It is additive and idempotent:

  • Models the listing reports no efforts for are left untouched.
  • A reasoningEfforts you already declared manually is never overwritten (manual wins).
  • Nothing is written when the produced section equals the current one (no settings churn, no loops).
  • A provider that's unreachable, 401s, or isn't openai-completions is skipped and logged — it never takes the route down.

Because pi-ai re-reads its profiles on every request, the change takes effect without a restart; just refresh the Web page.

Install

The plugin is a DSH bundle, installed into your existing web profile with the dsh CLI (a dsh.cmd ships with your install and is on PATH):

cd /d D:\
dsh plugin --profile web add ./dsh-vllm-effort-sync

This initializes/uses the web profile, links the package into its node_modules, and appends dsh-vllm-effort-sync to the profile's bundle list (because package.json declares dsh.bundle). Verify the layer:

dsh --profile web --dump-config   # look for the "# == dsh-vllm-effort-sync" layer

Then (re)start your web app:

node --import tsx/esm apps/cli/src/bin.ts web

Installing from npm (like any third-party bundle)

The package is a publish-ready npm bundle: main points at the compiled lib/index.js, types at lib/index.d.ts, and the @deepseek-ai/* imports are peer dependencies (resolved from the DSH runtime that loads it, exactly how @linxin666/dsh-web-ui-all / dsh-dafeiyu work). To publish:

# give it your own name first (npm already has "dsh-vllm-effort-sync")
npm login
npm publish        # runs prepack -> tsc build -> tarball containing lib/ + patch

Then anyone (including you, on any machine) installs by package name:

dsh plugin --profile web add <your-scope>/dsh-vllm-effort-sync

Or pack a local tarball to distribute without a registry:

pnpm pack          # -> dsh-vllm-effort-sync-0.1.0.tgz
dsh plugin --profile web add ./dsh-vllm-effort-sync-0.1.0.tgz

Note: prepack/prepare compile src/index.ts to lib/ with tsc. A full build needs the DSH type sources (this checkout) available at publish time; the shipped tarball already contains the built lib/, so consumers installing the tarball/npm package never need them.

Optional per‑plugin config lives in the profile's cordis.patch.yml:

- id: vllm-effort-sync
  name: dsh-vllm-effort-sync
  config:
    providers: [my]          # only sync this route (default: all with a baseURL)
    applyCompat: true        # add pi-ai compat switches (default true)
    syncOnStart: true        # sync at load (default true)
    watch: true              # re-sync on section changes (default true)

To remove:

dsh plugin --profile web remove dsh-vllm-effort-sync

What you should see

After the first sync pass, $DSH_HOME/settings.yaml gains, under your provider's model:

llm-pi-ai:
  providers:
    my:
      apiKeyEnv: MY_API_KEY
      api: openai-completions
      baseURL: http://8.159.147.156:30011/v1
      compat:
        thinkingFormat: deepseek
        supportsReasoningEffort: true
      models:
        - id: deepseek-v4-flash-vllm
          contextWindow: 1048576
          reasoningEfforts:
            off:
            low: low
            medium: medium
            high: high
            max: max

Then in the Web UI: reopen the conversation composer, use the model selector (the button that reads "选择模型 / model"), and deepseek-v4-flash-vllm now offers a 推理等级 / reasoning effort submenu with Off / Low / Medium / High / Max, and picking one records reasoningEffort next to the model and sends reasoning_effort on the wire.

Files

  • src/index.ts — the plugin source (imports only @deepseek-ai/cordis, @deepseek-ai/dsh-credentials, @deepseek-ai/dsh-settings, all peer deps).
  • lib/index.js + lib/index.d.ts — compiled build output (what package/entry/dsh.bundle points at).
  • tsconfig.build.json — tsc config (compiles src/ → lib/).
  • cordis.patch.yml — the bundle patch that inserts the vllm-effort-sync row.
  • package.json — the bundle manifest (dsh.bundle) with build/prepack scripts and peer deps.
  • README.md — this file.

Notes & limitations

  • It only inspects routes on the openai-completions protocol (the one vLLM/OpenAI‑compatible gateways speak). Other protocols are left alone.
  • Wire spellings come straight from the listing, so if your vLLM uses non‑standard effort names, what it advertises is what's written.
  • Resolving a key uses ctx.credentials.resolve(apiKeyEnv) then the process environment, mirroring the pi-ai adapter.
  • This is an additive settings sync; it does not modify DSH source and needs no Web rebuild.