Chuyển đến nội dung chính

dsh-vllm-effort-sync

Đã xác minh

dsh-vllm-effort-sync · v0.1.2 · MIT

DSH plugin: read reasoning_efforts from OpenAI-compatible (vLLM) providers' /models listing and sync them into the llm-pi-ai settings section so the Web composer shows a per-model reasoning-effort picker.

Cài đặt

dsh plugin add dsh-vllm-effort-sync

Xác nhận layer đã áp bằng dsh --profile default --dump-config — xem hướng dẫn cài plugin.

Mã nguồn

Phát hành lên npm mà không có repository công khai. Hãy kiểm tra nội dung package trước khi cài.

Thẻ

Tác giả

Readme

dsh-vllm-effort-sync

A DeepSeek Harness (DSH) plugin that makes the reasoning‑effort picker show up in the Web composer for OpenAI‑compatible (e.g. self‑hosted vLLM) models — by reading the reasoning_efforts each provider advertises in its GET /models listing and syncing it into the llm-pi-ai settings section.

Why it exists

Your settings.yaml connects a provider like this (via the llm-pi-ai adapter):

llm-pi-ai:
  providers:
    my:
      apiKeyEnv: MY_API_KEY
      api: openai-completions
      baseURL: http://8.159.147.156:30011/v1
      models:
        - id: deepseek-v4-flash-vllm
          contextWindow: 1048576

A hand‑declared model like deepseek-v4-flash-vllm has no reasoning capability by default. DSH itself never looks at reasoning_efforts in the /models reply (its built‑in discovery only reads id/name/contextWindow/maxTokens), so:

  • the Web model selector shows no "推理等级 / reasoning effort" menu, and
  • requests don't carry reasoning_effort.

…even though your vLLM genuinely advertises it:

"reasoning_efforts": { "off": null, "low": "low", "medium": "medium", "high": "high", "max": "max" }

This plugin closes that gap automatically.

What it does

On load (and reactively whenever the llm-pi-ai settings section changes), it:

  1. Reads the llm-pi-ai settings namespace.
  2. For every provider with a baseURL on the OpenAI‑compatible chat‑completions protocol, it resolves the API key (through your apiKeyEnv credential) and calls GET <baseURL>/models.
  3. Extracts each model's reasoning_efforts map (pi‑ai levels → wire spelling).
  4. Writes it back as a per‑model reasoningEfforts declaration plus the compat switches pi‑ai needs to dispatch reasoning_effort (thinkingFormat: deepseek, supportsReasoningEffort: true).

It is additive and idempotent:

  • Models the listing reports no efforts for are left untouched.
  • A reasoningEfforts you already declared manually is never overwritten (manual wins).
  • Nothing is written when the produced section equals the current one (no settings churn, no loops).
  • A provider that's unreachable, 401s, or isn't openai-completions is skipped and logged — it never takes the route down.

Because pi-ai re-reads its profiles on every request, the change takes effect without a restart; just refresh the Web page.

Install

The plugin is a DSH bundle, installed into your existing web profile with the dsh CLI (a dsh.cmd ships with your install and is on PATH):

cd /d D:\
dsh plugin --profile web add ./dsh-vllm-effort-sync

This initializes/uses the web profile, links the package into its node_modules, and appends dsh-vllm-effort-sync to the profile's bundle list (because package.json declares dsh.bundle). Verify the layer:

dsh --profile web --dump-config   # look for the "# == dsh-vllm-effort-sync" layer

Then (re)start your web app:

node --import tsx/esm apps/cli/src/bin.ts web

Installing from npm (like any third-party bundle)

The package is a publish-ready npm bundle: main points at the compiled lib/index.js, types at lib/index.d.ts, and the @deepseek-ai/* imports are peer dependencies (resolved from the DSH runtime that loads it, exactly how @linxin666/dsh-web-ui-all / dsh-dafeiyu work). To publish:

# give it your own name first (npm already has "dsh-vllm-effort-sync")
npm login
npm publish        # runs prepack -> tsc build -> tarball containing lib/ + patch

Then anyone (including you, on any machine) installs by package name:

dsh plugin --profile web add <your-scope>/dsh-vllm-effort-sync

Or pack a local tarball to distribute without a registry:

pnpm pack          # -> dsh-vllm-effort-sync-0.1.0.tgz
dsh plugin --profile web add ./dsh-vllm-effort-sync-0.1.0.tgz

Note: prepack/prepare compile src/index.ts to lib/ with tsc. A full build needs the DSH type sources (this checkout) available at publish time; the shipped tarball already contains the built lib/, so consumers installing the tarball/npm package never need them.

Optional per‑plugin config lives in the profile's cordis.patch.yml:

- id: vllm-effort-sync
  name: dsh-vllm-effort-sync
  config:
    providers: [my]          # only sync this route (default: all with a baseURL)
    applyCompat: true        # add pi-ai compat switches (default true)
    syncOnStart: true        # sync at load (default true)
    watch: true              # re-sync on section changes (default true)

To remove:

dsh plugin --profile web remove dsh-vllm-effort-sync

What you should see

After the first sync pass, $DSH_HOME/settings.yaml gains, under your provider's model:

llm-pi-ai:
  providers:
    my:
      apiKeyEnv: MY_API_KEY
      api: openai-completions
      baseURL: http://8.159.147.156:30011/v1
      compat:
        thinkingFormat: deepseek
        supportsReasoningEffort: true
      models:
        - id: deepseek-v4-flash-vllm
          contextWindow: 1048576
          reasoningEfforts:
            off:
            low: low
            medium: medium
            high: high
            max: max

Then in the Web UI: reopen the conversation composer, use the model selector (the button that reads "选择模型 / model"), and deepseek-v4-flash-vllm now offers a 推理等级 / reasoning effort submenu with Off / Low / Medium / High / Max, and picking one records reasoningEffort next to the model and sends reasoning_effort on the wire.

Files

  • src/index.ts — the plugin source (imports only @deepseek-ai/cordis, @deepseek-ai/dsh-credentials, @deepseek-ai/dsh-settings, all peer deps).
  • lib/index.js + lib/index.d.ts — compiled build output (what package/entry/dsh.bundle points at).
  • tsconfig.build.json — tsc config (compiles src/ → lib/).
  • cordis.patch.yml — the bundle patch that inserts the vllm-effort-sync row.
  • package.json — the bundle manifest (dsh.bundle) with build/prepack scripts and peer deps.
  • README.md — this file.

Notes & limitations

  • It only inspects routes on the openai-completions protocol (the one vLLM/OpenAI‑compatible gateways speak). Other protocols are left alone.
  • Wire spellings come straight from the listing, so if your vLLM uses non‑standard effort names, what it advertises is what's written.
  • Resolving a key uses ctx.credentials.resolve(apiKeyEnv) then the process environment, mirroring the pi-ai adapter.
  • This is an additive settings sync; it does not modify DSH source and needs no Web rebuild.