跳到主要内容

dsh-desktop-agent

已验证

@logictan/dsh-desktop-agent · v0.1.6 · MIT · Web 界面

DSH desktop agent: a native-vision decision loop that drives the host desktop through the Cua Driver tools. DSH 桌面 Agent 子插件。

安装

dsh plugin add @logictan/dsh-desktop-agent

用 dsh --profile default --dump-config 确认 layer 已生效 —— 参见安装指南。

源码

发布到 npm 但没有公开仓库。安装前请检查包内容。

标签

说明文档

@logictan/dsh-desktop-agent

A DSH desktop agent that drives one window on your own desktop through the Cua Driver tools, using screenshots as its primary sense.

Each step captures the target window, sends the picture to a vision-capable model from DSH's own model directory, and performs the click, key, or text entry the model chose. Nothing is hard-coded to a particular model and there is no second API key: every call goes through ctx.llm.

Why screenshots instead of the accessibility tree

The accessibility tree is the right sense for ordinary applications, and this plugin still uses it — but it cannot see a game. A game is a self-drawn surface: it exposes no meaningful actionable nodes, and most native games additionally filter input that is routed by process id. For those targets the only channel that carries any information is the picture itself.

So the plugin routes by capability rather than by configuration:

The resolved model declares image Channel
yes screenshot + the element anchors + a JSON action
no, or the field is absent the driver's element table and markdown tree

"Absent" is treated as "cannot see" on purpose. The host rewrites an image sent to a route that does not accept one into placeholder text, after which the model is describing a picture it never received — and answering confidently about it.

Both senses arrive together

The vision channel used to ask for the screenshot alone, which left the model estimating every coordinate from the picture — the documented failure mode of vision models, and a real one here: asked for the centre of the largest button in a 1567×894 screenshot, a live model answered (310, 1062), a y past the bottom of the image.

One capture now returns the screenshot AND the window's own controls, each with a token, a role, a label, and its frame in the same pixel space. A decision may address a control by its token instead of guessing a coordinate, and the token names the element the driver itself identified. A token that has been superseded by a newer capture is refused with stale_element_token rather than mis-clicking, which is one more reason the loop re-observes every step.

Measured cost of the extra walk, on the heaviest tree available here (a browser window): 0.5–3.7 s, returning ~174 anchors and ~18 KB at the 300-element cap. The same walk uncapped returned 1130 nodes and 68 KB, which buys nothing — an element with no label cannot be named in a decision, and the screenshot already shows it. Only elements carrying both a token and a non-empty label become anchors.

Requirements

None beyond installing this plugin. It publishes the cua_driver_native__* tools itself, from the Cua Driver native SDK it depends on, and dispatches to them.

A profile that still mounts the old @deepseek-ai/dsh-experimental-computer-use-cua-driver-native provider must drop that row and its @deepseek-ai/dsh-computer-use dependency first: both would publish the same tool names, and the second registration is refused with tool "cua_driver_native__click" is already registered.

Settings

Settings → Plugins → desktop-agent → configuration:

Field Meaning
Vision provider / model The route used for each decision, chosen from the models that declare image input. All empty = the session's own current route.
Vision reasoning effort Optional, and only offered for models that declare one.
Max steps Actions per run; defaults to 40.
Screenshot long edge Pixels; defaults to 1568. Lower saves tokens, but too low hides the controls.
Delivery background (default) never steals focus; foreground is required by surfaces that filter per-process-routed input, such as canvas apps and games.
Allow bring-to-front Off by default. When on, the model may use bring_to_front to raise the target window. It verifiably steals your foreground, so it is opt-in.

Tool

desktop_agent({ app, goal, windowId }) resolves one window, runs the observe/decide/act loop until DONE/BLOCKED or a cap is reached, and returns a structured trace. app accepts an application name or a window title; omitting it uses the frontmost window.

When to call it

For any desktop task that takes more than about two steps, call desktop_agent rather than driving the cua_driver_native__* tools yourself. The loop's ~40 captures and decisions stay inside this one call; driving the raw tools by hand puts every screenshot into the agent's own context instead. The raw tools remain the right choice for resolving a window, or for a single action already decided on.

For a task that lives in a web page, browser_agent is the better tool — it drives the DOM over CDP and does not depend on reading pixels. This tool is for native windows, desktop applications, and games.

What it does not do

It does not verify the goal. done is the model's own claim about the screenshot it was shown, and the trace reports it as such. The driver's verify_state is deliberately not used as a second opinion: measured against the one Electron window available it answered unknown (untrusted_source) after ~5.6 s, so it would add latency and a false sense of checking without checking anything.

Only on-screen windows are eligible. A covered window cannot be captured or acted on, so the tool fails with the window named rather than screenshotting whatever happened to be in front.

Scope

v1 drives one window at a time, in the foreground of that window's own desktop, and is meant for tasks whose decisions are measured in seconds — turn-based and strategy games, simulations, settings panels, file managers.

It is not for real-time action games: one vision decision costs 300–800 ms, an order of magnitude away from human reaction time. It does not read or write process memory, and it makes no attempt to evade anti-cheat protection.

Coordinates

Action coordinates are pixels in the screenshot the model was shown, measured from its top-left corner. The driver resolves them against the capture the caller last saw, so every step re-observes before it acts — a new capture of a window also invalidates the element tokens of the previous one.