evaluation 8
Evidence-first multi-model coding comparison workbench for DeepSeek Harness
dsh plugin add dsh-evidence-arenaControlled A/B comparisons and evidence-backed reports for DeepSeek Harness plugins and presets.
dsh plugin add dsh-plugin-compareCheck which DeepSeek Harness plugins actually loaded, and look up public evaluation boards on trapstreet.run
dsh plugin add @trapstreet/dsh-trapstreetPlugin value auditor for DeepSeek Harness: judge a plugin before install and audit installed ones — heuristic scan + LLM judge, with model-switch re-audit reminders. · DSH 插件价值裁判:装前判断值不值得装,装后审计是否还该留,模型切换时提醒复核。
dsh plugin add dsh-plugin-judgeDSH web plugin: a LiveBench tab in the Trajectory view (right of 对话/轨迹). Run LiveBench evaluations against every model configured in the DeepSeek Harness — pick provider/model, category, task, release and question range from dropdowns, watch progress, and
dsh plugin add dsh-livebench-panelDSH 试金石:给「自改造」补上评测这一环——用一套金标准用例,把改动前 vs 改动后跑一遍、打分、对比,好就留、不好就撤。
dsh plugin add dsh-touchstoneMeasure whether a dsh setup change actually helped: register repeatable cases, run them, and diff the results before and after you change rules, skills, or models.
dsh plugin add dsh-verdictDSH 插件:同一任务、同一代码基线,多 Variant(模型/Preset)并行对照实验、EvidenceReceipt 对比与胜出 Patch 导出
dsh plugin add dsh-bakeoff