agent-evaluation 8
OMK β Observe. Measure. Know. Evidence-backed knowledge changes for AI applications.
dsh plugin add oh-my-knowledgeHost-side Cordis control plugin for aeval: one-shot config injector (experiment variables, gateway-lease budgets, fork lineage, bundle descriptor), with a standalone status mode, a Web status chip, and a self-check command.
dsh plugin add dsh-eval-controlDeepSeek Harness plugin and Skill for Harbor Candidate and Historical Session evaluation workflows.
dsh plugin add dsh-harbor-evolutionAgent evaluation platform: benchmark YAML, headless run orchestration, trace-based metrics, and run reports
dsh plugin add dsh-evalLLM-as-a-Verifier for dsh with the Best-of-N conversation mode built in: rank N candidates with a fine-grained verifier (expected grade over the logprob distribution), on demand via the verify tool or automatically on every turn of a Best-of-N session. In
dsh plugin add @aispin/plugin-verifierReplay real DeepSeek Harness turns against Standard, Minimal, Anchored, or plugin candidates with frozen request-surface evidence
dsh plugin add @webwalkerhq/dsh-replay-labRetrospective agent-session evaluation for DeepSeek Harness: deterministic grade cards from persisted session logs, cross-session regression diffs, era comparison β no benchmark runs, no LLM judges.
dsh plugin add dsh-session-evalPaired experiments and promotion gates for DSH plugins.
dsh plugin add dsh-plugin-abtest