Skip to content

agent-evaluation 6

  • oh-my-knowledge v0.54.0 Verified 18

    OMK β€” Observe. Measure. Know. Evidence-backed knowledge changes for AI applications.

    dsh plugin add oh-my-knowledge
  • plugin-verifier v0.3.3 Verified Web UI 3

    LLM-as-a-Verifier for dsh with the Best-of-N conversation mode built in: rank N candidates with a fine-grained verifier (expected grade over the logprob distribution), on demand via the verify tool or automatically on every turn of a Best-of-N session. In

    dsh plugin add @aispin/plugin-verifier
  • dsh-replay-lab v0.1.6 Verified Web UI 2

    Replay real DeepSeek Harness turns against Standard, Minimal, Anchored, or plugin candidates with frozen request-surface evidence

    dsh plugin add @webwalkerhq/dsh-replay-lab
  • dsh-eval v0.3.0 Verified 2

    Agent evaluation platform: benchmark YAML, headless run orchestration, trace-based metrics, and run reports

    dsh plugin add dsh-eval
  • dsh-harbor-evolution v0.7.3 Verified Web UI 1

    DeepSeek Harness plugin and bundled Skill for safely evolving Cordis Candidates with Harbor.

    dsh plugin add dsh-harbor-evolution
  • dsh-plugin-abtest v0.1.0 Verified

    Paired experiments and promotion gates for DSH plugins.

    dsh plugin add dsh-plugin-abtest