DeepSeek Harness Plugins
DeepSeek Harness PluginIndexed

dsh-eval

DeepSeek Harness plugin for benchmarking and evaluating agent runs from YAML-defined test cases.

UI & ProductivityTerminal & TUIDeveloper ToolsSecurity & PolicySkills & WorkflowsRemote Execution

Install

$ dsh plugin --profile eval add dsh-eval

Repository Intelligence

AI-assisted, source-grounded explanation based on the public repository snapshot. It does not replace compatibility or security verification.

source-grounded
dsh-eval is an installable DeepSeek Harness evaluation plugin that runs benchmark cases through headless dsh profiles, harvests session logs as traces, calculates run metrics, and produces JSON and Markdown reports. It also supports paired run comparisons, LLM judging, recorded-log replay, and importing Codex or Claude Code logs.

Key Capabilities

  • Runs YAML-defined benchmarks with headless dsh subprocesses across cases and trials
  • Harvests persisted session logs as trial traces, including merged subagent logs
  • Calculates task, tool, tool-selection, step, token, context, latency, cost, retry, and invalid-tool-call metrics
  • Supports scripted task checks and expected-tool matching
  • Uses optional LLM judge configuration for final-answer scoring and hallucination flags
  • Generates JSON run artifacts and Markdown reports
  • Compares two runs with signed B-minus-A deltas and win/lose/tie statistics
  • Replays recorded logs and imports Codex or Claude Code session logs

Useful For

  • Regression testing DeepSeek Harness agents or skills
  • Comparing two agent, prompt, model, or configuration variants
  • Running repeatable benchmark suites in CI using recorded logs
  • Analyzing tool selection, task completion, latency, token use, and estimated cost
  • Converting Codex or Claude Code session logs into evaluation runs

Who It Fits

  • DeepSeek Harness plugin and agent developers
  • Teams maintaining agent benchmark suites
  • Engineers evaluating changes to agent configurations or models

Documented Limitations

  • The package targets official @deepseek-ai releases with 0.1.0-rc.6 peer dependencies.
  • Benchmark execution requires a dsh launcher and runs one headless dsh subprocess per case and trial.
  • Task success and tool-selection metrics depend on benchmark expected.check and expected.tool configuration.
  • LLM-based final-answer scores and hallucination flags require judge configuration.
  • Parallel trial execution and a web dashboard are listed as roadmap items, not current features.

DSH Compatibility

Version-specific runtime evidence collected by DSH Plugin. A missing result means we have not tested that combination yet.

Not tested yetNo runtime compatibility tests have been published yet.

Security Signals

Objective signals discovered from package metadata and source inspection. These are not a guarantee that a plugin is safe.

Package Manifest Not Found

No root package.json was captured in the latest GitHub snapshot.

info
License Declared

GitHub reports the repository license as MIT.

info
Source Available

Public GitHub source metadata is available for this registry snapshot.

info

Source & Registry Notes

Traceable source and registry metadata for this entry, kept separate from runtime verification.

Source repository
hccccc01333/dsh-eval
Registry source
Public GitHub repository
Source snapshot
8c04b327d06a
Artifact type
plugin
AI enrichment
gpt-5.6-terra · 2026-08-14
Prompt version
dsh-plugin-enrichment-v3

This project is independently indexed from public source information. DSH Plugin is not affiliated with DeepSeek or the plugin author. Always check the author repository before installation.

Repository Activity

GitHub stars
0
Forks
0
Open issues
0
Last commit
2026-08-14
Last release