dsh-eval
DeepSeek Harness plugin for benchmarking and evaluating agent runs from YAML-defined test cases.
Install
$ dsh plugin --profile eval add dsh-evalPlugin Overview & Capabilities
AI-assisted organization based on the public repository snapshot. The content must be grounded in source evidence and does not replace compatibility or security verification.
Key Capabilities
- Runs YAML-defined benchmarks with headless dsh subprocesses across cases and trials
- Harvests persisted session logs as trial traces, including merged subagent logs
- Calculates task, tool, tool-selection, step, token, context, latency, cost, retry, and invalid-tool-call metrics
- Supports scripted task checks and expected-tool matching
- Uses optional LLM judge configuration for final-answer scoring and hallucination flags
- Generates JSON run artifacts and Markdown reports
- Compares two runs with signed B-minus-A deltas and win/lose/tie statistics
- Replays recorded logs and imports Codex or Claude Code session logs
Useful For
- Regression testing DeepSeek Harness agents or skills
- Comparing two agent, prompt, model, or configuration variants
- Running repeatable benchmark suites in CI using recorded logs
- Analyzing tool selection, task completion, latency, token use, and estimated cost
- Converting Codex or Claude Code session logs into evaluation runs
Who It Fits
- DeepSeek Harness plugin and agent developers
- Teams maintaining agent benchmark suites
- Engineers evaluating changes to agent configurations or models
Documented Limitations
- The package targets official @deepseek-ai releases with 0.1.0-rc.6 peer dependencies.
- Benchmark execution requires a dsh launcher and runs one headless dsh subprocess per case and trial.
- Task success and tool-selection metrics depend on benchmark expected.check and expected.tool configuration.
- LLM-based final-answer scores and hallucination flags require judge configuration.
- Parallel trial execution and a web dashboard are listed as roadmap items, not current features.
DSH Compatibility
Version-specific runtime evidence collected by DSH Plugin. A missing result means we have not tested that combination yet.
Security Signals
Objective signals discovered from package metadata and source inspection. These are not a guarantee that a plugin is safe.
GitHub reports the repository license as MIT.
No root package.json was captured in the latest GitHub snapshot.
Public GitHub source metadata is available for this registry snapshot.
Source & Registry Notes
Public provenance, Registry classification, and the latest source check for this entry, kept separate from runtime verification.
- Source repository
- hccccc01333/dsh-eval
- Registry source
- GitHub · dsh-plugin topic
- Registry classification
- Plugin
- Source checked
- 47f39d7 · 2026-09-18
This project is independently indexed from public source information. DSH Plugin is not affiliated with DeepSeek or the plugin author. Always check the author repository before installation.
Repository Activity
- GitHub stars
- 2GitHub stars
- Forks
- 0
- Open issues
- 0
- Last commit
- 2026-08-14
- Last release
- No release detected
Related DSH Plugins
Ranked by overlapping capabilities, use cases, plugin type, categories, and DSH profile.
Persistent multi-model Agent teams and bounded DAG workflows with GUI controls for DeepSeek Harness.
View plugin →Second-model AI auto-review for DeepSeek Harness approval requests with fail-closed safety and session audit.
View plugin →Native SQLite-backed local project taskboard plugin and UI for DeepSeek Harness with Agent claim and review flows.
View plugin →Compatibility bridge and host ABI enabling unmodified Pi plugins to run natively on DeepSeek Harness.
View plugin →