dsh-eval
DeepSeek Harness plugin for benchmarking and evaluating agent runs from YAML-defined test cases.
Install
$ dsh plugin --profile eval add dsh-evalRepository Intelligence
AI-assisted, source-grounded explanation based on the public repository snapshot. It does not replace compatibility or security verification.
Key Capabilities
- Runs YAML-defined benchmarks with headless dsh subprocesses across cases and trials
- Harvests persisted session logs as trial traces, including merged subagent logs
- Calculates task, tool, tool-selection, step, token, context, latency, cost, retry, and invalid-tool-call metrics
- Supports scripted task checks and expected-tool matching
- Uses optional LLM judge configuration for final-answer scoring and hallucination flags
- Generates JSON run artifacts and Markdown reports
- Compares two runs with signed B-minus-A deltas and win/lose/tie statistics
- Replays recorded logs and imports Codex or Claude Code session logs
Useful For
- Regression testing DeepSeek Harness agents or skills
- Comparing two agent, prompt, model, or configuration variants
- Running repeatable benchmark suites in CI using recorded logs
- Analyzing tool selection, task completion, latency, token use, and estimated cost
- Converting Codex or Claude Code session logs into evaluation runs
Who It Fits
- DeepSeek Harness plugin and agent developers
- Teams maintaining agent benchmark suites
- Engineers evaluating changes to agent configurations or models
Documented Limitations
- The package targets official @deepseek-ai releases with 0.1.0-rc.6 peer dependencies.
- Benchmark execution requires a dsh launcher and runs one headless dsh subprocess per case and trial.
- Task success and tool-selection metrics depend on benchmark expected.check and expected.tool configuration.
- LLM-based final-answer scores and hallucination flags require judge configuration.
- Parallel trial execution and a web dashboard are listed as roadmap items, not current features.
DSH Compatibility
Version-specific runtime evidence collected by DSH Plugin. A missing result means we have not tested that combination yet.
Security Signals
Objective signals discovered from package metadata and source inspection. These are not a guarantee that a plugin is safe.
No root package.json was captured in the latest GitHub snapshot.
GitHub reports the repository license as MIT.
Public GitHub source metadata is available for this registry snapshot.
Source & Registry Notes
Traceable source and registry metadata for this entry, kept separate from runtime verification.
- Source repository
- hccccc01333/dsh-eval
- Registry source
- Public GitHub repository
- Source snapshot
- 8c04b327d06a
- Artifact type
- plugin
- AI enrichment
- gpt-5.6-terra · 2026-08-14
- Prompt version
- dsh-plugin-enrichment-v3
This project is independently indexed from public source information. DSH Plugin is not affiliated with DeepSeek or the plugin author. Always check the author repository before installation.
Repository Activity
- GitHub stars
- 0
- Forks
- 0
- Open issues
- 0
- Last commit
- 2026-08-14
- Last release
- —
Related DSH Plugins
Ranked by overlapping capabilities, use cases, plugin type, categories, and DSH profile.
Schedule standalone coding tasks in fresh DeepSeek Harness Agent sessions, with workspace-scoped controls and durable run history.
View plugin →一个适用于deepseek-harness的插件,功能是显示当前账户余额以及当前会话预估的费用消耗 | A plugin for deepseek-harness that displays the current account balance and the estimated cost consumption of the current session.
View plugin →Cross-agent, local-first persistent memory plugin for DeepSeek Harness (DSH), powered by Mnemon. It shares long-term memory across Mnemon-enabled agents and adds runtime memory, searchable project documents, semantic recall, knowledge graph, and a Sidebar UI.
View plugin →A curated list of plugins, skills, MCP servers, patch/profile layers, orchestrators & UIs for DeepSeek Harness (DSH). Visualization · PPT · Coding · Agents · Loops (auto-research) and more. #dsh
View plugin →