huangdaxianer/dsh-dual-model-eval

huangdaxianer★ 0TypeScript最后同步: 2026-08-19

在 GitHub 打开

Compare multiple coding models side by side in DeepSeek Harness with isolated Git worktrees, live traces, and adoptable results

README 摘要

DeepSeek Harness Multi-Model Evaluation ( dsh-dual-model-eval ) English 简体中文 Compare multiple coding models side by side inside DeepSeek Harness. dsh-dual-model-eval is an installable multi-model and coding-agent evaluation plugin: one prompt runs concurrently across selected LLM routes, every model works in an isolated Git worktree, and the normal Chat tab streams tool trajectories and renders comparable, adoptable results. Compatibility: the first release targets DeepSeek Harness 0.1.0-rc.7 . DeepSeek Harness is currently a developer preview, so plugin APIs may change between release candidates. Install Install the pinned release into the built-in web profile: Restart the Harness web process after installation: The repository contains committed, prebuilt lib/ artifacts. Installing from GitHub therefore does not require authorizing a dependency prepare script. What it adds - A Comparison test switch inside the existing model selector. - Multi-select for two to four configured model routes. - Concurrent execution from one shared Git commit in isolated worktrees. - Live, independently expandable tool-call trajectories for every model. - Compact elapsed-time and tool-count summaries,…

在 GitHub 查看完整 README →
工具/开发agent-evaluationcoding-agentdeepseekdeepseek-harnessdsh-plugingit-worktreellm-evaluationmodel-evaluation

分类