yan77-h/dsh-agent-evaluator
yan77-h★ 0JavaScriptLast synced: 2026-08-16
agent evaluation
README excerpt
dsh-agent-evaluator An agent evaluation plugin for DeepSeek Harness: import test sets, run agents, score results, and generate reports. Chinese Install If dsh is installed globally, use it instead of pnpm dsh . Cloning the repository or using a git URL also works. Usage Web UI: Settings → Agent Evaluation, or run /eval commands: Subcommands: import , list , drop , model , models , judge , run , resume , kill , runs , progress , report . Headless / CI: Features - Imports JSONL / JSON / CSV / TSV and directories, adapting common benchmark field names. - Runs each case in an isolated agent with concurrency, retries, and timeouts. - Scorers: exact , contains , llm . - Persists reports under $DSH EVAL HOME/eval/reports , with checkpoint/resume support. - Web panel provides model pickers, live progress bars, and pass/fail charts. Documentation - INSTALL.md / INSTALL.zh.md — installation - USAGE.md / USAGE.zh.md — full usage - AGENTS.md / AGENTS.zh.md — architecture and maintenance Development Integration tests and client builds require a deepseek-harness checkout; see AGENTS.md. MIT License, see LICENSE.
View full README on GitHub →Category
🎨 Best DeepSeek Harness Design Plugin. The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, HTML/PDF/PPTX/MP4 export. 🤖 Claude Code / Codex / Cursor / DeepSeek Harness / OpenCode & 20+ CLIs via BYOK.
★ 87,432
volcengine/OpenVikingSelf-evolving Context Database for AI Agents. Unify Agent Memory, Knowledge RAG and Skills.
★ 28,650
titanwings/colleague-skill将冰冷的离别化为温暖的 Skill,欢迎加入数字生命1.0!Transforming cold farewells into warm skills? It's giving rebirth era. Welcome to Digital Life 1.0. 🫶
★ 22,793