dsh-plugin-evaluation/dsh-plugin-evaluation-standards

dsh-plugin-evaluation★ 1JavaScript最后同步: 2026-08-20

在 GitHub 打开

Open evaluation datasets, test cases, and metrics for DSH plugins.

README 摘要

DSH Plugin Evaluation Datasets English 中文 日本語 A growing collection of evaluation datasets for DSH plugins. Each dataset is a profile (which metrics to use) and a cases file (test prompts and expected answers). Pick one that fits your plugin, run its cases, and use the results to understand how your plugin behaves. Start here 1. Browse the datasets. 2. Choose one that matches your plugin and the scenarios you want to cover. 3. Open its profile and cases files. 4. Run the cases against your plugin and review the results. Need a dataset that is not here yet? Use the AI-assisted authoring guide to draft one, then contribute it. Build this collection with us Plugin authors, users, and people who know real business scenarios are all welcome. You do not need a finished JSON dataset to participate: - Have a real scenario? Open an issue with how a user would ask, what the plugin should do, and the supporting facts or setup conditions. - Have a small set of cases? Submit a profile and cases following the contribution guide. - Maintain a dataset long term? Keep it in your own repository and add it to this catalog using the external dataset listing guide. Common tasks, tricky conditions, and c…

在 GitHub 查看完整 README →
工具/开发benchmarksdeepseek-harnessdeepseek-harness-plugindshdsh-pluginevaluation-datasetsllm-evaluationtest-cases

分类