satan9394/dsh-llm-eval
satan9394★ 0JavaScript最后同步: 2026-08-19
DSH skill: LLM 评估,忠实度/相关性/测试集/幻觉检测/回归守护(受 wshobson/agents 启发)
README 摘要
dsh-llm-eval LLM 评估:忠实度/相关性/正确性/完整性维度、测试集构建、幻觉专项检测、 回归守护。 受 wshobson/agents(38k★ MIT) 的 llm-application-dev/llm-evaluation 技能启发,改编为 DSH 中文原创精简版。 安装 使用 对 agent 说"评估这个 LLM 应用 / 检测幻觉", llm-eval 技能输出 评估报告(维度/测试集/通过率/失败模式)。 结构 License MIT。原创精简改编,灵感来自 wshobson/agents(MIT)。
在 GitHub 查看完整 README →分类
🌊 The original agent meta-harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
★ 68,716
volcengine/OpenVikingSelf-evolving Context Database for AI Agents. Unify Agent Memory, Knowledge RAG and Skills.
★ 31,761
titanwings/colleague-skill将冰冷的离别化为温暖的 Skill,欢迎加入数字生命1.0!Transforming cold farewells into warm skills? It's giving rebirth era. Welcome to Digital Life 1.0. 🫶
★ 23,757