lhwwxy/dsh-model-deploy
lhwwxy★ 0JavaScript最后同步: 2026-08-18
LLM model selection & deployment analysis tool for DeepSeek Harness: deployability, VRAM, TTFT, latency, throughput and power for 38 models × 20 GPUs/NPUs
README 摘要
dsh-model-deploy 🌐 语言切换 / Language: 简体中文 · English DeepSeek Harness(DSH)插件: LLM 模型选型部署分析 。 给它模型、GPU/NPU 型号、节点数、精度、上下文长度、并发数与互联带宽,它用一阶解析模型给出可审计的估算: - ✅/⚠️/❌ 能否部署 (显存/精度/互联检查)+ 修复建议 - 显存 :单卡与集群拆解(权重 / KV 缓存 / 激活 / 运行时) - 时延 :首 token(TTFT)与整体时延(空闲 / 满载两种口径) - 吞吐 :单请求与系统 tok/s、req/s、预填充 tok/s - 功耗 :空闲 / 典型 / 峰值(含 PUE)、每 token 能耗 - 自动 TP × PP × DP 并行策略与选择理由 安装 或从 GitHub 检出安装: --profile 必填。安装后重启会话(或 profile),工具 schema 才会进入 prompt 组装。 工具 model deploy analyze 主分析工具。示例调用: 字段: model (目录 id,或自定义模型 JSON 字符串)、 gpu (目录 id)、 gpusPerNode (1–8)、 nodes (1–64)、 precision ( fp32 bf16 fp16 fp8 int8 int4 fp4 )、 ctx / batch / prompt / output (tokens)、 intraMode ( auto custom )+ intraGBs 、 interMode ( none ib400 ib800 roce100 roce200 custom )+ interGBs 。 返回结构化报告( status 、 issues 、 strategy 、 memory 、 performance 、 power 、 assumptions ),渲染输出为简洁中文文本报告。 model deploy catalog 列出支持的模型(id、参数量、层数、上下文上限)与 GPU/NPU(id、显存、带宽、 FP16/FP8/INT8/FP4、互联、TDP),支持 query / vendor 过滤。先用它查 id,再调分析工具。 覆盖范围 - 38 个模型 — Qwen3 全系(含 Qwen3-Next/Coder/VL)、Llama 3.1/3.3/4、 DeepSeek-V3/V3.1/V3.2/V4-Flash/V4-Pro/R1、Kimi-K2/-K2-Thinking/-K3、 GLM-4.5-Air/4.6/5、Hunyuan-A13B、Baichuan-M2、Seed-OSS、GPT-OSS、Mixtral、 MiniCPM4、QwQ、Qwen2.5 … - 20 款 G…
在 GitHub 查看完整 README →分类
DeepSeek Harness: Everything is a Plugin.
★ 182,325
amruthpillai/reactive-resumeA one-of-a-kind resume builder that keeps your privacy in mind. Completely secure, customizable, portable, open-source and free forever. Try it out today!
★ 41,477
anywhere-labs/deepseek-harness-desktop为 DeepSeek Harness (DSH) 插件生态打造的现代化桌面端解决方案。万物皆「插件」,桌面本身也是「插件」。
★ 17,784