welsione/dsh-mmx-bridge
welsione★ 1JavaScript最后同步: 2026-08-15
MiniMax multimodal capability hub for DeepSeek Harness (DSH): image understanding (VLM), text/image-to-video, speech, music, audio cover, web search, quota — one mmx_multimodal model tool wrapping the mmx-cli.
README 摘要
dsh-mmx-bridge DeepSeek Harness(DSH)的 MiniMax 多模态桥接插件。 English · English README 功能 一个 mmx bridge 工具,覆盖 MiniMax 全部多模态能力: - 图片理解 (describe)· 文生图 (image) - 视频生成 (video)· 语音合成 (speech) - 音乐生成 (music)· 音频翻唱 (cover) - 联网搜索 (search)· 用量查询 (quota) 生成产物经 /mmx-files/ 同源提供(支持 Range),对话流直接内嵌图片预览与音视频播放器,结果携带可播放 URL。 效果展示 图片生成 语音合成 :--: :--: 图像识别(VLM 描述) 插件配置页 :--: :--: 安装 将下面的提示词复制给你的 AI 助手(Agent),它会按照 AGENT.md 完成安装: 请阅读仓库根目录的 AGENT.md 文档,按照其中的步骤,在当前 DSH profile 中完成 dsh-mmx-bridge 插件的安装、挂载与验证。 相关 DeepSeek Harness · MiniMax CLI · awesome-dsh-plugins 许可证 MIT
在 GitHub 查看完整 README →分类
The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网第一个 DeepSeek Harness 视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。
★ 2,047
Anionex/agent-vision-toolkit为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
★ 920
Anionex/dsh-vision-toolkit让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.
★ 457