xuxun-oss/dsh-vision-imagen
xuxun-oss★ 2JavaScript最后同步: 2026-08-20
DeepSeek Harness all-in-one: no model switching — regular DeepSeek auto-routes to vision & image gen. Multi-backend: Gemini + any OpenAI-compatible (GPT-4o, Qwen-VL, GLM-4V, gpt-image, DALL-E, Flux, OpenRouter). gemini_vision/gemini_generate_image/gemini_optimize_image with vision self-check. Better than modlens.
README 摘要
dsh-vision-imagen DeepSeek Harness 插件:为 DeepSeek 模型桥接 Google Gemini 的多模态识别与图像生成。 当 DeepSeek 模型需要看图、生图、改图时,自动调用 Gemini: 工具 作用 vision read 识别/读取/描述/分析图片(OCR、物体、图表),支持本地路径或 http(s) URL generate image 文本生图,生成后 必定用 Gemini 视觉模型自检反馈 (不达标时按优化 prompt 重绘) edit image 改图/优化已有图片:先分析原图 → 生成改进版 → 自检迭代 ✨ 核心卖点:全能一体,无需切换模型 - 一个 DeepSeek 会话搞定一切 :直接用常规 DeepSeek 模型(不必换成视觉模型),需要看图、生图、改图时,插件自动路由到对应的 Gemini 多模态/生图模型—— 全程零手动切换模型 。 - 比 modlens 等更完整 :modlens 只是「给纯文本模型装一只读图的眼睛」;本插件是 视觉识别 + 图像生成 + 图像编辑 + 自检反馈的完整闭环 ,并且直接复用你自己的 Gemini API Key,不依赖任何第三方桥接、无额外订阅。 - 自动调用,无需操心 :识别 → vision read ;生图 → generate image ;改图 → edit image 。DeepSeek 模型通过系统提示与工具描述自动选对工具,对用户完全透明。 - 自检闭环,质量把关 :生成/优化后,自动用 Gemini 视觉模型检查成品图,不达标按优化提示自动重绘,并把检查结论(画面描述 / 达标与否 / 问题清单)直接反馈出来。 🧩 环境要求(依赖) 依赖 要求 说明 DeepSeek Harness( dsh ) ≥ 0.1.0-rc.7 (rc 通道) 插件宿主。先全局安装 dsh,再用 dsh plugin --profile web add ... 加载本插件 Node.js ≥ 18 宿主半运行环境( engines ) @deepseek-ai/dsh-tools peer 依赖( ) 必需 :宿主半通过 defineTool 注册工具。由 dsh 宿主提供,或随 pnpm install 自动解析—— 无需手动安装 react ^18.2.0(peer) 浏览器半(设置页)所需,由 dsh Web 前端提供 @deepseek-ai/dsh-client-runtime rc 通道 客户端核心服务,声明于 dsh.client.inject ,dsh Web 构建时自动注入 后端 API Key 自备 Gemini(Google AI Studio 免费 Key)或所选 OpenAI 兼容后端 本插件运行时 只依赖 @deepseek…
在 GitHub 查看完整 README →分类
:rocket: The Ultimate Image Uploader for Efficient Creators. Supports Obsidian, Typora, VS Code etc. and 60+ image hosting services (S3, GitHub, Cloudflare R2, Imgur, Aliyun OSS...). Paste, upload, done.
★ 27,002
liustack/modlensThe first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。
★ 3,503
Tiger3807861189/J-Space-Cognition-Suite-V3.6J-Space Cognition Suite V3.6 - AI cognitive-enhancement Skills based on Anthropic's J-space global workspace research. | 哔哩哔哩:Tiger380 (UID 3494375382321675) — https://space.bilibili.com/3494375382321675
★ 3,013