hanchn/dsh-multimodal-router
hanchn★ 0JavaScriptLast synced: 2026-08-17
A zero-config, multi-provider vision tool for DeepSeek Harness with automatic local model discovery and privacy-aware remote fallback.
README excerpt
Multimodal Router for DeepSeek Harness 简体中文 A zero-config multimodal plugin for DeepSeek Harness. It adds image understanding and enables complete local realtime voice conversation when both a compatible model and the local voice runtime are ready. Why Multimodal Router? DeepSeek remains the reasoning agent. Multimodal Router delegates image perception to a local or explicitly configured multimodal model, then returns text evidence to the agent. The default path is private and requires no model name or endpoint configuration. Interface preview The conversation keeps the uploaded image, the user's question, and the model's answer together in one view. Features - Zero-config discovery of vision-capable Ollama models via model metadata - An image attachment button plus drag-and-drop and clipboard intake with automatic local analysis - Capability-gated voice conversation: local incremental text in the composer, silence-to-send, and spoken replies - Image and audio for Gemma 4 E2B, E4B, and 12B; image only for 26B and 31B - Original thumbnails remain visible in chat; images are archived locally by content hash - Local-first routing with remote fallback disabled by default - Multiple pri…
View full README on GitHub →Category
Prompt as Code | GPT-Image2 工业级提示词引擎与模板库,470+ 个案例逆向工程,20+ 套工业级模板,并提炼出Skills,持续更新中
★ 10,851
liustack/modlensThe first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。
★ 2,700
Alisa0808/vox-directorTurn one topic into a finished Vox-style paper-collage explainer/ad video — automated end to end on Atlas Cloud + ffmpeg. An agent skill.
★ 1,335