hanchn/dsh-multimodal-router

hanchn★ 0JavaScript最后同步: 2026-08-17

在 GitHub 打开

A zero-config, multi-provider vision tool for DeepSeek Harness with automatic local model discovery and privacy-aware remote fallback.

README 摘要

Multimodal Router for DeepSeek Harness 简体中文 A zero-config multimodal plugin for DeepSeek Harness. It adds image understanding and enables complete local realtime voice conversation when both a compatible model and the local voice runtime are ready. Why Multimodal Router? DeepSeek remains the reasoning agent. Multimodal Router delegates image perception to a local or explicitly configured multimodal model, then returns text evidence to the agent. The default path is private and requires no model name or endpoint configuration. Interface preview The conversation keeps the uploaded image, the user's question, and the model's answer together in one view. Features - Zero-config discovery of vision-capable Ollama models via model metadata - An image attachment button plus drag-and-drop and clipboard intake with automatic local analysis - Capability-gated voice conversation: local incremental text in the composer, silence-to-send, and spoken replies - Image and audio for Gemma 4 E2B, E4B, and 12B; image only for 26B and 31B - Original thumbnails remain visible in chat; images are archived locally by content hash - Local-first routing with remote fallback disabled by default - Multiple pri…

在 GitHub 查看完整 README →
内容/媒体deepseek-harnessdsh-pluginmultimodalollamaopenai-compatiblevisionagent

分类