AlloyPlane/dsh-eye-vision

AlloyPlane★ 1JavaScript最后同步: 2026-08-17

在 GitHub 打开

README 摘要

dsh-eye-vision Give text-only DeepSeek Harness models eyes — image understanding, OCR, and UI analysis through any OpenAI-compatible multimodal API . Fork of dsh-free-vision (MIT) v1.0.1, with: - custom-provider fix — CUSTOM MODEL NAME is now forwarded, so arbitrary OpenAI-compatible endpoints (GPT-4o, Qwen-VL, GLM-4V, youtu-vita, vLLM, Ollama…) actually start - allowed-directories whitelist — the vision engine can read images from your configured workspace roots, not just the engine CWD and home directory How it works The main model never needs image input support. Paste the image path, get answers. Features - image understand tool registered on ctx.tools , visible to every session in the profile - Any OpenAI-compatible endpoint via custom provider — bring your own multimodal API - Free-tier providers built in : qwen (Qwen3-VL-Flash), volcengine (Doubao), siliconflow (DeepSeek-OCR) - Multi-crop for large images (detail preservation) - Allowed-directories whitelist ( LUMA ALLOWED DIRS ) — read images from your workspace - Proxy vars stripped for direct mainland-China API access - Live settings — save via the settings route, no restart needed Installation Restart dsh web . The tool …

在 GitHub 查看完整 README →
内容/媒体aideepseek-harnessdshdsh-pluginimage-understandingllmmultimodalocr

分类