1710782766/dsh-llm-vision

1710782766★ 0TypeScript最后同步: 2026-08-17

在 GitHub 打开

Reliable vision + OCR for text-only models on DeepSeek Harness: describe_image (normal/critical) + extract_text tools, auto-preprocessing, retries, and a persistent answer cache.

README 摘要

dsh-llm-vision English 中文 Reliable vision + OCR for text-only models on DeepSeek Harness. Prompt engineering that makes screenshot QA trustworthy, the same reliability engineering (preprocessing / retries / persistent cache) — plus the DSH-native experience paste-bridge, live settings card, URL input, and attachment references. Status : v0.1.0 on GitHub and npm. Verified end-to-end against a live OpenAI-compatible vision endpoint (DashScope qwen3-vl-plus / qwen3.5-ocr ) in the real web GUI; 183 offline tests. Why Text-only models (DeepSeek V4, GLM text series, …) cannot see images. This plugin registers two model-facing tools backed by any OpenAI-compatible vision endpoint: Tool Purpose describe image Image understanding with two perspectives: normal (natural description) and critical (objective inspection that actively reports text misalignment, overlap, occlusion, wrapping anomalies, missing elements, and separates fact from guess). The critical lens is the antidote to vision models rationalizing rendering bugs — use it for page/UI problem reports and screenshot-vs-design comparisons. extract text OCR & document parsing through a dedicated OCR model — ID cards, invoices, receipts…

在 GitHub 查看完整 README →
内容/媒体deepseek-harnessdshdsh-pluginimage-understandingllmmultimodalocrvision

分类