welsione/dsh-mmx-multimodal

welsione★ 0JavaScriptLast synced: 2026-08-14

Open on GitHub

MiniMax multimodal capability hub for DeepSeek Harness (DSH): image understanding (VLM), text/image-to-video, speech, music, audio cover, web search, quota — one mmx_multimodal model tool wrapping the mmx-cli.

README excerpt

dsh-mmx-multimodal MiniMax multimodal capability hub for DeepSeek Harness (DSH). One model tool, mmx multimodal , covers MiniMax's whole multimodal stack — and can optionally take over the harness's built-in web search . What it does Registers a single DSH model tool that dispatches to the MiniMax mmx CLI (npm package mmx-cli ). Your agent gets image understanding, image/video generation, TTS, music, audio cover, web search and quota checks through one tool — no API wiring, no custom endpoints. Optionally, the plugin can also take over the harness's built-in web search tool with an mmx-backed implementation — see Web search shadow. Features action Capability Key parameters :-- :-- :-- describe Image understanding (VLM) image (path or URL), optional prompt image Text-to-image (default 3 images per call) prompt , optional aspectRatio video Text/image-to-video (waits for completion) prompt , optional image / duration / ratio speech Text-to-speech text , optional voice , out music Music generation prompt , optional lyrics or instrumental , out cover Audio cover (voice conversion) prompt + audio (reference), out search Web search q quota Usage / balance query — - Zero-config — sensible …

View full README on GitHub →
Media / Contentagent-toolai-agentcordisdeepseek-harnessdsh-pluginimage-generationminimaxmultimodal

Category