ocr 53
Plug-in vision for text-only LLMs, powered by the free Antigravity CLI
dsh plugin add @liustack/modlensEyes for text-only DeepSeek Harness agents: built-in free vision chain + pixel-level tools. Remote vision providers, including the default free fallback, receive image content unless strict local-only use is configured.
dsh plugin add dsh-vision-routerDeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, OCR, grounding, UI restoration, pixel diff, Artifacts, and Web UI.
dsh plugin add @anionex/dsh-vision-toolkitLocal IMAP invoice download, OCR, archive, and Excel summary bundle for DeepSeek Harness
dsh plugin add @ethanyoq/dsh-invoice-downloaderDeepWatch's capabilities, installable into an existing DeepSeek Harness profile
dsh plugin add @deepwatch/dsh-bundleUnified image understanding plugin for DeepSeek Harness (DSH). Visual twin adapter for native thumbnails + auto-analysis on any text-only model (incl. pi-ai providers); privacy/smart/strict routing; local tools (scan/OCR×4 engines (windows/macos/paddle/ra
dsh plugin add picturereaderLocal PDF, Office, image, and OCR document intelligence for DeepSeek Harness.
dsh plugin add dsh-docDeepSeek brain + automatic image transcription for DeepSeek Harness: a deepseek-vision provider route that transcribes attached images to text before delegating to the text-only DeepSeek adapter. Official deepseek-v4-flash-vision-exp by default (a pure-te
dsh plugin add dsh-vision-proxyLocal OCR fallback for DeepSeek Harness (Web): when the routed model cannot accept image input, a pasted image is saved locally and its text read by PP-OCRv5 + ONNX Runtime — fully offline, no vision model required. / DeepSeek Harness 本地 OCR 兜底插件(Web):当接入
dsh plugin add dsh-ocr-localEyes for text-only DeepSeek on DeepSeek Harness: image understanding via Doubao Web by default (zero cost, no API key), Antigravity IDE quota (flash/pro), Gemini API, or Cockpit proxy — model-invokable vision tool, wrapper adapters, evidence memory, and a
dsh plugin add dsh-vision-webGenerate and display images inline in the DSH chat via API channels or local CLIs (mmx / codex / agy), and read images into structured JSON evidence (OCR / layout / semantics) on any model — with backend probing and a bundled recovery skill.
dsh plugin add dsh-chat-imagineAuditable vision and cross-platform Computer Use runtime for DeepSeek Harness with source-preserving evidence.
dsh plugin add @dttxorg/deepseekeyesOn-demand vision for text-only DSH sessions: images become markers, and a vision_describe tool sends only image + question to an OpenAI-compatible vision model
dsh plugin add @dsh-extension/dsh-vision-bridgedsh-mineru: MinerU 文档解析插件 for DeepSeek Harness — 多模态全格式 (PDF/Word/PPT/Excel/HTML/图片) → 结构化 Markdown. 填 Token 走精准解析 API, 留空走 Agent 轻量解析 API.
dsh plugin add dsh-mineruDeepSeek Harness 原生视觉 Bundle:粘贴或拖入图片,通过托管的 deepseek-vision-mcp 调用 OpenAI 兼容视觉模型。
dsh plugin add dsh-plugin-deepseek-visionDeepSeek Harness native plugin: vision capabilities for text-only LLMs (describe, OCR, VQA, layout analysis, clipboard)
dsh plugin add dsh-plugin-deepeyeWorkspace-bound arbitrary file upload, reading, OCR, and rendering for DeepSeek Harness.
dsh plugin add dsh-open-fileEyes for text-only DeepSeek on DeepSeek Harness: image understanding via Antigravity IDE quota (default, flash/pro), OpenAI-compatible VLM endpoints, Gemini API, or local Ollama — model-invokable vision tool, wrapper adapters for deepseek/opencode-go, evi
dsh plugin add dsh-youreyesDeepSeek Harness plugin that turns UI screenshots into implementation-grade web specs using OCR, deterministic geometry, scene graphs, assets, and render comparison.
dsh plugin add dsh-ui-specAdaptive image routing for DeepSeek Harness: pass images directly to native multimodal models and transcribe them only for text-only models.
dsh plugin add dsh-vision-recognizerFlagship multimodal vision hub for DeepSeek Harness: ~40 tools, PDF drag-and-drop, LaTeX formulas, complex tables, QR codes, UI flow diagrams, and multi-model consensus.
dsh plugin add @goodandready/dsh-vision-bridgeGive DeepSeek Harness free eyes — natively. Paste or drop an image into the Web GUI composer: it appears as a thumbnail attachment (like a normal AI chat), and before the text-only DeepSeek model is called, the host automatically reads the image with the
dsh plugin add dsh-dseyesDeepSeek Harness plugin: a model-facing `vision` tool that describes and OCRs image files by calling the free Zhipu GLM vision API directly (no external CLI required).
dsh plugin add dsh-vision-free-eyesDeepSeek Harness out-of-tree plugin: give a text-only coding model eyes by routing images to a Qwen-VL (DashScope) route through ctx.llm and returning text.
dsh plugin add dsh-plugin-qwen-imageMicrosoft MarkItDown as a DeepSeek Harness tool: convert PDF, Word, Excel, PowerPoint, HTML, CSV, EPUB, or a URL into Markdown the model can read.
dsh plugin add dsh-markitdowndsh auto-vision bridge: paste an image into a text-only model's composer and a configured multimodal model transcribes it as text, automatically.
dsh plugin add @iroam2375/dsh-autovisionPoint-and-shoot screenshot capture for DeepSeek Harness: clipboard watcher + system floating window (comment & key-point, copy/save-doc/save-image) + zero-config OCR through the host's own multimodal model (Tongyi Qianwen fallback) + Obsidian per-day merg
dsh plugin add dsh-screenshot-captureVision toolkit for DeepSeek Harness -- glance, ground, detect, crop CLI tools for text-only agents to understand images
dsh plugin add dsh-plugin-vision-toolkitLightweight, route-preserving vision link for DSH: a configured vision model sees while the selected text model stays in control.
dsh plugin add dsh-vision-linkBridges images into text for DeepSeek Harness: when the session's selected model cannot see images, a configured vision model describes them and the descriptions enter the durable session history as folded context rows — your message stays untouched.
dsh plugin add dsh-auto-imageVisionForge: vision understanding + image generation for DeepSeek Harness (DSH)
dsh plugin add @lr611/visionforgeDeepSee: DeepSeek Harness vision, multi-model routing, and Gemini visual self-checks before delivery
dsh plugin add @chang416/deepseeMobileCode for the dsh web GUI: detect iOS/Android projects, run preview servers, and drive the simulator/emulator from the session — 41 agent tools (device_run, device_screen, device_ui_tree, device_scene, device_ui_rows, device_tap_row, device_tap_eleme
dsh plugin add dsh-mobilecodeFully-local OCR CLI for text-only LLMs: PaddleOCR-VL first-tier engine with tesseract fallback, dsh plugin included
dsh plugin add local-ocr-cliDrag-and-drop PDF, Office, images, and text/code files into DSH conversations. The host extracts (and OCRs) them into the prompt; attach_* tools cover notebook cells, PDF-page OCR, image describe, and save.
dsh plugin add @lucasxingg/dsh-file-attachdsh plugin: recognize attached images locally with Tesseract OCR and send only the recognized text to the model. Text models never receive image bytes; vision passthrough is opt-in.
dsh plugin add dsh-tesseract-ocrDSH 插件:通过硅基流动(SiliconFlow)视觉大模型识别/分析图片,支持本地文件路径、http(s) 图片 URL 与 base64 data URL。含持久化的粘贴识别面板(web 端)。
dsh plugin add dsh-siliconflow-visionLets a text-only agent call a multimodal model mid-task: the vision tool sends one image file to Qwen, Kimi, OpenAI, Claude, Gemini or a self-hosted endpoint and returns structured evidence — summary, verbatim OCR, layout, entities, and what the model cou
dsh plugin add dsh-plugin-visionDeepSeek Harness plugin: analyse images out of band with a vision model — pasted images are digested into text before admission and a describe_image tool covers image paths, all without ever changing the session's model.
dsh plugin add dsh-image-routerDeepSeek Harness plugin: content-aware PDF reading tools for vision models (pdf_scan / pdf_read_page / pdf_render_region). Detects figures (vector+raster), tables, math, and two-column layout per page, then routes text pages to extraction and figure/table
dsh plugin add dsh-pdf-readerPaste images into DeepSeek Harness chat and have them read by a free local backend (macOS Vision / Tesseract) before the text-only DeepSeek model answers
dsh plugin add dsh-ocr-bridgeLocal OCR for DeepSeek Harness: read text out of images with Apple Vision on macOS — no network, no API key, nothing to install. Falls back to tesseract elsewhere.
dsh plugin add dsh-ocr-freeLocal image understanding for the DeepSeek Harness: OCR, table-layout detection, and semantic description of textless images via macOS Vision, fully on-device (images never leave your Mac). 本地化识图,图片不出本机。
dsh plugin add @freespace8/dsh-free-visionBridge Apple's on-device Vision framework (macOS) into DeepSeek Harness: OCR, image classification, face detection, and document layout as local dsh tools. No network, no API key, no daemon.
dsh plugin add dsh-maclensGive text-only models eyes: an analyze_image tool for DeepSeek Harness, backed by free Chinese vision APIs (GLM-4V-Flash / Qwen-VL) or any OpenAI-compatible vision endpoint. 给纯文本模型装上眼睛。
dsh plugin add dsh-imgVision bundle + plugin for DeepSeek Harness: the describe_image tool reads local images and asks any configured multimodal route, with zero core changes
dsh plugin add dsh-bundle-vision--- description: "The out-of-tree ocr tool: recognizes a scanned PDF or image page by page, merges the pages in order only after the last one, and saves one Markdown document with the heading levels the model decided." kind: "package-reference" ---
dsh plugin add dsh-ab-ocrManaged document pipeline for the DeepSeek Harness: Markdown to DOCX/PDF, DOCX/PDF back to Markdown, and online sources as agent tools
dsh plugin add @yadsh/dsh-documentsDSH-Plugin for DeepSeek-Harness: fully-local image understanding & OCR powered by macOS Vision Framework
dsh plugin add @niyongsheng/free-vision-skillLocal-first vision for text-only DeepSeek Harness agents: route image understanding to a local OpenAI-compatible vision model, return structured JSON evidence (summary / OCR / layout / semantics / visual / uncertainty).
dsh plugin add dsh-vision-local