GH Repository · liustack
modlens
The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。
- stars
- 3,977
- 30-day movement
- +4415/day
- Related entries
- 61
- Connections
- 2
node/pnpmskill-packdshvisionimage-to-textvision-transformeragent-skillsclaude-codeclaude-skillscodexcordisdeepseekocrglmTypeScriptdsh-pluginopenclawtypescriptpi-agentharness-engineeringnodemultimodalhermes-agentharness
modlens is a skill-pack plugin that acts as a vision bridge for text-only coding agents such as DeepSeek, GLM, and other Harness-based agents. You paste an image and it returns structured JSON evidence covering OCR text, layout, and semantics.
It lets a text-only LLM agent read and reason about images it otherwise could not see.
Use it to
- Extract OCR text from screenshots into JSON
- Analyze UI screenshots for layout structure
- Add vision capability to a text-only agent
- Convert pasted images into semantic evidence for agent prompts
For Developers building or using text-only LLM coding agents
- Role
- skill-pack
- Language
- TypeScript
- Licence
- MIT
- Forks
- 120
- Open issues
- 3
- Last push
- 2026-09-07
- Latest release
- v2.0.0 · 2026-08-05
- Skills shipped
- 1
topicsvisionocragent-skillsmultimodaldeepseekimage-to-text