BigHugger
GH Repository · liustack

modlens

The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。

stars
3,977
30-day movement
+4415/day
Related entries
61
Connections
2
node/pnpmskill-packdshvisionimage-to-textvision-transformeragent-skillsclaude-codeclaude-skillscodexcordisdeepseekocrglmTypeScriptdsh-pluginopenclawtypescriptpi-agentharness-engineeringnodemultimodalhermes-agentharness

modlens is a skill-pack plugin that acts as a vision bridge for text-only coding agents such as DeepSeek, GLM, and other Harness-based agents. You paste an image and it returns structured JSON evidence covering OCR text, layout, and semantics.

It lets a text-only LLM agent read and reason about images it otherwise could not see.

Use it to

  • Extract OCR text from screenshots into JSON
  • Analyze UI screenshots for layout structure
  • Add vision capability to a text-only agent
  • Convert pasted images into semantic evidence for agent prompts

For Developers building or using text-only LLM coding agents

Role
skill-pack
Language
TypeScript
Licence
MIT
Forks
120
Open issues
3
Last push
2026-09-07
Latest release
v2.0.0 · 2026-08-05
Skills shipped
1
topicsvisionocragent-skillsmultimodaldeepseekimage-to-text