agent-vision-toolkit
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
- stars
- 1,207
- 30-day movement
- +114/day
- Related entries
- 61
- Connections
- 1
agent-vision-toolkit is a Python vision toolkit and skill set that lets text-only LLMs work with images. It covers image Q&A, multi-image understanding, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional integration into agents like Codex, Claude Code, Pi, Oh My Pi, and OpenCode.
Use it when your text-only model needs to see — answering questions about images, reading long screenshots, or driving a GUI.
Use it to
- Answer questions about pasted images
- OCR long screenshots for text-only models
- Restore frontend UI from screenshots
- Automate GUI interactions via an agent
- Integrate vision into Codex or Claude Code
For Developers running text-only LLM agents that need vision
- Role
- agent-app
- Language
- Python
- Licence
- MIT
- Forks
- 46
- Open issues
- 5
- Last push
- 2026-08-27
- Latest release
- v0.1.0 · 2026-08-06
- Skills shipped
- 1