GH Repository · 0xMassi
webclaw
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
- stars
- 2,351
- 30-day movement
- +186/day
- Related entries
- 61
- Connections
- 4
Rustjina-alternativeapify-alternativetls-fingerprintingdockerai-agentsmarkdownscraperapi-alternativeweb-scrapingmcp-servercrawl4ai-alternativellmrustself-hostedfirecrawl-alternativeweb-extractionclimcpweb-crawlerscrapingbee-alternativeai-scrapinghtml-to-markdown
webclaw is a Rust-based, local-first web content extraction tool for LLMs that scrapes, crawls, and extracts structured data. It ships as a CLI, a REST API, and an MCP server, converting web content (e.g., HTML to Markdown).
Reach for it when you want self-hosted web scraping and extraction feeding LLMs or agents without relying on hosted scraping services.
Use it to
- Serve web extraction to LLM clients via MCP
- Scrape and crawl sites from the CLI
- Run a self-hosted REST extraction API
- Convert HTML pages to Markdown for LLM context
For Developers building LLM agents needing self-hosted web extraction
- Role
- mcp-server
- Language
- Rust
- Licence
- AGPL-3.0
- Forks
- 231
- Open issues
- 2
- Last push
- 2026-09-14
- Latest release
- v0.1.0 · 2026-03-24
topicsweb-scrapingmcp-serverllmrustcliself-hosted