BigHugger
GH Repository · 0xMassi

webclaw

Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.

stars
2,351
30-day movement
+186/day
Related entries
61
Connections
4
Rustjina-alternativeapify-alternativetls-fingerprintingdockerai-agentsmarkdownscraperapi-alternativeweb-scrapingmcp-servercrawl4ai-alternativellmrustself-hostedfirecrawl-alternativeweb-extractionclimcpweb-crawlerscrapingbee-alternativeai-scrapinghtml-to-markdown

webclaw is a Rust-based, local-first web content extraction tool for LLMs that scrapes, crawls, and extracts structured data. It ships as a CLI, a REST API, and an MCP server, converting web content (e.g., HTML to Markdown).

Reach for it when you want self-hosted web scraping and extraction feeding LLMs or agents without relying on hosted scraping services.

Use it to

  • Serve web extraction to LLM clients via MCP
  • Scrape and crawl sites from the CLI
  • Run a self-hosted REST extraction API
  • Convert HTML pages to Markdown for LLM context

For Developers building LLM agents needing self-hosted web extraction

Role
mcp-server
Language
Rust
Licence
AGPL-3.0
Forks
231
Open issues
2
Last push
2026-09-14
Latest release
v0.1.0 · 2026-03-24
topicsweb-scrapingmcp-serverllmrustcliself-hosted