GH Repository · any4ai
AnyCrawl
AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Native multi-threading for bulk processing.
- stars
- 3,459
- 30-day movement
- +10/day
- Related entries
- 60
- Connections
- 1
node/pnpmaitoolscrawlscrapeserpragscrapingai-scrapingdockerhtml-to-markdownwebscraperTypeScriptdatanode
AnyCrawl is a Node.js/TypeScript crawler that converts websites into LLM-ready data, including HTML-to-Markdown conversion, and extracts structured search engine results pages (SERP) from engines such as Google, Bing and Baidu. It uses native multi-threading for bulk processing and is distributed as a Docker-based toolchain with Node and pnpm.
You need website content or structured SERP results in a form suitable for LLM and RAG pipelines.
Use it to
- Crawl websites into LLM-ready markdown data
- Extract structured SERP results from Google, Bing and Baidu
- Bulk-process many pages with multi-threading
- Feed crawled data into RAG pipelines
- Self-host the crawler via the provided Docker toolchain
For Developers building scraping and RAG data pipelines
- Role
- rag
- Language
- TypeScript
- Licence
- MIT
- Forks
- 366
- Open issues
- 5
- Last push
- 2026-09-15
- Latest release
- v0.0.1-alpha.1 · 2025-05-13
topicscrawlerscrapingserpraghtml-to-markdownnode