BigHugger
GH Repository · any4ai

AnyCrawl

AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Native multi-threading for bulk processing.

stars
3,459
30-day movement
+10/day
Related entries
60
Connections
1
node/pnpmaitoolscrawlscrapeserpragscrapingai-scrapingdockerhtml-to-markdownwebscraperTypeScriptdatanode

AnyCrawl is a Node.js/TypeScript crawler that converts websites into LLM-ready data, including HTML-to-Markdown conversion, and extracts structured search engine results pages (SERP) from engines such as Google, Bing and Baidu. It uses native multi-threading for bulk processing and is distributed as a Docker-based toolchain with Node and pnpm.

You need website content or structured SERP results in a form suitable for LLM and RAG pipelines.

Use it to

  • Crawl websites into LLM-ready markdown data
  • Extract structured SERP results from Google, Bing and Baidu
  • Bulk-process many pages with multi-threading
  • Feed crawled data into RAG pipelines
  • Self-host the crawler via the provided Docker toolchain

For Developers building scraping and RAG data pipelines

Role
rag
Language
TypeScript
Licence
MIT
Forks
366
Open issues
5
Last push
2026-09-15
Latest release
v0.0.1-alpha.1 · 2025-05-13
topicscrawlerscrapingserpraghtml-to-markdownnode