BigHugger
GH Repository · adbar

trafilatura

Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML

stars
6,827
30-day movement
+176/day
Related entries
60
Connections
1
pythonllmnews-aggregatorcorpus-toolshtml-to-markdownrss-feedscrapingteiragweb-scrapingnlphtml2textcorpus-buildertext-miningPythontext-extractiontext-cleaningcrawlertext-preprocessingnews-crawlerarticle-extractorreadability

Trafilatura is a Python library and command-line tool for gathering text and metadata from the web, covering crawling, scraping, and main-content extraction. It outputs results as CSV, JSON, HTML, Markdown, TXT, or XML.

Use it when you need to turn web pages into clean extracted text and metadata for downstream processing.

Use it to

  • Extract main article text from web pages
  • Crawl sites to build text corpora
  • Convert HTML to Markdown or plain text
  • Feed extracted text into RAG pipelines
  • Aggregate news from RSS feeds

For Developers building scraping, corpus, or RAG pipelines

Role
rag
Language
Python
Licence
Apache-2.0
Forks
431
Open issues
53
Last push
2026-09-11
Latest release
v0.1.0 · 2019-09-25
topicsweb-scrapingtext-extractioncrawlerragcorpus-buildingpython