GH Repository · adbar
trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
- stars
- 6,827
- 30-day movement
- +176/day
- Related entries
- 60
- Connections
- 1
pythonllmnews-aggregatorcorpus-toolshtml-to-markdownrss-feedscrapingteiragweb-scrapingnlphtml2textcorpus-buildertext-miningPythontext-extractiontext-cleaningcrawlertext-preprocessingnews-crawlerarticle-extractorreadability
Trafilatura is a Python library and command-line tool for gathering text and metadata from the web, covering crawling, scraping, and main-content extraction. It outputs results as CSV, JSON, HTML, Markdown, TXT, or XML.
Use it when you need to turn web pages into clean extracted text and metadata for downstream processing.
Use it to
- Extract main article text from web pages
- Crawl sites to build text corpora
- Convert HTML to Markdown or plain text
- Feed extracted text into RAG pipelines
- Aggregate news from RSS feeds
For Developers building scraping, corpus, or RAG pipelines
- Role
- rag
- Language
- Python
- Licence
- Apache-2.0
- Forks
- 431
- Open issues
- 53
- Last push
- 2026-09-11
- Latest release
- v0.1.0 · 2019-09-25
topicsweb-scrapingtext-extractioncrawlerragcorpus-buildingpython