data-scraper-agent
Build a fully automated AI-powered data collection agent for any public source — job boards, prices, news, GitHub, sports, anything. Runs on a schedule, enriches data with a free LLM (Gemini Flash), stores results in Notion/Sheets/Supabase, and learns from user feedback. Runs 100% free on GitHub Actions. Use when the user wants to monitor, collect, or track any public data automatically.
- installs 8w
- 1,792
- 30-day movement
- starts with the next reading
- Related entries
- 5
- Connections
- 0
A skill that guides an AI coding agent through building a scheduled, AI-enriched data collection agent in Python for public sources like APIs, HTML pages, RSS feeds, and JS-rendered sites. It covers a ten-step workflow from source connectors through Gemini Flash enrichment, storage (Notion/Sheets/Supabase), feedback learning, and a GitHub Actions runner.
Reach for it when you need to scaffold a free, unattended pipeline that monitors or collects public data on a schedule.
Use it to
- Build a scheduled scraper for a public API or website
- Enrich collected records with Gemini Flash in batches
- Store results in Notion, Sheets, or Supabase
- Add a feedback loop that learns from user decisions
- Set up a free GitHub Actions cron runner
For Developers building automated public-data collection bots
- Host repository
- affaan-m/ECC
- Installs, lifetime
- 3,000
- Installs, 8 weeks
- 1,792
- Host stars
- 261k
- Host language
- JavaScript