scrapy
scrapy
scrapy/scrapy: Scrapy, a fast high-level web crawling & scraping framework for Python.
Use case · decision ranking
Collect data repeatedly from authorized web sources.
11 reviewed matches, ranked by fit and deterministic project health.
scrapy
scrapy/scrapy: Scrapy, a fast high-level web crawling & scraping framework for Python.
firecrawl
firecrawl/firecrawl: The context API to search, scrape, and interact with the web at scale. 🔥.
apify
apify/crawlee: Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
D4Vinci
D4Vinci/Scrapling: 🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!.
getmaxun
getmaxun/maxun: 🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥.
MontFerret
MontFerret/ferret: Declarative data automation language and Go runtime for structured extraction workflows.
apify
apify/crawlee-python: Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
firecrawl
firecrawl/firecrawl-mcp-server: 🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
alirezamika
alirezamika/autoscraper: A Smart, Automatic, Fast and Lightweight Web Scraper for Python.
code4craft
code4craft/webmagic: A scalable web crawler framework for Java.
brightdata
brightdata/cli: Official Bright Data CLI - scrape, search, and extract structured web data directly from your terminal.