Use case · decision ranking

Data normalization pipelines

Normalize extracted or heterogeneous data for downstream systems.

10 reviewed matches, ranked by fit and deterministic project health.

#1

firecrawl

firecrawl

79Health
Editorial

firecrawl/firecrawl: The context API to search, scrape, and interact with the web at scale. 🔥.

Editorial
88% fitExtraction and transformation tooling can fit normalization before downstream analytics, search, or automation.Compare this repository →
171.9KTypeScriptAGPL-3.0data-extractionweb-scraping
#2

soimort

you-get

78Health
Editorial

soimort/you-get: :arrow_double_down: Dumb downloader that scrapes the web.

Editorial
88% fitExtraction and transformation tooling can fit normalization before downstream analytics, search, or automation.Compare this repository →
56.9KPythonNOASSERTIONdata-extraction
#3

airbytehq

airbyte

76Health
Editorial

airbytehq/airbyte: Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.

Editorial
88% fitExtraction and transformation tooling can fit normalization before downstream analytics, search, or automation.Compare this repository →
21.9KPythonNOASSERTIONapidata-extractionworkflow-automation
#4

clips

pattern

76Health
Editorial

clips/pattern: Web mining module for Python, with tools for scraping, natural language processing, machine learning, network analysis and visualization.

Editorial
88% fitExtraction and transformation tooling can fit normalization before downstream analytics, search, or automation.Compare this repository →
8.9KPythonBSD-3-Clausedata-extraction
#5

pathwaycom

pathway

75Health
Editorial

pathwaycom/pathway: Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.

Editorial
88% fitExtraction and transformation tooling can fit normalization before downstream analytics, search, or automation.Compare this repository →
62.4KPythonNOASSERTIONdata-extractionrag
#6

MontFerret

ferret

75Health
Editorial

MontFerret/ferret: Declarative data automation language and Go runtime for structured extraction workflows.

Editorial
88% fitExtraction and transformation tooling can fit normalization before downstream analytics, search, or automation.Compare this repository →
6KGoApache-2.0data-extractionweb-scraping
#7
75Health
Editorial

adbar/trafilatura: Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML.

Editorial
88% fitExtraction and transformation tooling can fit normalization before downstream analytics, search, or automation.Compare this repository →
6.7KPythonApache-2.0data-extraction
#8

alirezamika

autoscraper

71Health
Editorial

alirezamika/autoscraper: A Smart, Automatic, Fast and Lightweight Web Scraper for Python.

Editorial
88% fitExtraction and transformation tooling can fit normalization before downstream analytics, search, or automation.Compare this repository →
7.9KPythonMITdata-extractionweb-scraping
#9
60Health
Editorial

dzhng/deep-research: An AI-powered research assistant that performs iterative, deep research on any topic by combining search engines, web scraping, and large language models. The goal of this repo is to provide the simplest implementation of a deep research agent - e.g. an agent that can refine its research direction overtime and deep dive into a topic.

Editorial
88% fitExtraction and transformation tooling can fit normalization before downstream analytics, search, or automation.Compare this repository →
19.6KTypeScriptMITai-agentdata-extraction
#10

brightdata

cli

60Health
Editorial

brightdata/cli: Official Bright Data CLI - scrape, search, and extract structured web data directly from your terminal.

Editorial
88% fitExtraction and transformation tooling can fit normalization before downstream analytics, search, or automation.Compare this repository →
6.4KTypeScriptMITdata-extractionweb-scraping
Data normalization pipelines | ThingsO