Scheduler
Coordinates crawl requests, priorities, and retries.
Repository intelligence
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation. In ThingsO it is evaluated as a web crawling and scraping framework.
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation. In ThingsO it is evaluated as a web crawling and scraping framework.
Reliable web collection requires crawling, request management, parsing, retries, throttling, and adaptation to diverse page structures.
Provide crawler/scraper primitives for fetching pages, scheduling requests, extracting structured data, and controlling crawl behavior.
The project is useful when teams need the web-scraping capability without building every supporting primitive from scratch.
The baseline architecture for this web-scraping project is interpreted from its product category, while concrete runtime, technology, code paths, commands, and deployment evidence are compiled from the current repository snapshot.
Crawler engine with request scheduling, fetch/browser adapters, parsing/extraction logic, and output pipelines.
inferred · 80% confidenceSeed requests enter a scheduler, pages are fetched, parsers extract items and additional links, and outputs flow to downstream storage or processing.
inferred · 82% confidenceState behavior depends on the selected runtime/deployment; inspect the project’s execution modules and persistence configuration for durable-state requirements.
inferred · 55% confidencePersistence requirements are workload/deployment specific unless explicitly established by a captured manifest/container document.
inferred · 52% confidenceConcurrency is implementation/runtime specific; verify worker, async or parallel execution settings before capacity planning.
inferred · 52% confidenceScale according to the runtime’s supported process/service model and validate shared state, model hardware and external rate limits before horizontal replication.
inferred · 52% confidenceCoordinates crawl requests, priorities, and retries.
Retrieves page content through HTTP or browser execution.
Transforms page content into structured records or follow-up links.
Primary language reported by the current GitHub repository snapshot.
knownDeclared project dependency associated with backend framework.
knownDeclared project dependency associated with browser automation.
knownDefines dependency, packaging or build metadata.
knownContainer build or compose configuration is present in repository evidence.
knownRepository CI configuration automates checks, builds or release tasks.
knownThe semantic codebase map is derived from the captured repository tree. Key visible areas include docs, packages, docs/examples, packages/cli, packages/core.
docsProject documentation.
packagesReusable packages/modules that compose the project.
docs/examplesUsage examples/reference implementations.
packages/cliCommand-line interface implementation.
packages/coreCore domain or execution logic.
packages/basic-crawler/srcPrimary implementation source code.
packages/basic-crawler/testAutomated tests.
packages/browser-crawler/srcPrimary implementation source code.
Not established from available evidence.
The README provides executable setup/run commands; a representative captured command is `npx crawlee create my-crawler`.
known · 80% confidencenpx crawlee create my-crawlernpm startnpm install crawlee playwrightnpm install crawlee@nextPackage script `build` runs `turbo run build --filter=./packages/* && node ./scripts/typescript_fixes.mjs`.
known · 90% confidencePackage script `test` runs `vitest run --silent`.
known · 88% confidencePackage script `lint` runs `oxlint packages test docs --tsconfig=tsconfig.json --type-aware`.
known · 90% confidenceNot established from available evidence.
unknown · 0% confidenceCaptured CI configuration is present for automated repository checks/build/release tasks.
known · 82% confidenceA captured contribution/development document describes project contribution expectations.
known · 80% confidenceNot established from available evidence.
unknown · 0% confidenceExtend with spiders/crawlers, request middleware, parsers, extraction rules, pipelines, or browser adapters.
inferred · 72% confidenceNot established from available evidence.
unknown · 0% confidenceStart with documented public APIs and the codebase extension/provider/integration paths identified by the semantic tree map.
inferred · 58% confidenceNot established from available evidence.
Not established from available evidence.
Captured container configuration establishes a container-based development or deployment path.
known · 86% confidenceProduction topology is deployment-specific; validate stateful services, worker/runtime boundaries and external dependencies before high-availability scale-out.
inferred · 54% confidencePersistence requirements are workload/deployment specific unless explicitly established by a captured manifest/container document.
inferred · 52% confidenceConfiguration is supplied through the project’s documented runtime/application settings; inspect README and captured configuration files for exact keys.
inferred · 62% confidenceScale according to the runtime’s supported process/service model and validate shared state, model hardware and external rate limits before horizontal replication.
inferred · 52% confidenceNot established from available evidence.
unknown · 0% confidenceNot established from available evidence.
unknown · 0% confidenceRecovery planning should cover persistent state, generated artifacts and external integration credentials; exact procedures are deployment-specific.
inferred · 50% confidenceResource requirements depend on workload and selected runtime/model; benchmark the intended production workload before sizing infrastructure.
inferred · 50% confidenceNot established from available evidence.
unknown · 0% confidenceNot established from available evidence.
unknown · 0% confidenceUse the project’s supported secret/configuration mechanism and keep service credentials outside source control.
inferred · 52% confidenceNot established from available evidence.
unknown · 0% confidenceNot established from available evidence.
unknown · 0% confidenceNot established from available evidence.
unknown · 0% confidenceData can leave the deployment when configured external APIs, model providers or remote sources are used; exact flows depend on user configuration.
inferred · 50% confidenceNot established from available evidence.
unknown · 0% confidenceestablished with strong public adoption signals
inferred · 84% confidenceMaintained under GitHub owner `apify`; detailed governance/decision rights are not fully established by the bounded evidence pack.
inferred · 62% confidenceGitHub reports SPDX license `Apache-2.0`; verify repository license text and dependency obligations for the intended use.
known · 90% confidenceNot established from available evidence.
editorial / chatgpt-gpt-5.6-sol-manual · 78% overall confidence