For a residential proxy service for web scraping, the strongest matches are speedyapply/jobspy (Jobspy is a job board scraper that includes proxy), ultrafunkamsterdam/undetected-chromedriver (Undetected-chromedriver is a patched browser driver framework designed specifically) and any4ai/anycrawl (AnyCrawl is an open-source automated web crawler and headless). mendableai/firecrawl and firecrawl/firecrawl round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
Find the best web scraping proxies to bypass blocks. Compare top-rated tools by features, reliability, and cost to pick the right one for your project.
JobSpy is a job board scraper and listing aggregator designed to extract employment opportunities from multiple websites and compile them into a unified dataset. It functions as a job search automation tool that programmatically collects vacancies based on keywords, locations, and specific filters. The project serves as a web scraping framework that utilizes proxy routing and user-agent rotation to bypass rate limits and avoid server-side blocking during data extraction. It includes infrastructure for concurrent request aggregation and schema-based data normalization to ensure consistent form
Jobspy is a job board scraper that includes proxy rotation and user-agent rotation to avoid blocking, but it is specialized for job listings and lacks headless browser support and CAPTCHA solving, so it only partially fits your general web scraping needs.
Undetected-chromedriver is a framework for automated browser navigation designed to bypass anti-bot security measures. It functions by patching browser drivers at the binary level to obscure automation signals, allowing scripts to interact with protected websites without being flagged or blocked by security services. The project distinguishes itself through its ability to maintain stealth during automated sessions, including those executed in headless mode. It achieves this by injecting custom configurations to mimic human user behavior and by hooking into low-level browser debugging protocol
Undetected-chromedriver is a patched browser driver framework designed specifically to evade anti-bot security measures during automated scraping, making it a solid fit for your anti-detection needs — but it focuses primarily on stealth and headless support, so extra tooling is needed for proxy rotation, throttling, and CAPTCHA solving.
AnyCrawl is an AI-powered data extractor, automated web crawler, and headless browser orchestrator. It serves as a web content extraction API and a gateway that connects crawling and scraping tools to language models using a standardized API protocol. The project specializes in converting unstructured website content into structured JSON or markdown optimized for AI assistants. It utilizes language models and JSON schemas to pull specific information into validated formats and provides capabilities for AI page summarization and LLM-optimized content extraction. The system manages comprehensi
AnyCrawl is an open-source automated web crawler and headless browser orchestrator with per-request proxy assignments and exponential backoff retries, which provides some anti-bot capabilities; however, it lacks explicit user-agent rotation, CAPTCHA solving, and stealth plugins, so it fits the query as a matching tool but with a narrower set of anti-bot features.
Firecrawl is a headless browser automation tool and web crawling engine designed to extract structured data from the web. It functions as an API that transforms raw website content and documents into clean markdown and JSON formats to serve as context for large language models. The project distinguishes itself by using natural language prompts to translate human instructions into targeted data extraction tasks and browser actions. It can execute interactive page navigation, such as clicking and scrolling, and perform automated web research to retrieve structured data without manual interventi
Firecrawl is a headless browser automation and web scraping engine that includes proxy and fingerprint rotation for anti-detection, making it a relevant web scraping framework with anti-bot capabilities, though it may not cover every listed feature like CAPTCHA solving.
Firecrawl is a web data extraction platform designed to convert unstructured web content into clean, LLM-ready formats like markdown or JSON. It functions as an autonomous web crawler and scraper, capable of mapping entire domains, performing recursive navigation, and executing complex data gathering tasks. By leveraging headless browser orchestration, the system handles dynamic, JavaScript-heavy pages to ensure comprehensive data capture. The platform distinguishes itself through its focus on agentic workflows, providing a programmatic interface that allows autonomous agents to perform live
Firecrawl is an open-source web scraping and crawling platform with headless browser support for dynamic pages, making it a relevant scraping tool, but it does not prominently advertise anti-bot features like proxy rotation or CAPTCHA solving, so it may need extra configuration to fully match that requirement.
SeleniumBase is a Python-based framework designed for end-to-end web application testing and automated web scraping. It provides a unified interface for browser orchestration, managing browser lifecycles, and executing complex interaction sequences across multiple browser vendors and operating systems. The framework simplifies the development of automation workflows by handling driver provisioning, element synchronization, and project scaffolding. The project distinguishes itself through specialized stealth configurations that modify browser fingerprints to bypass anti-bot detection mechanism
SeleniumBase is a Python framework for automated web scraping that includes specialized stealth configurations to bypass anti-bot detection and supports headless browsers, but it does not natively provide proxy rotation, request throttling, or CAPTCHA solving.
Stealth is a distributed automation engine designed for headless browser orchestration, web scraping, and decentralized network coordination. It functions as a peer-to-peer framework that enables multiple nodes to align automated tasks, share cached web resources, and route traffic through a decentralized proxy architecture. By abstracting the underlying browser implementation, the project provides a unified interface for executing navigation and data extraction workflows across diverse environments. The platform distinguishes itself through its ability to synchronize browser sessions and aut
Stealth is an automatable web browser and scraper with built-in proxy support and a focus on privacy and anti-detection, which fits the need for a scraping tool that avoids blocks, though its exact support for features like CAPTCHA solving or user-agent rotation is not detailed in its description.
This project is a Python web scraping library and automated data collection suite. It provides tools for extracting structured data from websites, implementing web crawlers to navigate site links, and parsing HTML DOM structures to isolate specific elements and attributes. The toolkit includes a pipeline for processing unstructured text and cleaning raw web content to extract meaningful information. It also features capabilities for image data extraction and the integration of external APIs to retrieve structured data from remote endpoints. The system covers broad capability areas including
This is a web scraping library that includes proxy routing, headless browser automation, and anti-detection capabilities, making it fit the search for a scraping tool with anti-bot features, though specific measures like user-agent rotation and CAPTCHA solving aren't clearly confirmed.
Skyvern is an autonomous web navigation agent and browser-based workflow orchestrator that uses large language models to execute multi-step tasks on websites. By translating natural language instructions into actionable browser commands, the framework enables the automation of complex user workflows, including data extraction and interface interaction, without manual intervention. The platform distinguishes itself through a focus on secure, self-hosted infrastructure and stealth-oriented execution. It utilizes containerized browser isolation to maintain consistent environments and employs pro
Skyvern is an open-source browser automation framework that incorporates anti-bot evasion, proxy support, and stealth navigation, making it a capable tool for web scraping with anti-detection features.
This project is an MCP browser automation server that connects large language models to headless cloud browsers. It functions as an autonomous web workflow engine and an LLM web agent interface, enabling the translation of natural language instructions into browser actions and structured data retrieval. The system distinguishes itself through a managed headless browser cloud API that supports concurrent Chromium sessions with integrated stealth modes, CAPTCHA solving, and proxy traffic routing. It utilizes self-healing element selection to maintain automation resilience when page structures c
This is a browser automation server that provides headless Chromium sessions with integrated stealth modes, CAPTCHA solving, and proxy routing, which directly supports web scraping with anti-bot capabilities—though it is primarily designed for LLM agents rather than being a general-purpose scraping framework, it covers the core anti-detection features you need.
| Repositorio | Estrellas | Lenguaje | Licencia | Último push |
|---|---|---|---|---|
| speedyapply/jobspy | 3.7K | Python | MIT | |
| ultrafunkamsterdam/undetected-chromedriver | 12.4K | Python | gpl-3.0 | |
| any4ai/anycrawl | 2.7K | TypeScript | mit | |
| mendableai/firecrawl | 139.4K | TypeScript | AGPL-3.0 | |
| firecrawl/firecrawl | 133.5K | TypeScript | AGPL-3.0 | |
| seleniumbase/seleniumbase | 12.4K | Python | mit | |
| tholian-network/stealth | 1.1K | JavaScript | GPL-3.0 | |
| remitchell/python-scraping | 4.7K | Jupyter Notebook | — | |
| skyvern-ai/skyvern | 21.9K | Python | AGPL-3.0 | |
| browserbase/mcp-server-browserbase | 3.1K | TypeScript | apache-2.0 |