How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
Crawlab is a distributed web scraping platform designed to centralize the management, deployment, and execution of large-scale data extraction tasks. It functions as a control plane that orchestrates scraping scripts and automated workflows across multiple nodes, providing a unified environment for managing complex data collection operations. The platform distinguishes itself through a distributed architecture that coordinates worker nodes via a central master, utilizing real-time communication to maintain oversight of all active processes. It ensures operational consistency by isolating task
Geziyor, blazing fast web crawling & scraping framework for Go. Supports JS rendering.
Colly is a high-performance web scraping framework designed for the automated extraction of structured data from websites. It provides a programmable toolkit that manages the complexities of large-scale data collection, including concurrent request orchestration, automatic cookie handling, and robots.txt compliance. By utilizing an asynchronous execution model, the engine maintains high throughput while preventing resource exhaustion during recursive or distributed crawling tasks. The framework is distinguished by its modular, event-driven architecture, which allows developers to hook into sp
Antch, a fast, powerful and extensible web crawling & scraping framework for Go
A Unix-style personal search engine and web crawler for your digital footprint.
The main features of amirgamil/apollo are: Web Crawlers.
Projects with overlapping indexed features include: crawlab-team/crawlab — Crawlab is a distributed web scraping platform designed to centralize the management, deployment, and execution of… geziyor/geziyor — Geziyor, blazing fast web crawling & scraping framework for Go. Supports JS rendering. gocolly/colly — Colly is a high-performance web scraping framework designed for the automated extraction of structured data from… henrylee2cn/pholcus — Pholcus is a distributed web crawler framework written in Go designed for high-concurrency data extraction. It… jina-ai/reader — Reader is an AI data ingestion pipeline and web content parser designed to convert websites and documents into clean… antchfx/antch — Antch, a fast, powerful and extensible web crawling & scraping framework for Go.