9 open-source projects similar to a11ywatch/crawler, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.
The purpose of this project is to provide a nice DSL wrapper around the cumbersome htmlunit Java library. Here is an example taken from a unit test in this package:
scrala is a web crawling framework for scala, which is inspired by scrapy.
Nov 20 2017 -- A distributed open source search engine and spider/crawler written in C/C++ for Linux on Intel/AMD. From gigablast dot com, which has binaries for download. See the README.md file at the very bottom of this page for instructions.
EBOT (http://www.redaelli.org/matteo-blog/projects/ebot/) USAGE: see the wiki homepage (http://wiki.github.com/matteoredaelli/ebot/) and ebot_test.erl for details LICENSE: GPL v3+, see file LICENSE Copyright (c) 2010, 2011 - Matteo Redaelli at libero dot it
Web::Scraper - Web Scraping Toolkit using HTML and CSS Selectors or XPath expressions
Website | Guides | API Docs | Examples | Discord
HTTrack is a website mirroring tool that downloads entire websites to a local directory, preserving the relative link structure so the site remains fully navigable without an internet connection. It operates as an HTTP website downloader and offline browsing utility, capturing static copies of remote sites for long-term offline storage and archiving. The tool includes a resumable download state machine that persists progress, allowing interrupted transfers to continue without re-fetching already retrieved files. It also features an incremental synchronization engine that compares local and re