awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to a11ywatch/crawler

Projects sharing features with Crawler

9 open-source projects similar to a11ywatch/crawler, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • bplawler/crawlerbplawler avatar

    bplawler/crawler

    148View on GitHub↗

    The purpose of this project is to provide a nice DSL wrapper around the cumbersome htmlunit Java library. Here is an example taken from a unit test in this package:

    Scala
    View on GitHub↗148
  • gaocegege/scralagaocegege avatar

    gaocegege/scrala

    113View on GitHub↗

    scrala is a web crawling framework for scala, which is inspired by scrapy.

    Scala
    View on GitHub↗113
  • gigablast/open-source-search-enginegigablast avatar

    gigablast/open-source-search-engine

    1,601View on GitHub↗

    Nov 20 2017 -- A distributed open source search engine and spider/crawler written in C/C++ for Linux on Intel/AMD. From gigablast dot com, which has binaries for download. See the README.md file at the very bottom of this page for instructions.

    C++
    View on GitHub↗1,601
  • hadley/rvesthadley avatar

    hadley/rvest

    1,520View on GitHub↗

    Simple web scraping for R

    R
    View on GitHub↗1,520
  • matteoredaelli/ebotmatteoredaelli avatar

    matteoredaelli/ebot

    330View on GitHub↗

    EBOT (http://www.redaelli.org/matteo-blog/projects/ebot/) USAGE: see the wiki homepage (http://wiki.github.com/matteoredaelli/ebot/) and ebot_test.erl for details LICENSE: GPL v3+, see file LICENSE Copyright (c) 2010, 2011 - Matteo Redaelli at libero dot it

    Erlang
    View on GitHub↗330

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • miyagawa/web-scrapermiyagawa avatar

    miyagawa/web-scraper

    102View on GitHub↗

    Web::Scraper - Web Scraping Toolkit using HTML and CSS Selectors or XPath expressions

    Perl
    View on GitHub↗102
  • reggoodwin/ferritR

    reggoodwin/ferrit

    0View on GitHub↗
    View on GitHub↗0
  • spider-rs/spiderspider-rs avatar

    spider-rs/spider

    2,559View on GitHub↗

    Website | Guides | API Docs | Examples | Discord

    Rust
    View on GitHub↗2,559
  • xroche/httrackxroche avatar

    xroche/httrack

    4,360View on GitHub↗

    HTTrack is a website mirroring tool that downloads entire websites to a local directory, preserving the relative link structure so the site remains fully navigable without an internet connection. It operates as an HTTP website downloader and offline browsing utility, capturing static copies of remote sites for long-term offline storage and archiving. The tool includes a resumable download state machine that persists progress, allowing interrupted transfers to continue without re-fetching already retrieved files. It also features an incremental synchronization engine that compares local and re

    C
    View on GitHub↗4,360