awesome-repositories.com
Blog
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectAboutHow we rankPressMCP server
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
internetarchive avatar

internetarchive/heritrix3

0
View on GitHub↗
3,246 stars·789 forks·Java·5 viewsheritrix.readthedocs.io↗

Heritrix3

Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project. Heritrix (sometimes spelled heretrix, or misspelled or missaid as heratrix/heritix/heretix/heratix) is an archaic word for heiress (woman who inherits). Since our crawler seeks to collect…

Features

  • Java Crawling Frameworks - Extensible, web-scale, archival-quality crawler.

Star history

Star history chart for internetarchive/heritrix3Star history chart for internetarchive/heritrix3

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to Heritrix3

Similar open-source projects, ranked by how many features they share with Heritrix3.
  • crawlscript/webcollectorCrawlScript avatar

    CrawlScript/WebCollector

    3,091View on GitHub↗

    WebCollector is an open source web crawler framework based on Java.It provides some simple interfaces for crawling the Web,you can setup a multi-threaded web crawler in less than 5 minutes.

    Java
    View on GitHub↗3,091
  • digitalpebble/storm-crawlerDigitalPebble avatar

    DigitalPebble/storm-crawler

    980View on GitHub↗

    A scalable, mature and versatile web crawler based on Apache Storm

    Java
    View on GitHub↗980
  • norconex/collector-httpNorconex avatar

    Norconex/collector-http

    202View on GitHub↗

    Norconex HTTP Collector

    Java
    View on GitHub↗202
  • code4craft/webmagiccode4craft avatar

    code4craft/webmagic

    11,680View on GitHub↗

    Webmagic is a Java web crawling framework designed for building scalable automated crawlers to download and process large volumes of web pages. It functions as a distributed web crawler and dynamic content crawler, utilizing an XPath HTML parser to locate and extract specific data points from page structures. The framework distinguishes itself through its ability to handle dynamic content by rendering JavaScript and executing asynchronous requests to extract data from non-static pages. It also allows users to define and execute crawler logic via scripting languages, enabling the update of col

    Javacrawlerframeworkjava
    View on GitHub↗11,680
See all 12 alternatives to Heritrix3→

Frequently asked questions

What does internetarchive/heritrix3 do?

Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project. Heritrix (sometimes spelled heretrix, or misspelled or missaid as heratrix/heritix/heretix/heratix) is an archaic word for heiress (woman who inherits). Since our crawler seeks to collect…

What are the main features of internetarchive/heritrix3?

The main features of internetarchive/heritrix3 are: Java Crawling Frameworks.

What are some open-source alternatives to internetarchive/heritrix3?

Open-source alternatives to internetarchive/heritrix3 include: code4craft/webmagic — Webmagic is a Java web crawling framework designed for building scalable automated crawlers to download and process… crawlscript/webcollector — WebCollector is an open source web crawler framework based on Java.It provides some simple interfaces for crawling the… digitalpebble/storm-crawler — A scalable, mature and versatile web crawler based on Apache Storm. norconex/collector-http — Norconex HTTP Collector. pkwenda/webbee — 🐝 Web vertical crawler framework for fun. ssssssss-team/spider-flow — Spider-flow is a Java-based web crawling and data extraction platform that provides a centralized environment for…