awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目关于排名机制媒体报道MCP 服务器
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to internetarchive/heritrix3

Open-source alternatives to Heritrix3

12 open-source projects similar to internetarchive/heritrix3, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Heritrix3 alternative.

  • code4craft/webmagiccode4craft 的头像

    code4craft/webmagic

    11,680在 GitHub 上查看↗

    Webmagic is a Java web crawling framework designed for building scalable automated crawlers to download and process large volumes of web pages. It functions as a distributed web crawler and dynamic content crawler, utilizing an XPath HTML parser to locate and extract specific data points from page structures. The framework distinguishes itself through its ability to handle dynamic content by rendering JavaScript and executing asynchronous requests to extract data from non-static pages. It also allows users to define and execute crawler logic via scripting languages, enabling the update of col

    Javacrawlerframeworkjava
    在 GitHub 上查看↗11,680
  • crawlscript/webcollectorCrawlScript 的头像

    CrawlScript/WebCollector

    3,091在 GitHub 上查看↗

    WebCollector is an open source web crawler framework based on Java.It provides some simple interfaces for crawling the Web,you can setup a multi-threaded web crawler in less than 5 minutes.

    Java
    在 GitHub 上查看↗3,091
  • digitalpebble/storm-crawlerDigitalPebble 的头像

    DigitalPebble/storm-crawler

    980在 GitHub 上查看↗

    A scalable, mature and versatile web crawler based on Apache Storm

    Java
    在 GitHub 上查看↗980
  • norconex/collector-httpNorconex 的头像

    Norconex/collector-http

    202在 GitHub 上查看↗

    Norconex HTTP Collector

    Java
    在 GitHub 上查看↗202
  • pkwenda/webbeepkwenda 的头像

    pkwenda/webBee

    193在 GitHub 上查看↗

    🐝 Web vertical crawler framework for fun

    Java
    在 GitHub 上查看↗193

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Find more with AI search
  • ssssssss-team/spider-flowssssssss-team 的头像

    ssssssss-team/spider-flow

    11,277在 GitHub 上查看↗

    Spider-flow is a Java-based web crawling and data extraction platform that provides a centralized environment for managing automated information gathering. It functions as a no-code tool, allowing users to define complex data collection pipelines through a visual, drag-and-drop interface rather than manual programming. The platform distinguishes itself through a graph-based workflow orchestration system where users link discrete nodes to define navigation and parsing logic. It supports dynamic content crawling by integrating headless browsers to execute JavaScript and render page content that

    Javacrawlerjsoupspider
    在 GitHub 上查看↗11,277
  • uscdatascience/sparklerUSCDataScience 的头像

    USCDataScience/sparkler

    422在 GitHub 上查看↗

    A web crawler is a bot program that fetches resources from the web for the sake of building applications like search engines, knowledge bases, etc. Sparkler (contraction of Spark-Crawler) is a new web crawler that makes use of recent advancements in distributed computing and information…

    Java
    在 GitHub 上查看↗422
  • vida-nyu/acheViDA-NYU 的头像

    ViDA-NYU/ache

    484在 GitHub 上查看↗

    ACHE is a focused web crawler. It collects web pages that satisfy some specific criteria, e.g., pages that belong to a given domain or that contain a user-specified pattern. ACHE differs from generic crawlers in sense that it uses page classifiers to distinguish between relevant and irrelevant…

    Java
    在 GitHub 上查看↗484
  • xtuhcy/geccoxtuhcy 的头像

    xtuhcy/gecco

    2,513在 GitHub 上查看↗

    Gecco is a easy to use lightweight web crawler developed with java language.Gecco integriert jsoup, httpclient, fastjson, spring, htmlunit, redission ausgezeichneten framework,Let you only need to configure a number of jQuery style selector can be very quick to write a crawler.Gecco framework…

    Java
    在 GitHub 上查看↗2,513
  • yahoo/anthelionyahoo 的头像

    yahoo/anthelion

    2,832在 GitHub 上查看↗

    Anthelion is a Nutch plugin for focused crawling of semantic data. The project is an open-source project released under the Apache License 2.0.

    Java
    在 GitHub 上查看↗2,832
  • yasserg/crawler4jyasserg 的头像

    yasserg/crawler4j

    4,622在 GitHub 上查看↗

    Crawler4j is a multi-threaded Java web crawler and spider designed for high-volume web traversal and content extraction. It functions as a polite crawling framework that enables the discovery and indexing of HTML and binary content across multiple websites. The project distinguishes itself through a persistent crawling model that serializes session state to local storage, allowing the engine to resume indexing after a crash or interruption. It includes a politeness controller to regulate request frequency and delays, preventing server overloading and IP blocking. The system covers a broad ra

    Java
    在 GitHub 上查看↗4,622
  • zhegexiaohuozi/seimicrawlerzhegexiaohuozi 的头像

    zhegexiaohuozi/SeimiCrawler

    1,993在 GitHub 上查看↗

    SeimiCrawler

    Java
    在 GitHub 上查看↗1,993