awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
DigitalPebble avatar

DigitalPebble/storm-crawler

0
View on GitHub↗
980 stars·281 forks·Java·Apache-2.0·7 viewsstormcrawler.apache.org↗

Storm Crawler

A scalable, mature and versatile web crawler based on Apache Storm

Features

  • Java Crawling Frameworks - Low-latency, scalable crawler built on Apache Storm.
  • Streaming Applications - Web crawler SDK based on Apache Storm.

Star history

Star history chart for digitalpebble/storm-crawlerStar history chart for digitalpebble/storm-crawler

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Storm Crawler

These projects share indexed features with Storm Crawler. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • code4craft/webmagiccode4craft avatar

    code4craft/webmagic

    11,680View on GitHub↗

    Webmagic is a Java web crawling framework designed for building scalable automated crawlers to download and process large volumes of web pages. It functions as a distributed web crawler and dynamic content crawler, utilizing an XPath HTML parser to locate and extract specific data points from page structures. The framework distinguishes itself through its ability to handle dynamic content by rendering JavaScript and executing asynchronous requests to extract data from non-static pages. It also allows users to define and execute crawler logic via scripting languages, enabling the update of col

    Javacrawlerframeworkjava
    View on GitHub↗11,680
  • crawlscript/webcollectorCrawlScript avatar

    CrawlScript/WebCollector

    3,091View on GitHub↗

    WebCollector is an open source web crawler framework based on Java.It provides some simple interfaces for crawling the Web,you can setup a multi-threaded web crawler in less than 5 minutes.

    Java
    View on GitHub↗3,091
  • internetarchive/heritrix3internetarchive avatar

    internetarchive/heritrix3

    3,246View on GitHub↗

    Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project. Heritrix (sometimes spelled heretrix, or misspelled or missaid as heratrix/heritix/heretix/heratix) is an archaic word for heiress (woman who inherits). Since our crawler seeks to collect…

    Java
    View on GitHub↗3,246
  • aklivity/zillaaklivity avatar

    aklivity/zilla

    690View on GitHub↗

    🦎 A multi-protocol edge & service proxy. Seamlessly interface web apps, IoT clients, & microservices to Apache Kafka® via declaratively defined, stateless APIs.

    Java
    View on GitHub↗690
Compare all 15 related projects→

Frequently asked questions

What does digitalpebble/storm-crawler do?

A scalable, mature and versatile web crawler based on Apache Storm

What are the main features of digitalpebble/storm-crawler?

The main features of digitalpebble/storm-crawler are: Java Crawling Frameworks, Streaming Applications.

Which projects share features with digitalpebble/storm-crawler?

Projects with overlapping indexed features include: code4craft/webmagic — Webmagic is a Java web crawling framework designed for building scalable automated crawlers to download and process… crawlscript/webcollector — WebCollector is an open source web crawler framework based on Java.It provides some simple interfaces for crawling the… internetarchive/heritrix3 — Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project. Heritrix… javactrl/javactrl-kafka — Distributed, Scalable, Fault-Tolerant, Minimalistic Workflow Engine - No DAGs, No YAML, No Cumbersome Diagrams, Just… norconex/collector-http — Norconex HTTP Collector. aklivity/zilla — 🦎 A multi-protocol edge & service proxy. Seamlessly interface web apps, IoT clients, & microservices to Apache Kafka®…