awesome-repositories.com
Blog
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectAboutHow we rankPressMCP server
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Norconex avatar

Norconex/collector-http

0
View on GitHub↗
202 stars·71 forks·Java·Apache-2.0·2 viewsopensource.norconex.com/crawlers↗

Collector Http

Norconex HTTP Collector

Features

  • Java Crawling Frameworks - Full-featured HTTP crawler with data storage capabilities.

Star history

Star history chart for norconex/collector-httpStar history chart for norconex/collector-http

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to Collector Http

Similar open-source projects, ranked by how many features they share with Collector Http.
  • crawlscript/webcollectorCrawlScript avatar

    CrawlScript/WebCollector

    3,091View on GitHub↗

    WebCollector is an open source web crawler framework based on Java.It provides some simple interfaces for crawling the Web,you can setup a multi-threaded web crawler in less than 5 minutes.

    Java
    View on GitHub↗3,091
  • digitalpebble/storm-crawlerDigitalPebble avatar

    DigitalPebble/storm-crawler

    980View on GitHub↗

    A scalable, mature and versatile web crawler based on Apache Storm

    Java
    View on GitHub↗980
  • internetarchive/heritrix3internetarchive avatar

    internetarchive/heritrix3

    3,246View on GitHub↗

    Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project. Heritrix (sometimes spelled heretrix, or misspelled or missaid as heratrix/heritix/heretix/heratix) is an archaic word for heiress (woman who inherits). Since our crawler seeks to collect…

    Java
    View on GitHub↗3,246
  • code4craft/webmagiccode4craft avatar

    code4craft/webmagic

    11,680View on GitHub↗

    Webmagic is a Java web crawling framework designed for building scalable automated crawlers to download and process large volumes of web pages. It functions as a distributed web crawler and dynamic content crawler, utilizing an XPath HTML parser to locate and extract specific data points from page structures. The framework distinguishes itself through its ability to handle dynamic content by rendering JavaScript and executing asynchronous requests to extract data from non-static pages. It also allows users to define and execute crawler logic via scripting languages, enabling the update of col

    Javacrawlerframeworkjava
    View on GitHub↗11,680
See all 12 alternatives to Collector Http→

Frequently asked questions

What does norconex/collector-http do?

Norconex HTTP Collector

What are the main features of norconex/collector-http?

The main features of norconex/collector-http are: Java Crawling Frameworks.

What are some open-source alternatives to norconex/collector-http?

Open-source alternatives to norconex/collector-http include: code4craft/webmagic — Webmagic is a Java web crawling framework designed for building scalable automated crawlers to download and process… crawlscript/webcollector — WebCollector is an open source web crawler framework based on Java.It provides some simple interfaces for crawling the… digitalpebble/storm-crawler — A scalable, mature and versatile web crawler based on Apache Storm. internetarchive/heritrix3 — Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project. Heritrix… pkwenda/webbee — 🐝 Web vertical crawler framework for fun. ssssssss-team/spider-flow — Spider-flow is a Java-based web crawling and data extraction platform that provides a centralized environment for…