awesome-repositories.com
Blog
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectDespreCum realizăm clasamentulPresăServer MCP
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
ViDA-NYU avatar

ViDA-NYU/ache

0
View on GitHub↗
484 stele·136 fork-uri·Java·Apache-2.0·2 vizualizăriache.readthedocs.io↗

Ache

ACHE is a focused web crawler. It collects web pages that satisfy some specific criteria, e.g., pages that belong to a given domain or that contain a user-specified pattern. ACHE differs from generic crawlers in sense that it uses page classifiers to distinguish between relevant and irrelevant…

Features

  • Java Crawling Frameworks - Domain-specific web crawler for focused search.

Istoric stele

Graficul istoricului de stele pentru vida-nyu/acheGraficul istoricului de stele pentru vida-nyu/ache

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Start searching with AI

Alternative open-source pentru Ache

Proiecte open-source similare, clasificate după numărul de funcționalități comune cu Ache.
  • crawlscript/webcollectorAvatar CrawlScript

    CrawlScript/WebCollector

    3,091Vezi pe GitHub↗

    WebCollector is an open source web crawler framework based on Java.It provides some simple interfaces for crawling the Web,you can setup a multi-threaded web crawler in less than 5 minutes.

    Java
    Vezi pe GitHub↗3,091
  • digitalpebble/storm-crawlerAvatar DigitalPebble

    DigitalPebble/storm-crawler

    980Vezi pe GitHub↗

    A scalable, mature and versatile web crawler based on Apache Storm

    Java
    Vezi pe GitHub↗980
  • internetarchive/heritrix3Avatar internetarchive

    internetarchive/heritrix3

    3,246Vezi pe GitHub↗

    Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project. Heritrix (sometimes spelled heretrix, or misspelled or missaid as heratrix/heritix/heretix/heratix) is an archaic word for heiress (woman who inherits). Since our crawler seeks to collect…

    Java
    Vezi pe GitHub↗3,246
  • code4craft/webmagicAvatar code4craft

    code4craft/webmagic

    11,680Vezi pe GitHub↗

    Webmagic is a Java web crawling framework designed for building scalable automated crawlers to download and process large volumes of web pages. It functions as a distributed web crawler and dynamic content crawler, utilizing an XPath HTML parser to locate and extract specific data points from page structures. The framework distinguishes itself through its ability to handle dynamic content by rendering JavaScript and executing asynchronous requests to extract data from non-static pages. It also allows users to define and execute crawler logic via scripting languages, enabling the update of col

    Javacrawlerframeworkjava
    Vezi pe GitHub↗11,680
Vezi toate cele 12 alternative pentru Ache→

Întrebări frecvente

Ce face vida-nyu/ache?

ACHE is a focused web crawler. It collects web pages that satisfy some specific criteria, e.g., pages that belong to a given domain or that contain a user-specified pattern. ACHE differs from generic crawlers in sense that it uses page classifiers to distinguish between relevant and irrelevant…

Care sunt principalele funcționalități ale vida-nyu/ache?

Principalele funcționalități ale vida-nyu/ache sunt: Java Crawling Frameworks.

Care sunt câteva alternative open-source pentru vida-nyu/ache?

Alternativele open-source pentru vida-nyu/ache includ: code4craft/webmagic — Webmagic is a Java web crawling framework designed for building scalable automated crawlers to download and process… crawlscript/webcollector — WebCollector is an open source web crawler framework based on Java.It provides some simple interfaces for crawling the… digitalpebble/storm-crawler — A scalable, mature and versatile web crawler based on Apache Storm. internetarchive/heritrix3 — Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project. Heritrix… norconex/collector-http — Norconex HTTP Collector. pkwenda/webbee — 🐝 Web vertical crawler framework for fun.