awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetÀ proposNotre méthodologiePresseServeur MCP
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
xtuhcy avatar

xtuhcy/gecco

0
View on GitHub↗
2,513 stars·873 forks·Java·MIT·3 vues

Gecco

Gecco is a easy to use lightweight web crawler developed with java language.Gecco integriert jsoup, httpclient, fastjson, spring, htmlunit, redission ausgezeichneten framework,Let you only need to configure a number of jQuery style selector can be very quick to write a crawler.Gecco framework…

Features

  • Java Crawling Frameworks - Easy-to-use lightweight web crawler.

Historique des stars

Graphique de l'historique des stars pour xtuhcy/geccoGraphique de l'historique des stars pour xtuhcy/gecco

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Alternatives open source à Gecco

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec Gecco.
  • crawlscript/webcollectorAvatar de CrawlScript

    CrawlScript/WebCollector

    3,091Voir sur GitHub↗

    WebCollector is an open source web crawler framework based on Java.It provides some simple interfaces for crawling the Web,you can setup a multi-threaded web crawler in less than 5 minutes.

    Java
    Voir sur GitHub↗3,091
  • digitalpebble/storm-crawlerAvatar de DigitalPebble

    DigitalPebble/storm-crawler

    980Voir sur GitHub↗

    A scalable, mature and versatile web crawler based on Apache Storm

    Java
    Voir sur GitHub↗980
  • internetarchive/heritrix3Avatar de internetarchive

    internetarchive/heritrix3

    3,246Voir sur GitHub↗

    Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project. Heritrix (sometimes spelled heretrix, or misspelled or missaid as heratrix/heritix/heretix/heratix) is an archaic word for heiress (woman who inherits). Since our crawler seeks to collect…

    Java
    Voir sur GitHub↗3,246
  • code4craft/webmagicAvatar de code4craft

    code4craft/webmagic

    11,680Voir sur GitHub↗

    Webmagic is a Java web crawling framework designed for building scalable automated crawlers to download and process large volumes of web pages. It functions as a distributed web crawler and dynamic content crawler, utilizing an XPath HTML parser to locate and extract specific data points from page structures. The framework distinguishes itself through its ability to handle dynamic content by rendering JavaScript and executing asynchronous requests to extract data from non-static pages. It also allows users to define and execute crawler logic via scripting languages, enabling the update of col

    Javacrawlerframeworkjava
    Voir sur GitHub↗11,680
Voir les 12 alternatives à Gecco→

Questions fréquentes

Que fait xtuhcy/gecco ?

Gecco is a easy to use lightweight web crawler developed with java language.Gecco integriert jsoup, httpclient, fastjson, spring, htmlunit, redission ausgezeichneten framework,Let you only need to configure a number of jQuery style selector can be very quick to write a crawler.Gecco framework…

Quelles sont les fonctionnalités principales de xtuhcy/gecco ?

Les fonctionnalités principales de xtuhcy/gecco sont : Java Crawling Frameworks.

Quelles sont les alternatives open-source à xtuhcy/gecco ?

Les alternatives open-source à xtuhcy/gecco incluent : code4craft/webmagic — Webmagic is a Java web crawling framework designed for building scalable automated crawlers to download and process… crawlscript/webcollector — WebCollector is an open source web crawler framework based on Java.It provides some simple interfaces for crawling the… digitalpebble/storm-crawler — A scalable, mature and versatile web crawler based on Apache Storm. internetarchive/heritrix3 — Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project. Heritrix… norconex/collector-http — Norconex HTTP Collector. pkwenda/webbee — 🐝 Web vertical crawler framework for fun.