awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to uscdatascience/sparkler

Open-source alternatives to Sparkler

12 open-source projects similar to uscdatascience/sparkler, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Sparkler alternative.

  • code4craft/webmagicالصورة الرمزية لـ code4craft

    code4craft/webmagic

    11,680عرض على GitHub↗

    Webmagic is a Java web crawling framework designed for building scalable automated crawlers to download and process large volumes of web pages. It functions as a distributed web crawler and dynamic content crawler, utilizing an XPath HTML parser to locate and extract specific data points from page structures. The framework distinguishes itself through its ability to handle dynamic content by rendering JavaScript and executing asynchronous requests to extract data from non-static pages. It also allows users to define and execute crawler logic via scripting languages, enabling the update of col

    Javacrawlerframeworkjava
    عرض على GitHub↗11,680
  • crawlscript/webcollectorالصورة الرمزية لـ CrawlScript

    CrawlScript/WebCollector

    3,091عرض على GitHub↗

    WebCollector is an open source web crawler framework based on Java.It provides some simple interfaces for crawling the Web,you can setup a multi-threaded web crawler in less than 5 minutes.

    Java
    عرض على GitHub↗3,091
  • digitalpebble/storm-crawlerالصورة الرمزية لـ DigitalPebble

    DigitalPebble/storm-crawler

    980عرض على GitHub↗

    A scalable, mature and versatile web crawler based on Apache Storm

    Java
    عرض على GitHub↗980
  • internetarchive/heritrix3الصورة الرمزية لـ internetarchive

    internetarchive/heritrix3

    3,246عرض على GitHub↗

    Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project. Heritrix (sometimes spelled heretrix, or misspelled or missaid as heratrix/heritix/heretix/heratix) is an archaic word for heiress (woman who inherits). Since our crawler seeks to collect…

    Java
    عرض على GitHub↗3,246

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Find more with AI search
  • norconex/collector-httpالصورة الرمزية لـ Norconex

    Norconex/collector-http

    202عرض على GitHub↗

    Norconex HTTP Collector

    Java
    عرض على GitHub↗202
  • pkwenda/webbeeالصورة الرمزية لـ pkwenda

    pkwenda/webBee

    193عرض على GitHub↗

    🐝 Web vertical crawler framework for fun

    Java
    عرض على GitHub↗193
  • ssssssss-team/spider-flowالصورة الرمزية لـ ssssssss-team

    ssssssss-team/spider-flow

    11,277عرض على GitHub↗

    Spider-flow is a Java-based web crawling and data extraction platform that provides a centralized environment for managing automated information gathering. It functions as a no-code tool, allowing users to define complex data collection pipelines through a visual, drag-and-drop interface rather than manual programming. The platform distinguishes itself through a graph-based workflow orchestration system where users link discrete nodes to define navigation and parsing logic. It supports dynamic content crawling by integrating headless browsers to execute JavaScript and render page content that

    Javacrawlerjsoupspider
    عرض على GitHub↗11,277
  • vida-nyu/acheالصورة الرمزية لـ ViDA-NYU

    ViDA-NYU/ache

    484عرض على GitHub↗

    ACHE is a focused web crawler. It collects web pages that satisfy some specific criteria, e.g., pages that belong to a given domain or that contain a user-specified pattern. ACHE differs from generic crawlers in sense that it uses page classifiers to distinguish between relevant and irrelevant…

    Java
    عرض على GitHub↗484
  • xtuhcy/geccoالصورة الرمزية لـ xtuhcy

    xtuhcy/gecco

    2,513عرض على GitHub↗

    Gecco is a easy to use lightweight web crawler developed with java language.Gecco integriert jsoup, httpclient, fastjson, spring, htmlunit, redission ausgezeichneten framework,Let you only need to configure a number of jQuery style selector can be very quick to write a crawler.Gecco framework…

    Java
    عرض على GitHub↗2,513
  • yahoo/anthelionالصورة الرمزية لـ yahoo

    yahoo/anthelion

    2,832عرض على GitHub↗

    Anthelion is a Nutch plugin for focused crawling of semantic data. The project is an open-source project released under the Apache License 2.0.

    Java
    عرض على GitHub↗2,832
  • yasserg/crawler4jالصورة الرمزية لـ yasserg

    yasserg/crawler4j

    4,622عرض على GitHub↗

    Crawler4j is a multi-threaded Java web crawler and spider designed for high-volume web traversal and content extraction. It functions as a polite crawling framework that enables the discovery and indexing of HTML and binary content across multiple websites. The project distinguishes itself through a persistent crawling model that serializes session state to local storage, allowing the engine to resume indexing after a crash or interruption. It includes a politeness controller to regulate request frequency and delays, preventing server overloading and IP blocking. The system covers a broad ra

    Java
    عرض على GitHub↗4,622
  • zhegexiaohuozi/seimicrawlerالصورة الرمزية لـ zhegexiaohuozi

    zhegexiaohuozi/SeimiCrawler

    1,993عرض على GitHub↗

    SeimiCrawler

    Java
    عرض على GitHub↗1,993