awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to howie6879/aspider

Open-source alternatives to Aspider

21 open-source projects similar to howie6879/aspider, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Aspider alternative.

  • binux/pyspiderالصورة الرمزية لـ binux

    binux/pyspider

    16,809عرض على GitHub↗

    PySpider is a Python web crawling framework designed for automated data extraction. It provides a pipeline for periodically fetching web content, processing HTML, and persisting scraped information into database backends. The system features a web-based management interface for editing scraping scripts, monitoring task progress, and reviewing collected data. It includes a headless browser JavaScript renderer to capture rendered HTML from dynamic web pages and a distributed architecture that uses message queues to scale crawling workloads across multiple nodes. The framework also covers task

    Python
    عرض على GitHub↗16,809
  • chineking/colaالصورة الرمزية لـ chineking

    chineking/cola

    1,501عرض على GitHub↗

    A high-level distributed crawling framework.

    Python
    عرض على GitHub↗1,501
  • cocrawler/cocrawlerالصورة الرمزية لـ cocrawler

    cocrawler/cocrawler

    194عرض على GitHub↗

    CoCrawler is a versatile web crawler built using modern tools and concurrency.

    Python
    عرض على GitHub↗194
  • codelucas/newspaperالصورة الرمزية لـ codelucas

    codelucas/newspaper

    14,982عرض على GitHub↗

    Newspaper is a Python library designed for scraping, parsing, and analyzing web-based information. It functions as a framework for automated news aggregation and large-scale web content extraction, providing tools to download, clean, and structure text, metadata, and media from diverse online sources. The project distinguishes itself through a pipeline-oriented architecture that combines heuristic-based content extraction with natural language processing. It automatically identifies and isolates article bodies from web page boilerplate while simultaneously performing language detection, keywo

    HTMLcrawlercrawlingnews
    عرض على GitHub↗14,982

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Find more with AI search
  • douban/brownantالصورة الرمزية لـ douban

    douban/brownant

    157عرض على GitHub↗

    |Build Status| |Coverage Status| |PyPI Version| |PyPI Downloads| |Wheel Status|

    Python
    عرض على GitHub↗157
  • gaojiuli/gainالصورة الرمزية لـ gaojiuli

    gaojiuli/gain

    0عرض على GitHub↗

    Taken Over By Shad0w For Responsible Disclosure Kiwi BBP

    عرض على GitHub↗0
  • hickford/mechanicalsoupالصورة الرمزية لـ hickford

    hickford/MechanicalSoup

    4,868عرض على GitHub↗

    MechanicalSoup is a Python web automation library designed to simulate browser behavior. It functions as a toolkit for web scraping and automation, providing an HTML parsing engine and an HTTP session manager to interact with websites programmatically. The library enables headless web interaction by mimicking a real user session. It manages persistent state through cookie handling and automatic redirect following, allowing for programmatic website navigation and the simulation of complex browser interactions. Its capabilities cover automated form population and submission using CSS selectors

    Python
    عرض على GitHub↗4,868
  • holgerd77/django-dynamic-scraperH

    holgerd77/django-dynamic-scraper

    0عرض على GitHub↗

    django-dynamic-scraper

    عرض على GitHub↗0
  • iogf/sukhoiالصورة الرمزية لـ iogf

    iogf/sukhoi

    873عرض على GitHub↗

    Minimalist and powerful Web Crawler.

    Python
    عرض على GitHub↗873
  • istresearch/scrapy-clusterالصورة الرمزية لـ istresearch

    istresearch/scrapy-cluster

    1,224عرض على GitHub↗

    This Scrapy project uses Redis and Kafka to create a distributed on demand scraping cluster.

    Python
    عرض على GitHub↗1,224
  • jmcarp/robobrowserالصورة الرمزية لـ jmcarp

    jmcarp/robobrowser

    3,696عرض على GitHub↗

    Robobrowser is a Python web scraping library that provides a headless browser emulator and an HTML DOM parser. It is designed to programmatically navigate websites, interact with HTML forms, and extract data from web pages. The tool includes a web request caching mechanism to store previously fetched web content, reducing network traffic and increasing loading speeds for repeated requests. It covers capabilities for automated web navigation, programmatic web scraping, and web form automation, including the ability to populate input fields and trigger submission events. The system also manage

    Python
    عرض على GitHub↗3,696
  • jmg/crawleyالصورة الرمزية لـ jmg

    jmg/crawley

    191عرض على GitHub↗

    High Speed WebCrawler built on Eventlet. Supports databases engines like Postgre, Mysql, Oracle, Sqlite. Command line tools. Extract data using your favourite tool. XPath or Pyquery (A Jquery-like library for python). Cookie Handlers. Very easy to use (see the example).

    Python
    عرض على GitHub↗191
  • manning23/mspiderالصورة الرمزية لـ manning23

    manning23/MSpider

    345عرض على GitHub↗

    The information security department of 360 company has been recruiting for a long time and is interested in contacting the mailbox zhangxin1at360.cn.

    Python
    عرض على GitHub↗345
  • matiasb/demiurgeالصورة الرمزية لـ matiasb

    matiasb/demiurge

    118عرض على GitHub↗

    PyQuery-based scraping micro-framework.

    Python
    عرض على GitHub↗118
  • rivermont/spidyالصورة الرمزية لـ rivermont

    rivermont/spidy

    354عرض على GitHub↗

    Spidy (/spˈɪdi/) is the simple, easy to use command line web crawler. Given a list of web links, it uses the Python requests library to query the webpages. Spidy then uses lxml to extract all links from the page and adds them to its list. Pretty simple!

    Python
    عرض على GitHub↗354
  • rolando/scrapy-redisالصورة الرمزية لـ rolando

    rolando/scrapy-redis

    5,639عرض على GitHub↗

    This project is a distributed web crawling framework that enables the horizontal scaling of scraping tasks. It uses Redis as a centralized request queue manager and state store to coordinate crawl progress and request metadata across multiple server instances. The system distributes crawling workloads by sharing a single request queue and utilizes a distributed duplicate filter to prevent multiple workers from visiting the same page. It persists complex request state and metadata as JSON strings within the shared remote store. The framework also provides capabilities for distributed data pro

    Python
    عرض على GitHub↗5,639
  • scrapinghub/portiaالصورة الرمزية لـ scrapinghub

    scrapinghub/portia

    9,509عرض على GitHub↗

    Portia is a containerized scraping platform and visual web scraper that enables no-code data extraction. It serves as a Scrapy visual scraping tool and spider generator, allowing users to design and deploy web scrapers through a graphical interface instead of writing manual selector code. The system distinguishes itself by converting visual web page annotations into executable Scrapy spider code and structured JSON specifications. This visual-to-code mapping allows users to define scraping logic and extraction rules through a point-and-click interface, which can then be exported for use in ex

    Python
    عرض على GitHub↗9,509
  • scrapy/scrapelyالصورة الرمزية لـ scrapy

    scrapy/scrapely

    1,887عرض على GitHub↗

    Scrapely

    HTML
    عرض على GitHub↗1,887
  • scrapy/scrapyالصورة الرمزية لـ scrapy

    scrapy/scrapy

    62,274عرض على GitHub↗

    Scrapy is a comprehensive framework designed for automated web data extraction and large-scale crawling. It operates on an asynchronous, event-driven engine that manages non-blocking network requests and data processing tasks, allowing for the efficient retrieval of structured information from web documents using path-based selectors. The system distinguishes itself through a highly modular architecture that supports complex data collection workflows. Users can implement custom middleware and signal handlers to intercept and modify request flows, while a priority-based scheduler manages concu

    Pythoncrawlercrawlingframework
    عرض على GitHub↗62,274
  • soimort/you-getالصورة الرمزية لـ soimort

    soimort/you-get

    56,839عرض على GitHub↗

    This project is a command-line utility designed to fetch video, audio, and image content from a wide range of web platforms. It functions by parsing page metadata and utilizing modular, site-specific scripts to extract direct media stream URLs from complex web structures, enabling the local archiving of digital media for offline use. The tool distinguishes itself through its ability to handle authenticated content, allowing users to inject browser-stored session cookies to access restricted or private media. It also supports real-time media streaming by piping remote content directly into ext

    Python
    عرض على GitHub↗56,839
  • xianhu/pspiderالصورة الرمزية لـ xianhu

    xianhu/PSpider

    1,840عرض على GitHub↗

    A simple web spider frame written by Python, which needs Python3.8+

    Python
    عرض على GitHub↗1,840