awesome-repositories.com
المدونة
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
jaypyles avatar

jaypyles/ScraperrArchived

0
View on GitHub↗
4,897 نجوم·242 تفرعات·TypeScript·MIT·8 مشاهداتscraperr-docs.pages.dev↗

Scraperr

Scraperr is a self-hosted web scraping and crawling platform designed for extracting structured data from websites using XPath selectors. It functions as a containerized system for managing scraping jobs through a queue and analyzing the resulting content using artificial intelligence.

The project differentiates itself through its Kubernetes-native architecture, allowing for scalable deployment and management via package managers. It includes a crawling engine capable of domain-level spidering to discover linked pages and a data analyzer that uses artificial intelligence to query extracted web content.

The platform covers a broad range of capabilities, including automated data extraction, bulk web crawling, and media file downloading. It provides tools for visualizing scraped data in tables, configuring custom request headers to mimic browser identities, and exporting results into CSV or Markdown formats.

The application supports customizable installation parameters and version updates through Kubernetes deployment configurations.

Features

  • XPath Data Extractors - Uses XPath selectors to precisely target and extract specific data elements from web pages.
  • Web Scraping and Extraction - Provides a self-hosted platform for extracting structured data from websites using XPath selectors.
  • Automated Web Scraping - Provides a self-hosted platform for automatically extracting structured information from websites.
  • CSS and XPath Query Engines - Utilizes XPath expressions to locate and extract specific nodes from parsed HTML documents.
  • Domain-Restricted Crawling - Includes a crawling engine that limits link discovery to a specific root domain for focused data collection.
  • Kubernetes Application Deployments - Utilizes a Kubernetes-native architecture for scalable deployment and automated application management.
  • Job Queues - Manages bulk scraping tasks through a sequential job queue to ensure reliable data collection.
  • Data Extractions - Uses XPath selectors and DOM traversal to pull structured data from multiple URLs.
  • Domain Spidering - Implements a crawler that visits all linked pages within a domain to discover wide-ranging content.
  • Scraping Infrastructure Management - Provides infrastructure to submit and track multiple scraping tasks via a centralized job queue.
  • Web Crawling - Systematically discovers and indexes web content across entire domains for large-scale data collection.
  • Web Scrapers - Ships an automated system for navigating websites and extracting structured data via a job queue.
  • Selector Mappings - Enables precise data extraction by mapping target URLs to specific XPath selectors.
  • Content Query Analyzers - Integrates AI to allow users to query and analyze the extracted web content using natural language.
  • Web Content AI Analysis - Implements AI-driven analysis to query and retrieve information from extracted web content.
  • AI-Powered Intelligence - Leverages artificial intelligence to query and extract insights from large volumes of scraped web content.
  • CSV Data Exports - Exports scraped data into CSV and Markdown files for external analysis in spreadsheet applications.
  • AI Data Analysis Tools - Provides a capability to answer questions based on scraped web data using AI model APIs.
  • Kubernetes-Native Data Platforms - Implements a containerized scraping platform designed specifically for scalable orchestration on Kubernetes.
  • File-Based Data Exports - Serializes collected web data into structured formats such as CSV and Markdown for external use.
  • Scraped Data Exporters - Converts results from completed scraping jobs into structured CSV files.
  • Kubernetes Application Deployments - Provides a Kubernetes-native architecture for scalable deployment and management of the scraping platform.

سجل النجوم

مخطط تاريخ النجوم لـ jaypyles/scraperrمخطط تاريخ النجوم لـ jaypyles/scraperr

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

بدائل مفتوحة المصدر لـ Scraperr

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Scraperr.
  • any4ai/anycrawlالصورة الرمزية لـ any4ai

    any4ai/AnyCrawl

    2,742عرض على GitHub↗

    AnyCrawl is an AI-powered data extractor, automated web crawler, and headless browser orchestrator. It serves as a web content extraction API and a gateway that connects crawling and scraping tools to language models using a standardized API protocol. The project specializes in converting unstructured website content into structured JSON or markdown optimized for AI assistants. It utilizes language models and JSON schemas to pull specific information into validated formats and provides capabilities for AI page summarization and LLM-optimized content extraction. The system manages comprehensi

    TypeScriptai-scrapingaitoolscrawl
    عرض على GitHub↗2,742
  • rchipka/node-osmosisR

    rchipka/node-osmosis

    4,110عرض على GitHub↗

    This project is a Node.js web scraping framework designed to automate data extraction through a programmatic workflow of requests, parsing, and document interaction. It functions as a headless web crawler, an HTTP request manager, and a DOM parser and extractor. The framework distinguishes itself by combining a JavaScript execution engine to interact with dynamic content and a hybrid selection system that utilizes both CSS and XPath selectors. It includes specialized middleware for proxy rotation and cookie-jar session management to maintain authenticated states and manage automated traffic.

    JavaScript
    عرض على GitHub↗4,110
  • friendsofphp/goutteF

    FriendsOfPHP/Goutte

    9,201عرض على GitHub↗

    Goutte is a PHP web scraper and DOM crawler designed for extracting data from websites. It functions as an HTTP client wrapper that enables the retrieval of web pages and the parsing of HTML content. The project provides a web form automator to programmatically fill and submit HTML forms to remote servers. It also includes a mechanism for automated website crawling by following links to discover and archive web content. The system supports stateful session management to maintain cookies and headers across requests. It further covers HTML data extraction through DOM-based element selection an

    PHP
    عرض على GitHub↗9,201
  • asciimoo/collyالصورة الرمزية لـ asciimoo

    asciimoo/colly

    25,348عرض على GitHub↗

    Colly is a web scraping framework and concurrent crawler written in Go. It provides a system for traversing web pages, following links, and extracting structured data from HTML and XML documents. The framework includes a distributed scraping engine designed to spread data collection tasks across multiple instances to increase throughput. It ensures compliance with website owner policies by automatically reading and respecting robots.txt files. The system manages request lifecycles through domain-based rate limiting, concurrency controls, and session management via a stateful cookie jar. It s

    Go
    عرض على GitHub↗25,348
عرض جميع البدائل الـ 30 لـ Scraperr→

الأسئلة الشائعة

ما هي وظيفة jaypyles/scraperr؟

Scraperr is a self-hosted web scraping and crawling platform designed for extracting structured data from websites using XPath selectors. It functions as a containerized system for managing scraping jobs through a queue and analyzing the resulting content using artificial intelligence.

ما هي الميزات الرئيسية لـ jaypyles/scraperr؟

الميزات الرئيسية لـ jaypyles/scraperr هي: XPath Data Extractors, Web Scraping and Extraction, Automated Web Scraping, CSS and XPath Query Engines, Domain-Restricted Crawling, Kubernetes Application Deployments, Job Queues, Data Extractions.

ما هي البدائل مفتوحة المصدر لـ jaypyles/scraperr؟

تشمل البدائل مفتوحة المصدر لـ jaypyles/scraperr: any4ai/anycrawl — AnyCrawl is an AI-powered data extractor, automated web crawler, and headless browser orchestrator. It serves as a web… rchipka/node-osmosis — This project is a Node.js web scraping framework designed to automate data extraction through a programmatic workflow… friendsofphp/goutte — Goutte is a PHP web scraper and DOM crawler designed for extracting data from websites. It functions as an HTTP client… generalnewsextractor/generalnewsextractor — GeneralNewsExtractor is a specialized system for identifying and extracting structured news data through configurable… asciimoo/colly — Colly is a web scraping framework and concurrent crawler written in Go. It provides a system for traversing web pages,… ionicabizau/scrape-it — scrape-it is a Node.js web scraper and HTML parser designed to extract structured data from websites and HTML files.…