awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

17 مستودعات

Awesome GitHub RepositoriesData Ingestion

Tools and platforms for moving, streaming, and synchronizing data.

Explore 17 awesome GitHub repositories matching part of an awesome list · Data Ingestion. Refine with filters or upvote what's useful.

Awesome Data Ingestion GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • apache/pulsarالصورة الرمزية لـ apache

    apache/pulsar

    15,276عرض على GitHub↗

    Apache Pulsar is a cloud-native distributed pub-sub messaging system designed for high-performance data ingestion. It functions as a geo-replicated data streamer and a multi-tenant event streaming platform, providing a serverless stream processing engine and a tiered storage messaging broker. The system distinguishes itself by separating serving layers from storage layers to allow independent scaling of compute and data retention. It features native geo-replication to synchronize messages across different geographical regions and employs a multi-layered tenant isolation model using authentica

    Distributed pub-sub messaging platform with flexible messaging models.

    Java
    عرض على GitHub↗15,276
  • rudderlabs/rudder-serverالصورة الرمزية لـ rudderlabs

    rudderlabs/rudder-server

    4,437عرض على GitHub↗

    Rudder Server عبارة عن منصة بيانات عملاء وخط أنابيب توجيه أحداث مصمم لجمع وتحويل وتوجيه بيانات أحداث العملاء من مصادر مختلفة إلى مستودعات البيانات وأدوات الأعمال. يعمل كمحلل هوية عملاء، يربط المعرفات من مصادر متعددة لبناء رسم بياني موحد للهوية وملفات تعريف سلوكية شاملة للعملاء. يتميز النظام بقدرات ETL العكسية، التي تدفع شرائح العملاء والجماهير المعالجة من مستودعات البيانات مرة أخرى إلى تطبيقات الطرف الثالث التشغيلية. كما يوفر مستوى بيانات حاوية لنشر Kubernetes، مما يتيح إدارة البنية التحتية للبيانات ككود. تغطي المنصة مجموعة واسعة من قدرات إدارة البيانات، بما في ذلك تحويل الأحداث في الوقت الفعلي، والتحقق من المخطط عبر كتالوجات البيانات، وحوكمة الخصوصية. تشمل هذه أدوات لإدارة موافقة المستخدم، وفرض إقامة البيانات داخل مناطق جغرافية محددة، وإخفاء معلومات التعريف الشخصية أثناء النقل. تتم إدارة تثبيت ونشر مكونات مستوى البيانات باستخدام مخططات Helm.

    Open-source customer data infrastructure for event routing.

    Gobigquerycdpcustomer-data
    عرض على GitHub↗4,437
  • facebookarchive/scribeالصورة الرمزية لـ facebookarchive

    facebookarchive/scribe

    3,911عرض على GitHub↗

    Scribe is a distributed log aggregation system designed to collect and route real-time log data from numerous servers to centralized storage or analysis tools. It functions as a log data pipeline and scalable collector that gathers streaming data and writes it to local disks or remote endpoints. The system employs a log routing server model that organizes incoming streams into specific buckets based on predefined configuration mappings. It supports multi-hop log forwarding, allowing data to be routed through a chain of intermediate servers to centralize logs from diverse network segments. Re

    Aggregator for streaming log data.

    C++
    عرض على GitHub↗3,911
  • bruin-data/ingestrالصورة الرمزية لـ bruin-data

    bruin-data/ingestr

    3,714عرض على GitHub↗

    ingestr is a command-line tool for copying and syncing data between different database engines and third-party platforms without writing custom code. It functions as an ETL pipeline utility that extracts data from diverse sources and loads it into destinations. The tool features a schema-agnostic data loader that maps source fields to destination columns dynamically, removing the need for predefined static table definitions. It also operates as an incremental data synchronizer, updating destination tables by appending new records or merging changes to maintain current datasets. The system pr

    CLI utility for copying data between various sources.

    Go
    عرض على GitHub↗3,714
  • mozilla-services/hekaالصورة الرمزية لـ mozilla-services

    mozilla-services/heka

    3,403عرض على GitHub↗

    DEPRECATED: Data collection and processing made easy.

    Open-source stream processing and data collection system.

    Go
    عرض على GitHub↗3,403
  • linkedin/gobblinالصورة الرمزية لـ linkedin

    linkedin/gobblin

    2,267عرض على GitHub↗

    A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

    Universal framework for data ingestion and integration.

    Java
    عرض على GitHub↗2,267
  • pinterest/secorالصورة الرمزية لـ pinterest

    pinterest/secor

    1,858عرض على GitHub↗

    Secor is a service implementing Kafka log persistence

    Service for implementing Kafka log persistence.

    Java
    عرض على GitHub↗1,858
  • bruin-data/bruinالصورة الرمزية لـ bruin-data

    bruin-data/bruin

    1,620عرض على GitHub↗

    Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.

    End-to-end pipeline tool for ingestion, transformation, and quality.

    Goanalyticsbigquerydata-analysis
    عرض على GitHub↗1,620
  • netflix/suroالصورة الرمزية لـ Netflix

    Netflix/suro

    796عرض على GitHub↗

    Netflix's distributed Data Pipeline

    Log aggregation system based on Chukwa.

    Java
    عرض على GitHub↗796
  • gazette/coreالصورة الرمزية لـ gazette

    gazette/core

    793عرض على GitHub↗

    Build platforms that flexibly mix SQL, batch, and stream processing paradigms

    Distributed streaming infrastructure built on cloud storage.

    Go
    عرض على GitHub↗793
  • skizzehq/skizzeالصورة الرمزية لـ skizzehq

    skizzehq/skizze

    772عرض على GitHub↗

    A probabilistic data structure service and storage

    Probabilistic data store for counting and sketching.

    Go
    عرض على GitHub↗772
  • aklivity/zillaالصورة الرمزية لـ aklivity

    aklivity/zilla

    690عرض على GitHub↗

    🦎 A multi-protocol edge & service proxy. Seamlessly interface web apps, IoT clients, & microservices to Apache Kafka® via declaratively defined, stateless APIs.

    API gateway for event-driven architectures and streaming protocols.

    Java
    عرض على GitHub↗690
  • linkedin/white-elephantالصورة الرمزية لـ linkedin

    linkedin/white-elephant

    190عرض على GitHub↗

    Hadoop log aggregator and dashboard

    Log aggregation and dashboarding tool.

    Java
    عرض على GitHub↗190
  • sonalgoyal/hihoالصورة الرمزية لـ sonalgoyal

    sonalgoyal/hiho

    92عرض على GitHub↗

    Hadoop Data Integration with various databases, ftp servers, salesforce. Incremental update, dedup, append, merge your data on Hadoop.

    Framework for connecting disparate data sources to Hadoop.

    Java
    عرض على GitHub↗92
  • linkedin/kamikazeالصورة الرمزية لـ linkedin

    linkedin/kamikaze

    22عرض على GitHub↗

    DocId set compression and set operation library

    Utility for compressing sorted integer arrays.

    Java
    عرض على GitHub↗22
  • streamsets/datacollectorS

    streamsets/datacollector

    0عرض على GitHub↗

    Infrastructure for continuous big data ingestion.

    عرض على GitHub↗0
  • papertrail/kestrelP

    papertrail/kestrel

    0عرض على GitHub↗

    Distributed message queue system for high-throughput data.

    عرض على GitHub↗0
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Data Ingestion