awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

17 repository-uri

Awesome GitHub RepositoriesData Ingestion

Tools and platforms for moving, streaming, and synchronizing data.

Explore 17 awesome GitHub repositories matching part of an awesome list · Data Ingestion. Refine with filters or upvote what's useful.

Awesome Data Ingestion GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • apache/pulsarAvatar apache

    apache/pulsar

    15,276Vezi pe GitHub↗

    Apache Pulsar is a cloud-native distributed pub-sub messaging system designed for high-performance data ingestion. It functions as a geo-replicated data streamer and a multi-tenant event streaming platform, providing a serverless stream processing engine and a tiered storage messaging broker. The system distinguishes itself by separating serving layers from storage layers to allow independent scaling of compute and data retention. It features native geo-replication to synchronize messages across different geographical regions and employs a multi-layered tenant isolation model using authentica

    Distributed pub-sub messaging platform with flexible messaging models.

    Java
    Vezi pe GitHub↗15,276
  • rudderlabs/rudder-serverAvatar rudderlabs

    rudderlabs/rudder-server

    4,437Vezi pe GitHub↗

    Rudder Server is a customer data platform and event routing pipeline designed to collect, transform, and route customer event data from various sources to data warehouses and business tools. It functions as a customer identity resolver, linking identifiers from multiple sources to build a unified identity graph and comprehensive behavioral customer profiles. The system differentiates itself through reverse ETL capabilities, which push processed customer segments and audiences from data warehouses back into operational third-party applications. It also provides a containerized data plane for K

    Open-source customer data infrastructure for event routing.

    Gobigquerycdpcustomer-data
    Vezi pe GitHub↗4,437
  • facebookarchive/scribeAvatar facebookarchive

    facebookarchive/scribe

    3,911Vezi pe GitHub↗

    Scribe este un sistem distribuit de agregare a log-urilor, conceput pentru a colecta și ruta date de log în timp real de la numeroase servere către stocare centralizată sau instrumente de analiză. Funcționează ca un pipeline de date de log și un colector scalabil care adună datele în flux și le scrie pe discuri locale sau endpoint-uri la distanță. Sistemul folosește un model de server de rutare a log-urilor care organizează fluxurile primite în „buckets” specifice, bazate pe mapări de configurare predefinite. Suportă redirecționarea log-urilor prin mai multe hop-uri, permițând rutarea datelor printr-un lanț de servere intermediare pentru a centraliza log-urile din diverse segmente de rețea. Fiabilitatea și observabilitatea sunt gestionate prin buffering pe disc local, care stochează mesajele de ieșire în timpul defecțiunilor de rețea pentru a preveni pierderea datelor, și monitorizarea sănătății bazată pe contoare pentru a urmări numărul intern de mesaje și succesul livrării. Arhitectura include, de asemenea, streaming asincron de mesaje și rutare centralizată a log-urilor pentru a gestiona fluxurile de date cu throughput ridicat în rețea.

    Aggregator for streaming log data.

    C++
    Vezi pe GitHub↗3,911
  • bruin-data/ingestrAvatar bruin-data

    bruin-data/ingestr

    3,714Vezi pe GitHub↗

    ingestr is a command-line tool for copying and syncing data between different database engines and third-party platforms without writing custom code. It functions as an ETL pipeline utility that extracts data from diverse sources and loads it into destinations. The tool features a schema-agnostic data loader that maps source fields to destination columns dynamically, removing the need for predefined static table definitions. It also operates as an incremental data synchronizer, updating destination tables by appending new records or merging changes to maintain current datasets. The system pr

    CLI utility for copying data between various sources.

    Go
    Vezi pe GitHub↗3,714
  • mozilla-services/hekaAvatar mozilla-services

    mozilla-services/heka

    3,403Vezi pe GitHub↗

    DEPRECATED: Data collection and processing made easy.

    Open-source stream processing and data collection system.

    Go
    Vezi pe GitHub↗3,403
  • linkedin/gobblinAvatar linkedin

    linkedin/gobblin

    2,267Vezi pe GitHub↗

    A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

    Universal framework for data ingestion and integration.

    Java
    Vezi pe GitHub↗2,267
  • pinterest/secorAvatar pinterest

    pinterest/secor

    1,858Vezi pe GitHub↗

    Secor is a service implementing Kafka log persistence

    Service for implementing Kafka log persistence.

    Java
    Vezi pe GitHub↗1,858
  • bruin-data/bruinAvatar bruin-data

    bruin-data/bruin

    1,620Vezi pe GitHub↗

    Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.

    End-to-end pipeline tool for ingestion, transformation, and quality.

    Goanalyticsbigquerydata-analysis
    Vezi pe GitHub↗1,620
  • netflix/suroAvatar Netflix

    Netflix/suro

    796Vezi pe GitHub↗

    Netflix's distributed Data Pipeline

    Log aggregation system based on Chukwa.

    Java
    Vezi pe GitHub↗796
  • gazette/coreAvatar gazette

    gazette/core

    793Vezi pe GitHub↗

    Build platforms that flexibly mix SQL, batch, and stream processing paradigms

    Distributed streaming infrastructure built on cloud storage.

    Go
    Vezi pe GitHub↗793
  • skizzehq/skizzeAvatar skizzehq

    skizzehq/skizze

    772Vezi pe GitHub↗

    A probabilistic data structure service and storage

    Probabilistic data store for counting and sketching.

    Go
    Vezi pe GitHub↗772
  • aklivity/zillaAvatar aklivity

    aklivity/zilla

    690Vezi pe GitHub↗

    🦎 A multi-protocol edge & service proxy. Seamlessly interface web apps, IoT clients, & microservices to Apache Kafka® via declaratively defined, stateless APIs.

    API gateway for event-driven architectures and streaming protocols.

    Java
    Vezi pe GitHub↗690
  • linkedin/white-elephantAvatar linkedin

    linkedin/white-elephant

    190Vezi pe GitHub↗

    Hadoop log aggregator and dashboard

    Log aggregation and dashboarding tool.

    Java
    Vezi pe GitHub↗190
  • sonalgoyal/hihoAvatar sonalgoyal

    sonalgoyal/hiho

    92Vezi pe GitHub↗

    Hadoop Data Integration with various databases, ftp servers, salesforce. Incremental update, dedup, append, merge your data on Hadoop.

    Framework for connecting disparate data sources to Hadoop.

    Java
    Vezi pe GitHub↗92
  • linkedin/kamikazeAvatar linkedin

    linkedin/kamikaze

    22Vezi pe GitHub↗

    DocId set compression and set operation library

    Utility for compressing sorted integer arrays.

    Java
    Vezi pe GitHub↗22
  • streamsets/datacollectorS

    streamsets/datacollector

    0Vezi pe GitHub↗

    Infrastructure for continuous big data ingestion.

    Vezi pe GitHub↗0
  • papertrail/kestrelP

    papertrail/kestrel

    0Vezi pe GitHub↗

    Distributed message queue system for high-throughput data.

    Vezi pe GitHub↗0
  1. Home
  2. Part of an Awesome List
  3. Databases & Data
  4. Data Ingestion