awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to linkedin/gobblin

Open-source alternatives to Gobblin

30 open-source projects similar to linkedin/gobblin, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Gobblin alternative.

  • netflix/suroNetflix 的头像

    Netflix/suro

    796在 GitHub 上查看↗

    Netflix's distributed Data Pipeline

    Java
    在 GitHub 上查看↗796
  • facebookarchive/scribefacebookarchive 的头像

    facebookarchive/scribe

    3,911在 GitHub 上查看↗

    Scribe is a distributed log aggregation system designed to collect and route real-time log data from numerous servers to centralized storage or analysis tools. It functions as a log data pipeline and scalable collector that gathers streaming data and writes it to local disks or remote endpoints. The system employs a log routing server model that organizes incoming streams into specific buckets based on predefined configuration mappings. It supports multi-hop log forwarding, allowing data to be routed through a chain of intermediate servers to centralize logs from diverse network segments. Re

    C++
    在 GitHub 上查看↗3,911
  • linkedin/white-elephantlinkedin 的头像

    linkedin/white-elephant

    190在 GitHub 上查看↗

    Hadoop log aggregator and dashboard

    Java
    在 GitHub 上查看↗190
  • mozilla-services/hekamozilla-services 的头像

    mozilla-services/heka

    3,403在 GitHub 上查看↗

    DEPRECATED: Data collection and processing made easy.

    Go
    在 GitHub 上查看↗3,403
  • gazette/coregazette 的头像

    gazette/core

    793在 GitHub 上查看↗

    Build platforms that flexibly mix SQL, batch, and stream processing paradigms

    Go
    在 GitHub 上查看↗793

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Find more with AI search
  • papertrail/kestrelP

    papertrail/kestrel

    0在 GitHub 上查看↗
    在 GitHub 上查看↗0
  • pinterest/secorpinterest 的头像

    pinterest/secor

    1,858在 GitHub 上查看↗

    Secor is a service implementing Kafka log persistence

    Java
    在 GitHub 上查看↗1,858
  • aklivity/zillaaklivity 的头像

    aklivity/zilla

    690在 GitHub 上查看↗

    🦎 A multi-protocol edge & service proxy. Seamlessly interface web apps, IoT clients, & microservices to Apache Kafka® via declaratively defined, stateless APIs.

    Java
    在 GitHub 上查看↗690
  • rudderlabs/rudder-serverrudderlabs 的头像

    rudderlabs/rudder-server

    4,437在 GitHub 上查看↗

    Rudder Server is a customer data platform and event routing pipeline designed to collect, transform, and route customer event data from various sources to data warehouses and business tools. It functions as a customer identity resolver, linking identifiers from multiple sources to build a unified identity graph and comprehensive behavioral customer profiles. The system differentiates itself through reverse ETL capabilities, which push processed customer segments and audiences from data warehouses back into operational third-party applications. It also provides a containerized data plane for K

    Gobigquerycdpcustomer-data
    在 GitHub 上查看↗4,437
  • skizzehq/skizzeskizzehq 的头像

    skizzehq/skizze

    772在 GitHub 上查看↗

    A probabilistic data structure service and storage

    Go
    在 GitHub 上查看↗772
  • apache/pulsarapache 的头像

    apache/pulsar

    15,276在 GitHub 上查看↗

    Apache Pulsar is a cloud-native distributed pub-sub messaging system designed for high-performance data ingestion. It functions as a geo-replicated data streamer and a multi-tenant event streaming platform, providing a serverless stream processing engine and a tiered storage messaging broker. The system distinguishes itself by separating serving layers from storage layers to allow independent scaling of compute and data retention. It features native geo-replication to synchronize messages across different geographical regions and employs a multi-layered tenant isolation model using authentica

    Java
    在 GitHub 上查看↗15,276
  • sonalgoyal/hihosonalgoyal 的头像

    sonalgoyal/hiho

    92在 GitHub 上查看↗

    Hadoop Data Integration with various databases, ftp servers, salesforce. Incremental update, dedup, append, merge your data on Hadoop.

    Java
    在 GitHub 上查看↗92
  • streamsets/datacollectorS

    streamsets/datacollector

    0在 GitHub 上查看↗
    在 GitHub 上查看↗0
  • bruin-data/bruinbruin-data 的头像

    bruin-data/bruin

    1,620在 GitHub 上查看↗

    Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.

    Goanalyticsbigquerydata-analysis
    在 GitHub 上查看↗1,620
  • bruin-data/ingestrbruin-data 的头像

    bruin-data/ingestr

    3,714在 GitHub 上查看↗

    ingestr is a command-line tool for copying and syncing data between different database engines and third-party platforms without writing custom code. It functions as an ETL pipeline utility that extracts data from diverse sources and loads it into destinations. The tool features a schema-agnostic data loader that maps source fields to destination columns dynamically, removing the need for predefined static table definitions. It also operates as an incremental data synchronizer, updating destination tables by appending new records or merging changes to maintain current datasets. The system pr

    Go
    在 GitHub 上查看↗3,714
  • linkedin/kamikazelinkedin 的头像

    linkedin/kamikaze

    22在 GitHub 上查看↗

    DocId set compression and set operation library

    Java
    在 GitHub 上查看↗22
  • benthosdev/benthosbenthosdev 的头像

    benthosdev/benthos

    8,681在 GitHub 上查看↗

    Benthos is a stream processing engine and data integration pipeline used for routing, transforming, and connecting data streams between diverse sources and sinks. It functions as event routing middleware and a change data capture tool, streaming real-time database modifications as discrete events for downstream processing. The system utilizes a declarative pipeline configuration, where data flow and processing logic are defined in a single static file. It features a specialized domain-specific language for mapping, filtering, and enriching data payloads, allowing for complex transformations w

    Go
    在 GitHub 上查看↗8,681
  • apache/iggyapache 的头像

    apache/iggy

    4,382在 GitHub 上查看↗

    Iggy is a distributed message streaming platform and multi-protocol message broker that functions as a persistent distributed log store. It provides infrastructure for publishing and consuming binary messages using an append-only log, ensuring high availability and data consistency across nodes through Viewstamped Replication. The platform is distinguished by its specialized LLM streaming infrastructure, which uses a server protocol to connect large language models to streaming data and system controls. This includes standardized protocols for context management and data bridging via HTTP or

    Rustapachehttpiggy
    在 GitHub 上查看↗4,382
  • kroxylicious/kroxyliciouskroxylicious 的头像

    kroxylicious/kroxylicious

    288在 GitHub 上查看↗

    Kroxylicious, the snappy open source proxy for Apache Kafka®

    Java
    在 GitHub 上查看↗288
  • linkedin/datafuL

    linkedin/datafu

    0在 GitHub 上查看↗
    在 GitHub 上查看↗0
  • airbnb/kafkatairbnb 的头像

    airbnb/kafkat

    502在 GitHub 上查看↗

    KafkaT-ool

    Ruby
    在 GitHub 上查看↗502
  • nameetp/pdfmuxNameetP 的头像

    NameetP/pdfmux

    69在 GitHub 上查看↗

    PDF extraction that checks its own work. #2 reading order accuracy — zero AI, zero GPU, zero cost.

    Python
    在 GitHub 上查看↗69
  • oryxproject/oryxoryxproject 的头像

    oryxproject/oryx

    1,783在 GitHub 上查看↗

    Oryx 2: Lambda architecture on Apache Spark, Apache Kafka for real-time large scale machine learning

    Java
    在 GitHub 上查看↗1,783
  • pujansrt/data-geniepujansrt 的头像

    pujansrt/data-genie

    16在 GitHub 上查看↗

    High performant ETL engine written in TypeScript

    TypeScript
    在 GitHub 上查看↗16
  • sohu-co/kafka-nodeSOHU-Co 的头像

    SOHU-Co/kafka-node

    2,652在 GitHub 上查看↗

    Node.js client for Apache Kafka 0.8 and later.

    JavaScript
    在 GitHub 上查看↗2,652
  • twitter/elephant-birdtwitter 的头像

    twitter/elephant-bird

    1,133在 GitHub 上查看↗

    Twitter's collection of LZO and Protocol Buffer-related Hadoop, Pig, Hive, and HBase code.

    Java
    在 GitHub 上查看↗1,133
  • twitter/herontwitter 的头像

    twitter/heron

    3,632在 GitHub 上查看↗

    Apache Heron (Incubating) is a realtime, distributed, fault-tolerant stream processing engine from Twitter

    Java
    在 GitHub 上查看↗3,632
  • uber/kafka-loggeruber 的头像

    uber/kafka-logger

    45在 GitHub 上查看↗

    A kafka logger for winston

    JavaScript
    在 GitHub 上查看↗45
  • wurstmeister/kafka-dockerwurstmeister 的头像

    wurstmeister/kafka-docker

    6,968在 GitHub 上查看↗

    This project provides a containerized distribution of Apache Kafka for deploying distributed messaging brokers and event streaming platforms. It functions as a cluster orchestrator that enables the launch of interconnected brokers to establish high-throughput data pipelines. The system uses environment variables to automate topic provisioning and configure broker parameters during the container boot sequence. It manages network listener mapping and advertised hostnames to facilitate client connectivity across different networks. Capability areas include cluster deployment, broker scaling man

    Shell
    在 GitHub 上查看↗6,968
  • yahoo/kafka-manageryahoo 的头像

    yahoo/kafka-manager

    11,926在 GitHub 上查看↗

    Kafka Manager is a web-based management interface and monitoring tool for Apache Kafka clusters. It serves as a central control plane for topic administration, consumer monitoring, and cluster health inspection. The project provides specialized utilities for data rebalancing and partition reassignment to distribute workloads across brokers. It also includes tools to optimize partition leadership by electing preferred replicas. The platform covers a broad range of administrative capabilities, including the creation and configuration of message topics, tracking of consumer offsets, and the col

    Scala
    在 GitHub 上查看↗11,926