awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to linkedin/white-elephant

Open-source alternatives to White Elephant

30 open-source projects similar to linkedin/white-elephant, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best White Elephant alternative.

  • apache/pulsarAvatar von apache

    apache/pulsar

    15,276Auf GitHub ansehen↗

    Apache Pulsar is a cloud-native distributed pub-sub messaging system designed for high-performance data ingestion. It functions as a geo-replicated data streamer and a multi-tenant event streaming platform, providing a serverless stream processing engine and a tiered storage messaging broker. The system distinguishes itself by separating serving layers from storage layers to allow independent scaling of compute and data retention. It features native geo-replication to synchronize messages across different geographical regions and employs a multi-layered tenant isolation model using authentica

    Java
    Auf GitHub ansehen↗15,276
  • facebookarchive/scribeAvatar von facebookarchive

    facebookarchive/scribe

    3,911Auf GitHub ansehen↗

    Scribe is a distributed log aggregation system designed to collect and route real-time log data from numerous servers to centralized storage or analysis tools. It functions as a log data pipeline and scalable collector that gathers streaming data and writes it to local disks or remote endpoints. The system employs a log routing server model that organizes incoming streams into specific buckets based on predefined configuration mappings. It supports multi-hop log forwarding, allowing data to be routed through a chain of intermediate servers to centralize logs from diverse network segments. Re

    C++
    Auf GitHub ansehen↗3,911
  • mozilla-services/hekaAvatar von mozilla-services

    mozilla-services/heka

    3,403Auf GitHub ansehen↗

    DEPRECATED: Data collection and processing made easy.

    Go
    Auf GitHub ansehen↗3,403
  • netflix/suroAvatar von Netflix

    Netflix/suro

    796Auf GitHub ansehen↗

    Netflix's distributed Data Pipeline

    Java
    Auf GitHub ansehen↗796

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Find more with AI search
  • papertrail/kestrelP

    papertrail/kestrel

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • pinterest/secorAvatar von pinterest

    pinterest/secor

    1,858Auf GitHub ansehen↗

    Secor is a service implementing Kafka log persistence

    Java
    Auf GitHub ansehen↗1,858
  • gazette/coreAvatar von gazette

    gazette/core

    793Auf GitHub ansehen↗

    Build platforms that flexibly mix SQL, batch, and stream processing paradigms

    Go
    Auf GitHub ansehen↗793
  • rudderlabs/rudder-serverAvatar von rudderlabs

    rudderlabs/rudder-server

    4,437Auf GitHub ansehen↗

    Rudder Server is a customer data platform and event routing pipeline designed to collect, transform, and route customer event data from various sources to data warehouses and business tools. It functions as a customer identity resolver, linking identifiers from multiple sources to build a unified identity graph and comprehensive behavioral customer profiles. The system differentiates itself through reverse ETL capabilities, which push processed customer segments and audiences from data warehouses back into operational third-party applications. It also provides a containerized data plane for K

    Gobigquerycdpcustomer-data
    Auf GitHub ansehen↗4,437
  • skizzehq/skizzeAvatar von skizzehq

    skizzehq/skizze

    772Auf GitHub ansehen↗

    A probabilistic data structure service and storage

    Go
    Auf GitHub ansehen↗772
  • aklivity/zillaAvatar von aklivity

    aklivity/zilla

    690Auf GitHub ansehen↗

    🦎 A multi-protocol edge & service proxy. Seamlessly interface web apps, IoT clients, & microservices to Apache Kafka® via declaratively defined, stateless APIs.

    Java
    Auf GitHub ansehen↗690
  • sonalgoyal/hihoAvatar von sonalgoyal

    sonalgoyal/hiho

    92Auf GitHub ansehen↗

    Hadoop Data Integration with various databases, ftp servers, salesforce. Incremental update, dedup, append, merge your data on Hadoop.

    Java
    Auf GitHub ansehen↗92
  • streamsets/datacollectorS

    streamsets/datacollector

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • bruin-data/bruinAvatar von bruin-data

    bruin-data/bruin

    1,620Auf GitHub ansehen↗

    Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.

    Goanalyticsbigquerydata-analysis
    Auf GitHub ansehen↗1,620
  • bruin-data/ingestrAvatar von bruin-data

    bruin-data/ingestr

    3,714Auf GitHub ansehen↗

    ingestr is a command-line tool for copying and syncing data between different database engines and third-party platforms without writing custom code. It functions as an ETL pipeline utility that extracts data from diverse sources and loads it into destinations. The tool features a schema-agnostic data loader that maps source fields to destination columns dynamically, removing the need for predefined static table definitions. It also operates as an incremental data synchronizer, updating destination tables by appending new records or merging changes to maintain current datasets. The system pr

    Go
    Auf GitHub ansehen↗3,714
  • linkedin/gobblinAvatar von linkedin

    linkedin/gobblin

    2,267Auf GitHub ansehen↗

    A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

    Java
    Auf GitHub ansehen↗2,267
  • linkedin/kamikazeAvatar von linkedin

    linkedin/kamikaze

    22Auf GitHub ansehen↗

    DocId set compression and set operation library

    Java
    Auf GitHub ansehen↗22
  • apache/iggyAvatar von apache

    apache/iggy

    4,382Auf GitHub ansehen↗

    Iggy is a distributed message streaming platform and multi-protocol message broker that functions as a persistent distributed log store. It provides infrastructure for publishing and consuming binary messages using an append-only log, ensuring high availability and data consistency across nodes through Viewstamped Replication. The platform is distinguished by its specialized LLM streaming infrastructure, which uses a server protocol to connect large language models to streaming data and system controls. This includes standardized protocols for context management and data bridging via HTTP or

    Rustapachehttpiggy
    Auf GitHub ansehen↗4,382
  • pujansrt/data-genieAvatar von pujansrt

    pujansrt/data-genie

    16Auf GitHub ansehen↗

    High performant ETL engine written in TypeScript

    TypeScript
    Auf GitHub ansehen↗16
  • sohu-co/kafka-nodeAvatar von SOHU-Co

    SOHU-Co/kafka-node

    2,652Auf GitHub ansehen↗

    Node.js client for Apache Kafka 0.8 and later.

    JavaScript
    Auf GitHub ansehen↗2,652
  • twitter/hdfs-duAvatar von twitter

    twitter/hdfs-du

    228Auf GitHub ansehen↗

    HDFS-DU is an interactive visualization of the Hadoop distributed file system. The project aims to monitor different snapshots for the entire HDFS system in an interactive way, showing the size of the folders, the rate at which the size increases / decreases, and to highlight inefficient file…

    JavaScript
    Auf GitHub ansehen↗228
  • uber/kafka-loggerAvatar von uber

    uber/kafka-logger

    45Auf GitHub ansehen↗

    A kafka logger for winston

    JavaScript
    Auf GitHub ansehen↗45
  • wurstmeister/kafka-dockerAvatar von wurstmeister

    wurstmeister/kafka-docker

    6,968Auf GitHub ansehen↗

    This project provides a containerized distribution of Apache Kafka for deploying distributed messaging brokers and event streaming platforms. It functions as a cluster orchestrator that enables the launch of interconnected brokers to establish high-throughput data pipelines. The system uses environment variables to automate topic provisioning and configure broker parameters during the container boot sequence. It manages network listener mapping and advertised hostnames to facilitate client connectivity across different networks. Capability areas include cluster deployment, broker scaling man

    Shell
    Auf GitHub ansehen↗6,968
  • yahoo/kafka-managerAvatar von yahoo

    yahoo/kafka-manager

    11,926Auf GitHub ansehen↗

    Kafka Manager is a web-based management interface and monitoring tool for Apache Kafka clusters. It serves as a central control plane for topic administration, consumer monitoring, and cluster health inspection. The project provides specialized utilities for data rebalancing and partition reassignment to distribute workloads across brokers. It also includes tools to optimize partition leadership by electing preferred replicas. The platform covers a broad range of administrative capabilities, including the creation and configuration of message topics, tracking of consumer offsets, and the col

    Scala
    Auf GitHub ansehen↗11,926
  • airbnb/kafkatAvatar von airbnb

    airbnb/kafkat

    502Auf GitHub ansehen↗

    KafkaT-ool

    Ruby
    Auf GitHub ansehen↗502
  • yelp/mrjobAvatar von Yelp

    Yelp/mrjob

    2,611Auf GitHub ansehen↗

    mrjob: the Python MapReduce library

    Python
    Auf GitHub ansehen↗2,611
  • apache/incubator-gobblinAvatar von apache

    apache/incubator-gobblin

    2,267Auf GitHub ansehen↗

    A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

    Java
    Auf GitHub ansehen↗2,267
  • awslabs/aws-data-wranglerAvatar von awslabs

    awslabs/aws-data-wrangler

    4,107Auf GitHub ansehen↗

    This project is an AWS pandas integration library and data pipeline framework designed to simplify the movement and transformation of data between local memory and AWS storage and analytics services. It functions as a cloud data lake toolkit and storage file manager, allowing users to read, write, and transform structured data across various cloud environments. The library distinguishes itself as a distributed compute orchestrator capable of managing clusters in environments such as EMR to process datasets that exceed the memory limits of a single machine. It also provides specialized capabil

    Python
    Auf GitHub ansehen↗4,107
  • bahador-r/db2lakeAvatar von bahador-r

    bahador-r/db2lake

    2Auf GitHub ansehen↗
    TypeScript
    Auf GitHub ansehen↗2
  • bwhite/hadoopyAvatar von bwhite

    bwhite/hadoopy

    243Auf GitHub ansehen↗

    Brandyn White Andrew Miller

    C
    Auf GitHub ansehen↗243
  • confluentinc/bottledwater-pgC

    confluentinc/bottledwater-pg

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0