awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to mozilla-services/heka

Open-source alternatives to Heka

30 open-source projects similar to mozilla-services/heka, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Heka alternative.

  • linkedin/white-elephantlinkedin avatar

    linkedin/white-elephant

    190View on GitHub↗

    Hadoop log aggregator and dashboard

    Java
    View on GitHub↗190
  • netflix/suroNetflix avatar

    Netflix/suro

    796View on GitHub↗

    Netflix's distributed Data Pipeline

    Java
    View on GitHub↗796
  • apache/pulsarapache avatar

    apache/pulsar

    15,276View on GitHub↗

    Apache Pulsar is a cloud-native distributed pub-sub messaging system designed for high-performance data ingestion. It functions as a geo-replicated data streamer and a multi-tenant event streaming platform, providing a serverless stream processing engine and a tiered storage messaging broker. The system distinguishes itself by separating serving layers from storage layers to allow independent scaling of compute and data retention. It features native geo-replication to synchronize messages across different geographical regions and employs a multi-layered tenant isolation model using authentica

    Java
    View on GitHub↗15,276
  • aklivity/zillaaklivity avatar

    aklivity/zilla

    690View on GitHub↗

    🦎 A multi-protocol edge & service proxy. Seamlessly interface web apps, IoT clients, & microservices to Apache Kafka® via declaratively defined, stateless APIs.

    Java
    View on GitHub↗690
  • streamsets/datacollectorS

    streamsets/datacollector

    0View on GitHub↗
    View on GitHub↗0

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • rudderlabs/rudder-serverrudderlabs avatar

    rudderlabs/rudder-server

    4,437View on GitHub↗

    Rudder Server is a customer data platform and event routing pipeline designed to collect, transform, and route customer event data from various sources to data warehouses and business tools. It functions as a customer identity resolver, linking identifiers from multiple sources to build a unified identity graph and comprehensive behavioral customer profiles. The system differentiates itself through reverse ETL capabilities, which push processed customer segments and audiences from data warehouses back into operational third-party applications. It also provides a containerized data plane for K

    Gobigquerycdpcustomer-data
    View on GitHub↗4,437
  • pinterest/secorpinterest avatar

    pinterest/secor

    1,858View on GitHub↗

    Secor is a service implementing Kafka log persistence

    Java
    View on GitHub↗1,858
  • papertrail/kestrelP

    papertrail/kestrel

    0View on GitHub↗
    View on GitHub↗0
  • sonalgoyal/hihosonalgoyal avatar

    sonalgoyal/hiho

    92View on GitHub↗

    Hadoop Data Integration with various databases, ftp servers, salesforce. Incremental update, dedup, append, merge your data on Hadoop.

    Java
    View on GitHub↗92
  • gazette/coregazette avatar

    gazette/core

    793View on GitHub↗

    Build platforms that flexibly mix SQL, batch, and stream processing paradigms

    Go
    View on GitHub↗793
  • bruin-data/ingestrbruin-data avatar

    bruin-data/ingestr

    3,714View on GitHub↗

    ingestr is a command-line tool for copying and syncing data between different database engines and third-party platforms without writing custom code. It functions as an ETL pipeline utility that extracts data from diverse sources and loads it into destinations. The tool features a schema-agnostic data loader that maps source fields to destination columns dynamically, removing the need for predefined static table definitions. It also operates as an incremental data synchronizer, updating destination tables by appending new records or merging changes to maintain current datasets. The system pr

    Go
    View on GitHub↗3,714
  • facebookarchive/scribefacebookarchive avatar

    facebookarchive/scribe

    3,911View on GitHub↗

    Scribe is a distributed log aggregation system designed to collect and route real-time log data from numerous servers to centralized storage or analysis tools. It functions as a log data pipeline and scalable collector that gathers streaming data and writes it to local disks or remote endpoints. The system employs a log routing server model that organizes incoming streams into specific buckets based on predefined configuration mappings. It supports multi-hop log forwarding, allowing data to be routed through a chain of intermediate servers to centralize logs from diverse network segments. Re

    C++
    View on GitHub↗3,911
  • bruin-data/bruinbruin-data avatar

    bruin-data/bruin

    1,620View on GitHub↗

    Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.

    Goanalyticsbigquerydata-analysis
    View on GitHub↗1,620
  • skizzehq/skizzeskizzehq avatar

    skizzehq/skizze

    772View on GitHub↗

    A probabilistic data structure service and storage

    Go
    View on GitHub↗772
  • linkedin/kamikazelinkedin avatar

    linkedin/kamikaze

    22View on GitHub↗

    DocId set compression and set operation library

    Java
    View on GitHub↗22
  • linkedin/gobblinlinkedin avatar

    linkedin/gobblin

    2,267View on GitHub↗

    A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

    Java
    View on GitHub↗2,267
  • apache/iggyapache avatar

    apache/iggy

    4,382View on GitHub↗

    Iggy is a distributed message streaming platform and multi-protocol message broker that functions as a persistent distributed log store. It provides infrastructure for publishing and consuming binary messages using an append-only log, ensuring high availability and data consistency across nodes through Viewstamped Replication. The platform is distinguished by its specialized LLM streaming infrastructure, which uses a server protocol to connect large language models to streaming data and system controls. This includes standardized protocols for context management and data bridging via HTTP or

    Rustapachehttpiggy
    View on GitHub↗4,382
  • chrissnell/gopherwxchrissnell avatar

    chrissnell/gopherwx

    32View on GitHub↗

    RemoteWeather is a professional weather monitoring system that collects data from your weather station, stores it for historical analysis, and shares it through multiple channels including a beautiful live website, weather services, and amateur radio networks.

    Go
    View on GitHub↗32
  • gatling/gatlinggatling avatar

    gatling/gatling

    6,923View on GitHub↗

    Gatling is a load testing framework and traffic generation engine used to measure response times and error rates under heavy load. It functions as an as-code testing library, allowing users to define high-volume traffic simulations and performance tests through programming languages rather than graphical interfaces. The system enables multi-language load simulation and the ability to model concurrent user traffic to identify infrastructure bottlenecks and stability limits. It supports a test-as-code workflow, where version-controlled scripts are integrated into build pipelines as performance

    Scala
    View on GitHub↗6,923
  • centreon/centreoncentreon avatar

    centreon/centreon

    159View on GitHub↗

    Open source part of the centreon monorepoc

    PHP
    View on GitHub↗159
  • fullcontact/crankshaftdfullcontact avatar

    fullcontact/crankshaftd

    6View on GitHub↗

    crankshaftd

    Go
    View on GitHub↗6
  • faradayrf/aprs2influxdbFaradayRF avatar

    FaradayRF/aprs2influxdb

    29View on GitHub↗

    This program interfaces ham radio APRS-IS servers and saves packet data into an influxdb database. aprs2influxdb handles the connection, parsing, and saving of data into an influxdb database from APRS-IS using line protocol formatted strings. Periodically, a status message is also sent to the…

    Python
    View on GitHub↗29
  • ccpgames/aggregatedccpgames avatar

    ccpgames/aggregateD

    14View on GitHub↗

    aggregateD

    Go
    View on GitHub↗14
  • att-innovate/charmanderatt-innovate avatar

    att-innovate/charmander

    67View on GitHub↗

    Charmander Scheduler Lab

    Shell
    View on GitHub↗67
  • etsy/statsd-jvm-profileretsy avatar

    etsy/statsd-jvm-profiler

    335View on GitHub↗

    statsd-jvm-profiler is a JVM agent profiler that sends profiling data to StatsD. Inspired by riemann-jvm-profiler, it was primarily built for profiling Hadoop jobs, but can be used with any JVM process.

    Java
    View on GitHub↗335
  • fulldecent/google-sheets-etlfulldecent avatar

    fulldecent/google-sheets-etl

    22View on GitHub↗

    Live import all your Google Sheets to your data warehouse

    PHP
    View on GitHub↗22
  • edenhill/librdkafkaedenhill avatar

    edenhill/librdkafka

    991View on GitHub↗

    The Apache Kafka C/C++ library

    C
    View on GitHub↗991
  • edenhill/kafkacatedenhill avatar

    edenhill/kafkacat

    5,761View on GitHub↗

    Kafkacat is a suite of command-line utilities for interacting with Apache Kafka clusters. It provides a non-JVM binary for producing and consuming messages, inspecting cluster metadata, and debugging the Kafka protocol via the terminal. The tool functions as a producer and consumer capable of pushing data from files or standard input and reading messages from specific topics and partitions. It includes a metadata inspector to retrieve cluster state and partition configurations in plain text or JSON, as well as a protocol debugger for inspecting message offsets, timestamps, and binary payloads

    C
    View on GitHub↗5,761
  • google/cadvisorgoogle avatar

    google/cadvisor

    19,202View on GitHub↗

    cAdvisor is a container resource monitoring agent and performance analyzer that collects and exports CPU, memory, network, and disk usage statistics from running containers. It functions as a telemetry tool for discovering containers across various runtimes and serves as a Prometheus-compatible metrics exporter. The agent distinguishes itself by analyzing Linux control groups to provide visibility into resource consumption and limits. It utilizes kernel perf events and NUMA statistics for low-level hardware performance tracking and diagnostics, and it can identify out-of-memory kill events th

    Go
    View on GitHub↗19,202
  • dataspoclab/dataspoc-pipedataspoclab avatar

    dataspoclab/dataspoc-pipe

    2View on GitHub↗

    Data ingestion engine — Singer taps to Parquet in cloud buckets

    Python
    View on GitHub↗2