Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.
🦎 A multi-protocol edge & service proxy. Seamlessly interface web apps, IoT clients, & microservices to Apache Kafka® via declaratively defined, stateless APIs.
Apache Pulsar is a cloud-native distributed pub-sub messaging system designed for high-performance data ingestion. It functions as a geo-replicated data streamer and a multi-tenant event streaming platform, providing a serverless stream processing engine and a tiered storage messaging broker. The system distinguishes itself by separating serving layers from storage layers to allow independent scaling of compute and data retention. It features native geo-replication to synchronize messages across different geographical regions and employs a multi-layered tenant isolation model using authentica
ingestr is a command-line tool for copying and syncing data between different database engines and third-party platforms without writing custom code. It functions as an ETL pipeline utility that extracts data from diverse sources and loads it into destinations. The tool features a schema-agnostic data loader that maps source fields to destination columns dynamically, removing the need for predefined static table definitions. It also operates as an incremental data synchronizer, updating destination tables by appending new records or merging changes to maintain current datasets. The system pr
DEPRECATED: Data collection and processing made easy.
الميزات الرئيسية لـ mozilla-services/heka هي: Data Ingestion, Data Ingestion Pipelines, Data Collection Agents.
تشمل البدائل مفتوحة المصدر لـ mozilla-services/heka: facebookarchive/scribe — Scribe is a distributed log aggregation system designed to collect and route real-time log data from numerous servers… aklivity/zilla — 🦎 A multi-protocol edge & service proxy. Seamlessly interface web apps, IoT clients, & microservices to Apache Kafka®… bruin-data/ingestr — ingestr is a command-line tool for copying and syncing data between different database engines and third-party… bruin-data/bruin — Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end… apache/pulsar — Apache Pulsar is a cloud-native distributed pub-sub messaging system designed for high-performance data ingestion. It… gazette/core — Build platforms that flexibly mix SQL, batch, and stream processing paradigms.