awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to streamsets/datacollector

Open-source alternatives to Datacollector

30 open-source projects similar to streamsets/datacollector, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Datacollector alternative.

  • linkedin/gobblinالصورة الرمزية لـ linkedin

    linkedin/gobblin

    2,267عرض على GitHub↗

    A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

    Java
    عرض على GitHub↗2,267
  • linkedin/kamikazeالصورة الرمزية لـ linkedin

    linkedin/kamikaze

    22عرض على GitHub↗

    DocId set compression and set operation library

    Java
    عرض على GitHub↗22
  • linkedin/white-elephantالصورة الرمزية لـ linkedin

    linkedin/white-elephant

    190عرض على GitHub↗

    Hadoop log aggregator and dashboard

    Java
    عرض على GitHub↗190
  • aklivity/zillaالصورة الرمزية لـ aklivity

    aklivity/zilla

    690عرض على GitHub↗

    🦎 A multi-protocol edge & service proxy. Seamlessly interface web apps, IoT clients, & microservices to Apache Kafka® via declaratively defined, stateless APIs.

    Java
    عرض على GitHub↗690
  • mozilla-services/hekaالصورة الرمزية لـ mozilla-services

    mozilla-services/heka

    3,403عرض على GitHub↗

    DEPRECATED: Data collection and processing made easy.

    Go
    عرض على GitHub↗3,403
  • bruin-data/bruinالصورة الرمزية لـ bruin-data

    bruin-data/bruin

    1,620عرض على GitHub↗

    Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.

    Goanalyticsbigquerydata-analysis
    عرض على GitHub↗1,620

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Find more with AI search
  • netflix/suroالصورة الرمزية لـ Netflix

    Netflix/suro

    796عرض على GitHub↗

    Netflix's distributed Data Pipeline

    Java
    عرض على GitHub↗796
  • papertrail/kestrelP

    papertrail/kestrel

    0عرض على GitHub↗
    عرض على GitHub↗0
  • pinterest/secorالصورة الرمزية لـ pinterest

    pinterest/secor

    1,858عرض على GitHub↗

    Secor is a service implementing Kafka log persistence

    Java
    عرض على GitHub↗1,858
  • facebookarchive/scribeالصورة الرمزية لـ facebookarchive

    facebookarchive/scribe

    3,911عرض على GitHub↗

    Scribe is a distributed log aggregation system designed to collect and route real-time log data from numerous servers to centralized storage or analysis tools. It functions as a log data pipeline and scalable collector that gathers streaming data and writes it to local disks or remote endpoints. The system employs a log routing server model that organizes incoming streams into specific buckets based on predefined configuration mappings. It supports multi-hop log forwarding, allowing data to be routed through a chain of intermediate servers to centralize logs from diverse network segments. Re

    C++
    عرض على GitHub↗3,911
  • rudderlabs/rudder-serverالصورة الرمزية لـ rudderlabs

    rudderlabs/rudder-server

    4,437عرض على GitHub↗

    Rudder Server is a customer data platform and event routing pipeline designed to collect, transform, and route customer event data from various sources to data warehouses and business tools. It functions as a customer identity resolver, linking identifiers from multiple sources to build a unified identity graph and comprehensive behavioral customer profiles. The system differentiates itself through reverse ETL capabilities, which push processed customer segments and audiences from data warehouses back into operational third-party applications. It also provides a containerized data plane for K

    Gobigquerycdpcustomer-data
    عرض على GitHub↗4,437
  • skizzehq/skizzeالصورة الرمزية لـ skizzehq

    skizzehq/skizze

    772عرض على GitHub↗

    A probabilistic data structure service and storage

    Go
    عرض على GitHub↗772
  • bruin-data/ingestrالصورة الرمزية لـ bruin-data

    bruin-data/ingestr

    3,714عرض على GitHub↗

    ingestr is a command-line tool for copying and syncing data between different database engines and third-party platforms without writing custom code. It functions as an ETL pipeline utility that extracts data from diverse sources and loads it into destinations. The tool features a schema-agnostic data loader that maps source fields to destination columns dynamically, removing the need for predefined static table definitions. It also operates as an incremental data synchronizer, updating destination tables by appending new records or merging changes to maintain current datasets. The system pr

    Go
    عرض على GitHub↗3,714
  • sonalgoyal/hihoالصورة الرمزية لـ sonalgoyal

    sonalgoyal/hiho

    92عرض على GitHub↗

    Hadoop Data Integration with various databases, ftp servers, salesforce. Incremental update, dedup, append, merge your data on Hadoop.

    Java
    عرض على GitHub↗92
  • gazette/coreالصورة الرمزية لـ gazette

    gazette/core

    793عرض على GitHub↗

    Build platforms that flexibly mix SQL, batch, and stream processing paradigms

    Go
    عرض على GitHub↗793
  • apache/pulsarالصورة الرمزية لـ apache

    apache/pulsar

    15,276عرض على GitHub↗

    Apache Pulsar is a cloud-native distributed pub-sub messaging system designed for high-performance data ingestion. It functions as a geo-replicated data streamer and a multi-tenant event streaming platform, providing a serverless stream processing engine and a tiered storage messaging broker. The system distinguishes itself by separating serving layers from storage layers to allow independent scaling of compute and data retention. It features native geo-replication to synchronize messages across different geographical regions and employs a multi-layered tenant isolation model using authentica

    Java
    عرض على GitHub↗15,276
  • apache/iggyالصورة الرمزية لـ apache

    apache/iggy

    4,382عرض على GitHub↗

    Iggy is a distributed message streaming platform and multi-protocol message broker that functions as a persistent distributed log store. It provides infrastructure for publishing and consuming binary messages using an append-only log, ensuring high availability and data consistency across nodes through Viewstamped Replication. The platform is distinguished by its specialized LLM streaming infrastructure, which uses a server protocol to connect large language models to streaming data and system controls. This includes standardized protocols for context management and data bridging via HTTP or

    Rustapachehttpiggy
    عرض على GitHub↗4,382
  • airbnb/kafkatالصورة الرمزية لـ airbnb

    airbnb/kafkat

    502عرض على GitHub↗

    KafkaT-ool

    Ruby
    عرض على GitHub↗502
  • yahoo/kafka-managerالصورة الرمزية لـ yahoo

    yahoo/kafka-manager

    11,926عرض على GitHub↗

    Kafka Manager is a web-based management interface and monitoring tool for Apache Kafka clusters. It serves as a central control plane for topic administration, consumer monitoring, and cluster health inspection. The project provides specialized utilities for data rebalancing and partition reassignment to distribute workloads across brokers. It also includes tools to optimize partition leadership by electing preferred replicas. The platform covers a broad range of administrative capabilities, including the creation and configuration of message topics, tracking of consumer offsets, and the col

    Scala
    عرض على GitHub↗11,926
  • apache/incubator-gobblinالصورة الرمزية لـ apache

    apache/incubator-gobblin

    2,267عرض على GitHub↗

    A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

    Java
    عرض على GitHub↗2,267
  • awslabs/aws-data-wranglerالصورة الرمزية لـ awslabs

    awslabs/aws-data-wrangler

    4,107عرض على GitHub↗

    This project is an AWS pandas integration library and data pipeline framework designed to simplify the movement and transformation of data between local memory and AWS storage and analytics services. It functions as a cloud data lake toolkit and storage file manager, allowing users to read, write, and transform structured data across various cloud environments. The library distinguishes itself as a distributed compute orchestrator capable of managing clusters in environments such as EMR to process datasets that exceed the memory limits of a single machine. It also provides specialized capabil

    Python
    عرض على GitHub↗4,107
  • bahador-r/db2lakeالصورة الرمزية لـ bahador-r

    bahador-r/db2lake

    2عرض على GitHub↗
    TypeScript
    عرض على GitHub↗2
  • confluentinc/bottledwater-pgC

    confluentinc/bottledwater-pg

    0عرض على GitHub↗
    عرض على GitHub↗0
  • dataspoclab/dataspoc-pipeالصورة الرمزية لـ dataspoclab

    dataspoclab/dataspoc-pipe

    2عرض على GitHub↗

    Data ingestion engine — Singer taps to Parquet in cloud buckets

    Python
    عرض على GitHub↗2
  • edenhill/kafkacatالصورة الرمزية لـ edenhill

    edenhill/kafkacat

    5,761عرض على GitHub↗

    Kafkacat is a suite of command-line utilities for interacting with Apache Kafka clusters. It provides a non-JVM binary for producing and consuming messages, inspecting cluster metadata, and debugging the Kafka protocol via the terminal. The tool functions as a producer and consumer capable of pushing data from files or standard input and reading messages from specific topics and partitions. It includes a metadata inspector to retrieve cluster state and partition configurations in plain text or JSON, as well as a protocol debugger for inspecting message offsets, timestamps, and binary payloads

    C
    عرض على GitHub↗5,761
  • edenhill/librdkafkaالصورة الرمزية لـ edenhill

    edenhill/librdkafka

    991عرض على GitHub↗

    The Apache Kafka C/C++ library

    C
    عرض على GitHub↗991
  • fulldecent/google-sheets-etlالصورة الرمزية لـ fulldecent

    fulldecent/google-sheets-etl

    22عرض على GitHub↗

    Live import all your Google Sheets to your data warehouse

    PHP
    عرض على GitHub↗22
  • kreuzberg-dev/kreuzbergالصورة الرمزية لـ kreuzberg-dev

    kreuzberg-dev/kreuzberg

    8,527عرض على GitHub↗

    Kreuzberg is a document extraction engine that converts PDFs, Office files, images, and over 90 other formats into clean, structured text and metadata. It is built around a compiled Rust core that can be used as a native library, a command-line tool, a REST API server, or a WebAssembly module for browser-based processing. The system is designed to run entirely on self-hosted infrastructure, with no data leaving the user's environment. What distinguishes Kreuzberg is its breadth of integration surfaces and its pipeline architecture. It exposes extraction capabilities through native bindings fo

    Rustdocument-intelligenceelixirffi
    عرض على GitHub↗8,527
  • kroxylicious/kroxyliciousالصورة الرمزية لـ kroxylicious

    kroxylicious/kroxylicious

    288عرض على GitHub↗

    Kroxylicious, the snappy open source proxy for Apache Kafka®

    Java
    عرض على GitHub↗288
  • mgillr/crdt-mergeالصورة الرمزية لـ mgillr

    mgillr/crdt-merge

    3عرض على GitHub↗

    Conflict-free merge for DataFrames, JSON, ML models & distributed agents — powered by CRDTs. The first merge library where every operation is mathematically guaranteed to converge.

    Python
    عرض على GitHub↗3