awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
linkedin avatar

linkedin/gobblin

0
View on GitHub↗
2,267 stars·749 forks·Java·Apache-2.0·15 viewsgobblin.apache.org↗

Gobblin

A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

Features

  • Big Data Frameworks - Universal data ingestion framework for Hadoop.
  • Data Ingestion - Universal framework for data ingestion and integration.
  • Data Ingestion and Integration - Universal framework for data ingestion into Hadoop.
  • Data Ingestion Pipelines - Universal framework for data ingestion.

Star history

Star history chart for linkedin/gobblinStar history chart for linkedin/gobblin

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Gobblin

These projects share indexed features with Gobblin. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • netflix/suroNetflix avatar

    Netflix/suro

    796View on GitHub↗

    Netflix's distributed Data Pipeline

    Java
    View on GitHub↗796
  • apache/pulsarapache avatar

    apache/pulsar

    15,276View on GitHub↗

    Apache Pulsar is a cloud-native distributed pub-sub messaging system designed for high-performance data ingestion. It functions as a geo-replicated data streamer and a multi-tenant event streaming platform, providing a serverless stream processing engine and a tiered storage messaging broker. The system distinguishes itself by separating serving layers from storage layers to allow independent scaling of compute and data retention. It features native geo-replication to synchronize messages across different geographical regions and employs a multi-layered tenant isolation model using authentica

    Java
    View on GitHub↗15,276
  • aklivity/zillaaklivity avatar

    aklivity/zilla

    690View on GitHub↗

    🦎 A multi-protocol edge & service proxy. Seamlessly interface web apps, IoT clients, & microservices to Apache Kafka® via declaratively defined, stateless APIs.

    Java
    View on GitHub↗690
  • bruin-data/bruinbruin-data avatar

    bruin-data/bruin

    1,620View on GitHub↗

    Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.

    Goanalyticsbigquerydata-analysis
    View on GitHub↗1,620
Compare all 30 related projects→

Frequently asked questions

What does linkedin/gobblin do?

A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems.

What are the main features of linkedin/gobblin?

The main features of linkedin/gobblin are: Big Data Frameworks, Data Ingestion, Data Ingestion and Integration, Data Ingestion Pipelines.

Which projects share features with linkedin/gobblin?

Projects with overlapping indexed features include: netflix/suro — Netflix's distributed Data Pipeline. bruin-data/ingestr — ingestr is a command-line tool for copying and syncing data between different database engines and third-party… aklivity/zilla — 🦎 A multi-protocol edge & service proxy. Seamlessly interface web apps, IoT clients, & microservices to Apache Kafka®… apache/pulsar — Apache Pulsar is a cloud-native distributed pub-sub messaging system designed for high-performance data ingestion. It… bruin-data/bruin — Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end… facebookarchive/scribe — Scribe is a distributed log aggregation system designed to collect and route real-time log data from numerous servers…