awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com

Data engineering roadmap

Ranking updated Jul 30, 2026

For data engineering resources, the first results are igorbarinov/awesome-data-engineering (This repository is a comprehensive, curated awesome-list containing a vast collection of data engineering tools, learning resources, architectures, and tutorials that directly match the search intent), datatalksclub/data-engineering-zoomcamp (This repository is a comprehensive, open-source educational curriculum for learning data engineering that covers major tools like Spark, Kafka, and dbt through hands-on pipeline and infrastructure projects) and datastacktv/data-engineer-roadmap (This repository is a comprehensive, curated roadmap and learning resource directory that covers data pipelines, big data processing, data warehousing, and infrastructure topics tailored specifically for data engineering). data-engineering-community/data-engineering-wiki and dataexpert-io/data-engineer-handbook round out the shortlist. Compare the match explanations and check the project documentation against your requirements.

Hand-picked data engineering resources and roadmaps to help you learn skills, compare tools, and pick the right tools.

Data engineering roadmap

Find the best repos with AI.We'll search the best matching repositories with AI.
  • igorbarinov/awesome-data-engineeringigorbarinov avatar

    igorbarinov/awesome-data-engineering

    8,306View on GitHub↗

    This repository is a comprehensive, curated awesome-list containing a vast collection of data engineering tools, learning resources, architectures, and tutorials that directly match the search intent.

    Big Data
    View on GitHub↗8,306
  • datatalksclub/data-engineering-zoomcampDataTalksClub avatar

    DataTalksClub/data-engineering-zoomcamp

    42,483View on GitHub↗

    This project is an open-source educational curriculum designed to provide comprehensive training in data engineering. It focuses on building scalable data pipelines and managing cloud-native infrastructure through a structured, self-paced program that combines technical explanations with hands-on practical exercises. The curriculum distinguishes itself by emphasizing industry-standard methodologies, specifically teaching students how to implement infrastructure as code and manage data workflows through orchestration tools. By utilizing container-based environment isolation and declarative con

    This repository is a comprehensive, open-source educational curriculum for learning data engineering that covers major tools like Spark, Kafka, and dbt through hands-on pipeline and infrastructure projects.

    Jupyter NotebookData Pipeline Orchestrators
    View on GitHub↗42,483
  • datastacktv/data-engineer-roadmapdatastacktv avatar

    datastacktv/data-engineer-roadmap

    12,747View on GitHub↗

    This project is a collection of specialized study guides and roadmaps centered on computer science, data engineering, and machine learning fundamentals. It provides a structured curriculum of technical competencies, tools, and skills required to transition into professional data engineering roles. The project features a data engineering skill map that visually organizes databases, processing architectures, and infrastructure tools. It also includes a machine learning learning path covering supervised and unsupervised learning techniques alongside model operations. The curriculum covers broad

    This repository is a comprehensive, curated roadmap and learning resource directory that covers data pipelines, big data processing, data warehousing, and infrastructure topics tailored specifically for data engineering.

    Career Development PathsComputer Science EducationHierarchical Knowledge Structures
    View on GitHub↗12,747
  • data-engineering-community/data-engineering-wikidata-engineering-community avatar

    data-engineering-community/data-engineering-wiki

    1,985View on GitHub↗

    The data engineering wiki is a crowdsourced knowledge base and reference guide assembled through collaborative contributions from practitioners. It functions as a structured repository of learning paths, architectural decision guides, and software evaluations for data systems, compiled from plain-text source markup files into a searchable static documentation site. The content is organized into strict conceptual hierarchies covering core engineering concepts, security and governance, and infrastructure tools. Contributors and readers can explore foundational architectural patterns, storage s

    This repository is a comprehensive community-driven wiki that curates tools, learning materials, and architectures covering data pipelines, storage, and processing for data engineering.

    CSSCommunity Knowledge BasesArchitectural Decision GuidesCommunity Curation Workflows
    View on GitHub↗1,985
  • dataexpert-io/data-engineer-handbookDataExpert-io avatar

    DataExpert-io/data-engineer-handbook

    41,758View on GitHub↗

    This project is a comprehensive, community-driven knowledge base designed to support individuals pursuing careers in data engineering. It functions as a centralized learning hub that aggregates industry best practices, technical documentation, and educational resources to assist with both professional development and the design of robust data pipeline architectures. The repository distinguishes itself by providing a structured technical career roadmap that includes curated learning paths, interview preparation strategies, and practical project examples. By indexing a diverse range of media—in

    This repository is a comprehensive, curated collection of tools, tutorials, architectures, and learning materials tailored specifically for data engineering.

    Jupyter NotebookAwesome ListData Engineering CurriculaData Architecture Patterns
    View on GitHub↗41,758
  • andkret/cookbookandkret avatar

    andkret/Cookbook

    15,161View on GitHub↗

    Cookbook is a comprehensive knowledge base and reference repository for data engineering. It serves as a centralized directory for data architecture patterns, professional career roadmaps, and a curated collection of public datasets. The project provides a structured guide for transitioning into specialized data engineering roles through skill-matrix mapping and technical interview preparation. It further distinguishes itself by documenting real-world industry case studies and decomposing large-scale industrial implementations into repeatable architectural patterns. The repository covers a b

    This repository provides a curated collection of data engineering knowledge bases, architecture patterns, and learning materials, though it focuses more on guides and roadmaps than an exhaustive tool directory.

    PythonData Architecture PatternsArchitectural Case StudiesArchitecture Reference Catalogs
    View on GitHub↗15,161
  • mrsuichuan/data-warehouse-learningMrSuiChuan avatar

    MrSuiChuan/data-warehouse-learning

    1,154View on GitHub↗

    Data warehouse learning is a reference implementation of a real-time stream processing system and open-source data lakehouse architecture. It combines stream processing engines, open lakehouse formats, and analytical data warehouses into a complete e-commerce data warehouse system built for both offline and real-time analytics pipelines. The project implements hybrid data warehouse architectures utilizing multi-layer storage models and stream-batch processing pipelines. It features change data capture pipelines that stream database transaction logs into messaging systems, progressive data tra

    This repository provides a comprehensive learning collection and practical code for building real-time and offline data warehouses, covering major big data processing and storage frameworks.

    JavaData WarehousingStreaming Data Lakehouses
    View on GitHub↗1,154
  • danielbeach/data-engineering-practicedanielbeach avatar

    danielbeach/data-engineering-practice

    2,726View on GitHub↗

    Data engineering practice repository providing tutorials, distributed processing engines, and Python data pipeline automation scripts. The system encompasses automated data validation, distributed compute aggregation, embedded columnar querying, lazy evaluation planning, partitioned storage export, and cloud storage retrieval. The capability surface covers cloud integration and storage, data engineering and pipelines, data processing and analytics, data quality and testing, database and storage, file management, and monitoring and observability.

    This repository provides a hands-on collection of Python data pipeline scripts, distributed processing engines, and validation tutorials rather than a directory of external links, but it covers many of the requested data engineering concepts.

    PythonDistributed Computing EnginesPractice Problem RepositoriesAnalytical Query Engines
    View on GitHub↗2,726
  • databricks/spark-the-definitive-guidedatabricks avatar

    databricks/Spark-The-Definitive-Guide

    3,099View on GitHub↗

    This project is an educational resource and technical manual for Apache Spark, focused on the architecture and practical application of large-scale data processing. It serves as a guide for big data engineering and distributed computing, covering the principles of parallel processing and fault-tolerant data distribution. The material provides instructional content on designing distributed ETL pipelines and implementing data analysis workflows. It includes tutorials for polyglot data processing, offering patterns and examples for using Python, Scala, and Java within a unified environment. The

    This repository is a comprehensive technical guide and educational resource specifically for Apache Spark rather than a broader directory of data engineering tools and frameworks.

    ScalaBig Data ProcessingDistributed ComputingETL Workflows
    View on GitHub↗3,099
  • orchest/orchestorchest avatar

    orchest/orchest

    4,138View on GitHub↗

    Orchest is a data pipeline orchestrator and containerized workflow manager. It provides a platform for designing, scheduling, and executing complex data processing sequences through a combination of a graphical interface and scripting. The platform distinguishes itself by using containers to manage software dependencies, ensuring consistent execution across different environments. It features a polyglot task scheduler capable of triggering jobs written in multiple programming languages and includes a version control system that tracks historical snapshots of project configurations and code.

    Orchest is a data pipeline orchestrator and workflow manager rather than a curated educational directory or resource collection for learning data engineering.

    TypeScriptData OrchestrationData Pipeline OrchestrationData Pipelines
    View on GitHub↗4,138
  • apache/beamapache avatar

    apache/beam

    8,612View on GitHub↗

    Apache Beam is a distributed data pipeline framework and unified data processing model designed to handle both bounded batch data and unbounded real-time streams. It provides a system for building scalable, data-parallel workflows that operate across compute clusters using a single programming model. The framework utilizes a cross-runner pipeline abstraction that decouples the data processing logic from the underlying execution backend, allowing the same pipeline to run on different distributed compute engines. It supports multi-language pipeline development by translating high-level code fro

    Apache Beam is a powerful distributed data processing framework and pipeline tool, but it is a specific software library rather than the curated directory of tools, tutorials, and learning materials the visitor is looking for.

    JavaDistributed ComputingDistributed Data Processing FrameworksData Pipelines
    View on GitHub↗8,612
Compare the top 10 at a glance
RepositoryStarsLanguageLicenseLast push
igorbarinov/awesome-data-engineering8.3K—cc0-1.0Feb 10, 2026
datatalksclub/data-engineering-zoomcamp42.5KJupyter Notebook—Jun 10, 2026
datastacktv/data-engineer-roadmap
12.7K
—
—
Jan 25, 2022
data-engineering-community/data-engineering-wiki2KCSSCC0-1.0May 28, 2026
dataexpert-io/data-engineer-handbook41.8KJupyter Notebook—Apr 2, 2026
andkret/cookbook15.2KPythonApache-2.0Jun 12, 2026
mrsuichuan/data-warehouse-learning1.2KJavaArtistic-2.0Apr 26, 2026
danielbeach/data-engineering-practice2.7KPython—Jan 8, 2025
databricks/spark-the-definitive-guide3.1KScalaotherAug 26, 2020
orchest/orchest4.1KTypeScriptApache-2.0Jun 6, 2023

Related searches

  • an open source framework for data pipelines
  • a framework for building scalable data pipelines
  • Data science resources
  • Rust data engineering
  • MLOps resources
  • a step-by-step path to becoming a data engineer
  • Software architecture resources
  • Developer resource lists