awesome-repositories.com
Blog
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectAboutHow we rankPressMCP server
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
DataTalksClub avatar

DataTalksClub/data-engineering-zoomcamp

0
View on GitHub↗
42,483 stars·8,411 forks·Jupyter Notebook·22 viewsairtable.com/appzbS8Pkg9PL254a/shr6oVXeQvSI5HuWD↗

Data Engineering Zoomcamp

This project is an open-source educational curriculum designed to provide comprehensive training in data engineering. It focuses on building scalable data pipelines and managing cloud-native infrastructure through a structured, self-paced program that combines technical explanations with hands-on practical exercises.

The curriculum distinguishes itself by emphasizing industry-standard methodologies, specifically teaching students how to implement infrastructure as code and manage data workflows through orchestration tools. By utilizing container-based environment isolation and declarative configuration, the program ensures that learners gain experience with reproducible deployments and consistent development environments across distributed systems.

The training covers a broad range of technical topics, including the design of automated data processing tasks and the configuration of cloud resources. The materials are organized into modular, progressive units that build foundational knowledge before advancing to complex engineering workflows.

The course materials are hosted in a centralized repository, which facilitates community-supported updates and collaborative improvements to the educational assets.

Features

  • Data Engineering Curricula - A comprehensive technical syllabus focused on building scalable data pipelines, managing cloud infrastructure, and mastering modern distributed computing workflows.
  • Data Engineering - Focuses on building scalable data pipelines and storage systems using modern cloud infrastructure.
  • Data Pipeline Architectures - Designing and managing automated workflows that handle the movement, transformation, and scheduling of data across complex distributed systems.
  • Cloud Infrastructure Courses - A practical guide to provisioning and managing cloud resources using declarative configuration files and containerized execution environments for data-intensive applications.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI
  • Technical Training - Provides structured guides and hands-on practice to build professional data engineering skills.
  • Data Pipeline Orchestrators - "Teaches the use of automated workflow tools to schedule, monitor, and manage the execution of complex data processing tasks."
  • Infrastructure as Code - Automates cloud resource provisioning using declarative configuration files to ensure reproducible deployments.
  • Open-Source Learning Programs - A structured collection of learning materials and practical exercises designed to teach technical skills through a self-paced, community-supported program.
  • CI/CD and Orchestration - Data engineering course covering dbt.
  • Data Engineering - Course on data engineering fundamentals.
  • Infrastructure as Code Practices - Automating the provisioning and management of cloud resources through declarative configuration files to ensure consistent and reproducible deployments.
  • Curricula - Organizes complex technical topics into progressive, modular learning units.
  • Containerization - Teaches the use of container images to ensure consistent execution environments across development and production.
  • Development Environments - Ensures consistent execution environments by isolating dependencies within portable containers.
  • Curriculum Modules - Module 1: Containerization and Infrastructure as Code - Introduction to GCP - Docker and Docker Compose - Running PostgreSQL with Docker - Infrastructure setup with Terraform - Homework #### Module 2: Workflow Orche
  • Professional Development - Provides hands-on training in industry-standard tools for real-world engineering tasks.
  • Star history

    Star history chart for datatalksclub/data-engineering-zoomcampStar history chart for datatalksclub/data-engineering-zoomcamp

    Frequently asked questions

    What does datatalksclub/data-engineering-zoomcamp do?

    This project is an open-source educational curriculum designed to provide comprehensive training in data engineering. It focuses on building scalable data pipelines and managing cloud-native infrastructure through a structured, self-paced program that combines technical explanations with hands-on practical exercises.

    What are the main features of datatalksclub/data-engineering-zoomcamp?

    The main features of datatalksclub/data-engineering-zoomcamp are: Data Engineering Curricula, Data Engineering, Data Pipeline Architectures, Cloud Infrastructure Courses, Technical Training, Data Pipeline Orchestrators, Infrastructure as Code, Open-Source Learning Programs.

    What are some open-source alternatives to datatalksclub/data-engineering-zoomcamp?

    Open-source alternatives to datatalksclub/data-engineering-zoomcamp include: dagster-io/dagster — Dagster is a data orchestration platform designed to manage the entire lifecycle of data assets through declarative… dataexpert-io/data-engineer-handbook — This project is a comprehensive, community-driven knowledge base designed to support individuals pursuing careers in… andkret/cookbook — Cookbook is a comprehensive knowledge base and reference repository for data engineering. It serves as a centralized… apache/airflow — Airflow is a platform for programmatically authoring, scheduling, and monitoring complex data pipelines. It functions… kestra-io/kestra — Kestra is a declarative workflow orchestrator designed to manage complex task dependencies and automated processes… prefecthq/prefect — Prefect is a workflow orchestration platform designed to define, schedule, and monitor complex data pipelines as…

    Open-source alternatives to Data Engineering Zoomcamp

    Similar open-source projects, ranked by how many features they share with Data Engineering Zoomcamp.
    • dagster-io/dagsterdagster-io avatar

      dagster-io/dagster

      14,974View on GitHub↗

      Dagster is a data orchestration platform designed to manage the entire lifecycle of data assets through declarative modeling and version-controlled code. It functions as a workflow engine that treats data assets as first-class primitives, allowing teams to define, schedule, and monitor complex pipelines while maintaining clear visibility into lineage, dependencies, and data quality. The platform distinguishes itself by using a code-as-configuration framework that enables standard software engineering practices, such as unit testing and local mocking, to be applied directly to data workflows.

      Pythonanalyticsdagsterdata-engineering
      View on GitHub↗14,974
    • dataexpert-io/data-engineer-handbookDataExpert-io avatar

      DataExpert-io/data-engineer-handbook

      41,758View on GitHub↗

      This project is a comprehensive, community-driven knowledge base designed to support individuals pursuing careers in data engineering. It functions as a centralized learning hub that aggregates industry best practices, technical documentation, and educational resources to assist with both professional development and the design of robust data pipeline architectures. The repository distinguishes itself by providing a structured technical career roadmap that includes curated learning paths, interview preparation strategies, and practical project examples. By indexing a diverse range of media—in

      Jupyter Notebookapachesparkawesomebigdata
      View on GitHub↗41,758
    andkret/cookbookandkret avatar

    andkret/Cookbook

    15,161View on GitHub↗

    Cookbook is a comprehensive knowledge base and reference repository for data engineering. It serves as a centralized directory for data architecture patterns, professional career roadmaps, and a curated collection of public datasets. The project provides a structured guide for transitioning into specialized data engineering roles through skill-matrix mapping and technical interview preparation. It further distinguishes itself by documenting real-world industry case studies and decomposing large-scale industrial implementations into repeatable architectural patterns. The repository covers a b

    Python
    View on GitHub↗15,161
  • apache/airflowapache avatar

    apache/airflow

    45,902View on GitHub↗

    Airflow is a platform for programmatically authoring, scheduling, and monitoring complex data pipelines. It functions as a workflow automation engine that manages the lifecycle of recurring business processes by executing code-defined task dependencies. By representing workflows as directed acyclic graphs, the system ensures that task execution order and data flow are explicitly defined and reliably maintained across distributed computing environments. The platform distinguishes itself through a highly modular, provider-based architecture that decouples core orchestration logic from external

    Pythonairflowapacheapache-airflow
    View on GitHub↗45,902
  • See all 30 alternatives to Data Engineering Zoomcamp→