30 resources · 27 free
Data Engineering
Build data pipelines and warehouses — ETL, big data, and production data systems.
Data Engineering Zoomcamp
A free 9-week course by DataTalksClub that teaches data engineering fundamentals by building an end-to-end data pipeline from scratch. You learn Docker, SQL, dbt, Spark, Kafka, and orchestration through video lectures, homework, and a course platform with deadlines. Aimed at beginners and career switchers who want production-ready pipeline skills. Backed by a large community with 44k+ GitHub stars and active Slack support.
dbt Learn
dbt's official learning platform with free courses such as dbt Fundamentals, which teaches analytics engineering with hands-on practice in a hosted environment. You learn to build and test data models, write SQL in Jinja, manage transformations, and deploy dbt projects. Designed for analysts and data engineers working with modern data warehouses. Self-paced, browser-based, and free to enroll.
Apache Kafka Documentation
The official documentation for Apache Kafka, the leading open-source event streaming platform. It covers core concepts like topics, partitions, producers, consumers, and Kafka Streams, plus a quickstart that gets you running a local broker and hands-on with real events. An essential reference and study resource for anyone learning streaming data engineering. Maintained by the Apache Kafka project with documentation for all recent versions.
The Data Engineering Cookbook
A free ebook by Andreas Kretz that teaches the fundamentals of becoming a data engineer. It walks through core data engineering concepts, the day-to-day tasks of the role, and the tooling landscape, with practical recipes and career guidance. Hosted on GitHub with 15k+ stars under an Apache-2.0 license. Suited to newcomers and juniors mapping out a self-study path.
Data Engineering Wiki
A community-built knowledge base for learning data engineering, with interlinked notes on core concepts, an FAQ, practical guides such as Getting Started With Data Engineering, and a curated learning-resources collection. Great for building intuition about pipelines, batch versus streaming, OLTP versus OLAP, and data modeling. Maintained by the data engineering community and mirrored on GitHub.
Introduction to Data Engineering (DataCamp)
DataCamp's flagship introductory course on data engineering, covering the role of the data engineer, pipelines, ETL/ELT, data warehouses, and batch versus streaming processing. You learn through short videos and interactive coding exercises, with the first chapter free and the full course behind a subscription. Best for complete beginners and analysts moving toward engineering roles. DataCamp is one of the largest data-science learning platforms, with thousands of companies using its curriculum.
Google Cloud Skills Boost
Google Cloud's official hands-on training platform hosting hundreds of courses and interactive labs across BigQuery, Dataflow, Pub/Sub, Dataproc, and data analytics. Many courses are free, and free monthly credits let you run real labs in live cloud environments. You learn by doing, which suits aspiring data engineers and analysts preparing for Google Cloud certifications such as Professional Data Engineer. Labs scale from beginner to advanced. Free with limits.
Get Started with Data Engineering on Azure
Microsoft Learn's official learning path for Azure data engineering fundamentals, covering data ingestion, storage, transformation, and visualization with services like Azure Data Factory and Azure Synapse. Modules combine short readings with free interactive sandboxes you can launch right in the browser. Entirely free and self-paced. Designed for beginners starting their Azure data engineering journey and a stepping stone toward the DP-203 certification.
dbt Fundamentals (dbt Learn)
Free official course from dbt Labs that teaches analytics engineering: models, sources, tests, documentation, and deployment using dbt Core and dbt Cloud. Lessons mix short videos with hands-on practice projects and a final quiz you can pass to earn a certificate. The core curriculum has no cost. Great for analysts and data engineers who want to build version-controlled, testable transformation pipelines on top of a warehouse.
Data Engineering Wiki
Community-maintained wiki of the data engineering community and subreddit, curating learning resources, book recommendations, course lists, tool comparisons, and career guidance vetted by practicing engineers. Free to read and open source on GitHub. A solid first stop when deciding what to learn next or comparing tools and certifications. Kept current by the r/dataengineering community and its maintainers.
Airbyte Tutorials
Airbyte's official tutorial library teaching data integration and ELT in practice: connecting sources and destinations, configuring syncs, incremental loading, and orchestrating pipelines with Airflow and dbt. Step-by-step guides include setup instructions and real connector examples you can follow with a free Airbyte Cloud trial or the open-source edition. Suited to engineers learning modern data movement tooling. Airbyte is one of the most widely adopted open-source ELT platforms.
Snowflake Quickstarts
Snowflake's official library of hands-on quickstart guides covering warehouse fundamentals, loading and transforming data, Snowpark, data sharing, and AI/ML workflows. Each quickstart provides step-by-step instructions and sample code you can run against a free trial account. Ideal for engineers and analysts learning the Snowflake data cloud from scratch. Guides are kept current by Snowflake and range from beginner to advanced, with a companion GitHub repository of assets.
Confluent Developer Courses
Free official courses from Confluent on Apache Kafka and stream processing, including Kafka fundamentals, Kafka Streams, ksqlDB, Connect, and schema management. Each course combines video lessons with hands-on labs and quizzes, and you can earn certificates for free. Built by the company behind Kafka's commercial distribution, so content reflects real production patterns. Excellent for engineers building streaming data pipelines and event-driven systems.
Prove your interest
Unlock this bonus resource by showing you are serious.
Data with Zach (YouTube Channel)
YouTube channel of Zach Wilson, a former Meta and Netflix data engineer, featuring a deep data engineering interview preparation series, SQL walkthroughs, resume reviews, and honest career guidance. Playlists organize content by topic, and live sessions are free to join. Aimed at aspiring and junior data engineers targeting top-tier roles. One of the most-followed data engineering educators on the platform.
Prove your interest
Unlock this bonus resource by showing you are serious.
The Seattle Data Guy (Blog)
Blog of Ben Rogojan, a senior data engineer, offering tutorials and essays on data engineering, machine learning pipelines, cloud architecture, and career growth. Posts walk through real production architectures with code and diagrams, plus pragmatic advice on breaking into the field. Free to read, with a companion newsletter. Best for practitioners who want concrete examples and industry perspective from someone who has built large-scale data systems.
Prove your interest
Unlock this bonus resource by showing you are serious.
Data Engineering Podcast
Weekly podcast hosted by Tobias Macey featuring in-depth interviews with data engineers and platform teams about pipelines, orchestration, storage engines, and data platform design. Episodes are free on every major podcast app, and the website includes show notes and transcripts. Great for staying current on real-world tooling and design tradeoffs while commuting. One of the longest-running shows dedicated specifically to data engineering.
Prove your interest
Unlock this bonus resource by showing you are serious.
Google Cloud Skills Boost — Data Engineer Path
The Professional Data Engineer learning path on Google Cloud Skills Boost is a curated program of on-demand courses, interactive labs, and skill badges for designing and operating data systems on Google Cloud. It teaches building data pipelines, data lakes and warehouses, BigQuery, Dataflow, and operationalizing machine learning models, with hands-on labs in real cloud environments. Prepares learners for the Google Cloud Professional Data Engineer certification. Free with limits.
Prove your interest
Unlock this bonus resource by showing you are serious.
Snowflake Developer Guides
Snowflake Developer Guides are interactive, hands-on tutorials that walk you through core Snowflake skills — creating stages, tables, and warehouses, loading structured and semi-structured data, cloning, Time Travel, and role-based access — inside a free 30-day trial account. Each guide is a self-paced lab with prerequisites and progress tracking, covering data engineering, streaming, and notebooks topics. Free to use; suited to developers and data engineers learning modern data warehousing.
Prove your interest
Unlock this bonus resource by showing you are serious.
Apache Airflow Tutorial
Official step-by-step tutorial from the Apache Airflow project: install Airflow, write your first DAG in Python, and run task pipelines with the scheduler. Teaches core concepts like tasks, dependencies, and operators through a complete walkable example. For aspiring data engineers building their first orchestration pipelines. The canonical starting point for the industry-standard workflow tool.
Prove your interest
Unlock this bonus resource by showing you are serious.
Databricks Academy
Free self-paced training hub from the company that created Apache Spark, with a full Data Engineer path: data ingestion with Delta Lake, building pipelines with Spark, declarative pipelines, and governance with Unity Catalog. Hands-on labs run on a free Databricks account, with optional badges. For beginners and intermediates learning lakehouse data engineering. Free on-demand courses built by the Spark founding team.
Prove your interest
Unlock this bonus resource by showing you are serious.
Data Engineering Weekly
Free weekly newsletter curating the best data engineering articles, tooling releases, and discussions on Airflow, Spark, dbt, streaming, warehouses, and lakehouses. Each issue links deep-dive posts and community debates so you keep pace with the field. For data engineers and analysts who want a steady learning diet without searching. A compact way to stay current on modern data stack practice.
Prove your interest
Unlock this bonus resource by showing you are serious.
Kimball Group
The definitive free resource on dimensional modeling from the creators of the Kimball methodology: articles and design tips on star schemas, fact and dimension tables, slowly changing dimensions, and data warehouse architecture. Learn the modeling patterns behind most production warehouses. For data engineers and BI professionals designing analytics-ready schemas. The canonical reference for data warehouse design.
Prove your interest
Unlock this bonus resource by showing you are serious.
Streaming 101: The World Beyond Batch
Tyler Akidau's classic free essay clarifying streaming terminology: bounded vs unbounded data, event time vs processing time, and windowing strategies for real-time pipelines. Written by the engineer behind Google's MillWheel and Cloud Dataflow. For data engineers designing streaming systems or moving from batch to real-time. The most-cited conceptual foundation for stream processing.
Prove your interest
Unlock this bonus resource by showing you are serious.
Delta Lake
Official documentation and learning resources for Delta Lake, the open-source storage layer that turns data lakes into lakehouses with ACID transactions, schema enforcement, and time travel. Tutorials walk through creating tables, reading and writing data, and optimizing pipelines. For data engineers building reliable data lake platforms. The standard foundation for modern lakehouse architecture.
Prove your interest
Unlock this bonus resource by showing you are serious.
Dagster University
Free structured courses from Dagster Labs on asset-based data orchestration: Dagster Essentials, AI-Driven Data Engineering, and Dagster and ETL with quizzes, labs, and certificates. You learn to model pipelines as data assets and orchestrate production data platforms. For data engineers moving beyond cron-style scheduling. Hands-on, self-paced, and completely free.
Prove your interest
Unlock this bonus resource by showing you are serious.
dbt Community
Free community hub for the dbt ecosystem: Slack workspace, Discourse forums, local meetups, and demo days where practitioners share analytics-engineering patterns. Ask questions about models, tests, and deployments and learn from thousands of working data professionals. For analysts and engineers using or learning dbt. The fastest way to get expert answers and see real-world dbt practice.
Prove your interest
Unlock this bonus resource by showing you are serious.
Great Expectations Docs
Official documentation for Great Expectations, the open-source data quality and pipeline observability framework: tutorials on validating data, writing expectations, and building data docs. Learn how to catch bad data before it reaches downstream consumers. For data engineers adding quality gates to pipelines. The leading free resource on data validation and pipeline observability.
Prove your interest
Unlock this bonus resource by showing you are serious.
SparkByExamples
Large free tutorial library for Apache Spark and PySpark: hundreds of worked examples on RDDs, DataFrames, SQL, streaming, and performance tuning, each with copy-paste code. Learn Spark by studying real, runnable snippets rather than theory. For developers and data engineers learning Spark in Python or Scala. A practical complement to official docs, with no account or paywall.
Prove your interest
Unlock this bonus resource by showing you are serious.
Redpanda University
Free training platform on event streaming and Kafka-compatible data pipelines: courses on streaming fundamentals, Kafka building blocks, producers and consumers in Python, Java, and Node, and stream processing with hands-on exercises. Register once to unlock the whole library. For developers and data engineers moving into real-time data work. Free, interactive, and built for hands-on learning.
Prove your interest
Unlock this bonus resource by showing you are serious.
Apache Flink
Official documentation and learning materials for Apache Flink, the leading open-source stream-processing engine: concepts, tutorials, and examples on event-time processing, windows, state, and exactly-once semantics. Learn to build real-time analytics and streaming pipelines at scale. For data engineers working with continuous data. The canonical reference for stateful stream processing.
Prove your interest
Unlock this bonus resource by showing you are serious.
18 more resources in Data Engineering
MMIC members see the full catalog — every resource, no limits.