Analytics · Intermediate
dbt and Data Pipelines
Rebuild a pipeline's transformation layer with dbt-core on PostgreSQL 16 — sources, models, tests, snapshots, docs, environments and CI — against a result you already know is right.
About this course
ULC-014 builds the Riverbank Media pipeline with hand-written SQL and a small Python runner. That is the right way to learn what a pipeline does, and the wrong way to maintain fifty models with a team. This elective rebuilds the *transformation* half of that same pipeline with **dbt-core 1.11** and the **dbt-postgres** adapter, on the same `riverbank` database, so you can see exactly what dbt adds — and prove it, because the pipeline's own mart is still sitting there to compare against. You install dbt into a virtual environment from a pinned offline wheelhouse, declare the pipeline's clean layer as sources, and build staging views and mart tables with `source()` and `ref()`. You make the fact table incremental and prove a second run inserts nothing. You snapshot a dimension, apply a change set to the source export, re-ingest, and watch the snapshot close a version and open a new one. You test every key and relationship, plant a fault that turns a test red, read the stored failures, fix it and watch it go green. You document every model and column, generate the lineage graph, and declare the exposure that says who reads the mart. Then you put it in a pipeline. You extend ULC-014's runner to call `dbt build` and record the phase in its load log — through a written change record with a rollback you have executed. You add a production target in its own schema, write a CI script, and prove it fails closed by planting a model that cannot compile. You measure a mart query, add post-hook indexes, and make the incremental strategy explicit. You enforce a contract on the fact table and configure source freshness. Every command, model, test and number in these lessons was executed on PostgreSQL 16 with dbt-core 1.11.15 and dbt-postgres 1.11.0, and the output you see is the output that was observed. The final project is an engagement: Riverbank Media asks you, as its first analytics engineer, to deliver and hand over the transformation layer — graded by automated checks against the warehouse your models actually built, and by a rubric against your handover.
- Content time
- 14 h 15 min
- Lessons
- 10
- Certificate
- Yes
- on completion
Lesson 1 is free. Enroll in a career path to access its full courses.
Lesson 1 is a free preview — read it without an account.

Outline
Lessons
Lesson 1: What dbt is, and the project skeletonFree preview
Install dbt-core and dbt-postgres from the lab's offline wheelhouse, scaffold a project, connect it to the Riverbank warehouse without writing a password anywhere, and put it under Git.
1 h 15 minLesson 2: Sources and staging models
Declare the pipeline's clean layer as a dbt source, build the three staging models with source(), and learn why staging is deliberately the dullest layer in the project.
1 h 20 minLesson 3: Marts and materializations
Build an ephemeral intermediate model, two dimensions, a fact and the published mart — then prove the mart agrees with the one the ULC-014 pipeline builds, row for row.
1 h 30 minLesson 4: Incremental models and snapshots
Make the fact incremental on a watermark, prove a second pass inserts nothing and keep the artefacts that prove it, then snapshot the user source and watch history appear when the export changes.
1 h 35 minLesson 5: Tests that fail closed
Test every key and relationship, store the failing rows, choose severity deliberately — then plant a fault, watch the suite go red, read the stored failures, fix it and prove it green.
1 h 35 minLesson 6: Documentation and lineage
Describe every model and column, generate the documentation site, declare the report that reads the mart as an exposure, and use the DAG to answer "what would this change break?
1 h 15 minLesson 7: Orchestrating dbt from the runner
Write the change record first, then make the ULC-014 pipeline call dbt build, record the phase in its own load log, fail the run when dbt fails — and execute the rollback to prove it exists.
1 h 25 minLesson 8: Environments, targets and CI
Add a production target in its own schema, write a CI script whose exit status is the contract, and prove it fails closed by planting a model that cannot compile.
1 h 25 minLesson 9: Performance of transformations
Measure two mart queries with EXPLAIN, compare what the merge and delete+insert strategies actually run, index the fact's keys from a post-hook, measure again — and write down what the numbers mean.
1 h 30 minLesson 10: Contracts, freshness and observability
Enforce a contract on the published fact and watch dbt refuse an undeclared column, measure whether the source is still arriving, and write down what is worth waking somebody for.
1 h 25 min
Where it leads