Database Platforms · Advanced

Data Architecture Foundations

Choose between the stores, the models and the movement patterns a data platform can be built from, and prove every choice on PostgreSQL 16 rather than asserting it.

About this course

You can run a database, build a pipeline and design a star. This course is about the decisions that come before all three: which store, which model, which integration pattern, what the retention and security design has to be, and how to write that down so another engineer can challenge it. You work on a three-store estate on your own lab machine — **Ultishop**, an order system of record; **Riverbank Media**, a streaming platform with a landing, clean, warehouse and quarantine pipeline; and **Meridian**, a finished retail star you cannot even read as yourself. Nothing is taken on trust. The inventory is collected from the stores' own catalogs. The storage-model argument is settled by modelling the same catalogue normalised and as documents and measuring what each costs, including the measurement that decides it, which is a write rather than a read. The warehouse-style argument is settled by building a small Data Vault and loading two billing snapshots through it. Then the platform gets built. Change data capture from the order system to a second PostgreSQL cluster, with a replication role that can read three tables and a password that never enters the catalog — and a seeded fault that stops the apply worker dead, so you diagnose a stalled stream from the catalog and the log rather than from a diagram. A partitioned open-format lake, read back through a foreign table. A catalog and column-level lineage that is validated against the database it describes. A certificate that a client can actually verify, field encryption whose key never enters the database, a masked view tested from both directions, and retention applied by dropping partitions — with a certificate change that stops the server from starting, because that is the one you want to have rehearsed. The last lesson is the work you will be given most often: somebody else's package, with a proof of concept that loads and returns correct numbers, and five claims in its document that are not true. You review it against a checklist and record every finding as a query anybody can re-run. The final project is an engagement for a four-site veterinary practice with clinical, personal, activity and financial data, delivered as an architecture package plus a proof of concept whose checks pass.

Content time
17 h 40 min
Lessons
10
Lab
Yes
provisioned for you
Certificate
Yes
on completion
Choose a career path

Lesson 1 is free. Enroll in a career path to access its full courses.

Lesson 1 is a free preview — read it without an account.

Database Platforms — the kind of infrastructure this course is practised on

Outline

Lessons

10 lessons · 17 h 40 min
  1. Lesson 1: Mapping the data estateFree preview

    What a data architect actually produces, and why the first artefact is a map built from the catalogs of the stores rather than from anybody's memory.

    1 h 25 min
  2. Lesson 2: Requirements and quality attributes

    Turning "we need it to be fast and safe" into a register of freshness, volume, retention, recovery, consistency and classification that the rest of the platform has to obey.

    1 h 25 min
  3. Lesson 3: Choosing a storage model

    Relational, document, key-value, columnar, time-series and graph — compared by building the same catalogue two ways in PostgreSQL and measuring what each one costs.

    1 h 45 min
  4. Lesson 4: Warehouse styles and the Data Vault

    Kimball, Inmon and Data Vault 2.0 compared by building hubs, a link and a satellite over Riverbank and loading two billing snapshots through them.

    1 h 45 min
  5. Lesson 5: Integration patterns and change data capture

    Batch, CDC, streaming and APIs compared, then logical replication built end to end from the order system to an analytics cluster — including the fault that stops it dead.

    2 h 5 min
  6. Lesson 6: Lake and lakehouse foundations

    Open formats, a partitioned layout that a path can prune, and a foreign table that reads one day of it from PostgreSQL — plus an honest account of what makes a lake a lakehouse.

    1 h 35 min
  7. Lesson 7: Governance, catalog and lineage

    Ownership that means something, definitions that travel with the objects they define, and column-level lineage recorded as data so it can be checked rather than believed.

    1 h 35 min
  8. Lesson 8: Security, privacy and retention

    The controls a restricted class implies — a certificate that verifies, field encryption with the key outside the database, a masked view over a least-privileged role, and deletion by dropping partitions.

    2 h 15 min
  9. Lesson 9: Scalability, resilience and retention

    A capacity method that starts with a measurement, partitioning that the planner actually prunes, and a destructive change that goes through a change record with a rollback that was run first.

    1 h 55 min
  10. Lesson 10: Decision records and the architecture review

    Writing decisions that can be challenged, migrating a published shape without breaking its consumers, and reviewing somebody else's package so that every finding is a query anybody can re-run.

    1 h 55 min

Hands-on

Your lab

Real virtual machines on the Ultiblob cluster, reached from your browser. You administer them; we provision and destroy them.

  1. vm-01db01linux

Provisioned for you when you launch the lab from the course. The machines are yours for the access window; release them and launch again whenever you like.

Where it leads

Part of these career paths