Software · Advanced

System Design Foundations

Design a service under load, growth and failure - and settle every argument with a running prototype you measured rather than a diagram you drew.

About this course

System design is usually taught with boxes and arrows and graded with adjectives. This course does neither. You build one service on your own lab machine — an order platform for a synthetic wholesale grocer — and every design decision in it is made, implemented, measured and written down, in that order. You start by writing requirements with objectives that have indicators, targets and windows, and a load estimate derived from the business rather than from a benchmark. Then you write a load generator that keeps its raw results, and you take a baseline. From then on, nothing is asserted. An index is added because the page query was measured before and after. Work is moved to a queue because the two designs were measured against each other with the partner answering slowly. A cache is added because the read-to-write ratio justifies it and the measurement shows what a hit actually saves — and the cache is proved to serve one database query for twelve concurrent readers rather than twelve. Failure is taught by causing it. You make the shipping partner hang and watch a request that has no deadline; you add the deadline, the bounded retry budget with jitter, the circuit breaker and a degraded answer, and then you make it hang again and watch the difference. You spike the service above its concurrency limit and require it to refuse work with `Retry-After` while liveness keeps answering. You let a message exhaust its retry budget into a dead-letter table and then replay it. You hold a lock open while a migration runs, and find out what a `lock_timeout` is for. You pause WAL replay on a read replica so the stale read is a fact rather than a race. You stop an instance in the middle of a polling test. The written work is graded as written work: requirements, decision records, a capacity plan that names cost levers instead of inventing cost figures, a change record, a runbook, a review of somebody else's flawed proposal where every finding is reproducible, and a design document. The last thing you build is an acceptance suite for your own design — and you run it against a deliberately broken copy of the service, because a suite that passes on both proves nothing.

Content time
20 h
Lessons
10
Certificate
Yes
on completion
Choose a career path

Lesson 1 is free. Enroll in a career path to access its full courses.

Lesson 1 is a free preview — read it without an account.

Software — the kind of infrastructure this course is practised on

Outline

Lessons

10 lessons · 20 h
  1. Lesson 1: Requirements, objectives and a baseline you measuredFree preview

    Turn a business brief into testable objectives and a sizing estimate, then build the smallest working prototype and measure it, so every later decision has something to argue against.

    2 h
  2. Lesson 2: Interfaces, idempotency and an index you can justify

    Design the write so a retry is safe and the read so a page is stable, then add the one index the page query needs - and prove it was worth adding by measuring before and after.

    1 h 50 min
  3. Lesson 3: Decoupling with an outbox and a work queue

    Move the confirmation off the request path with a transactional outbox and workers that claim rows with SKIP LOCKED, and measure what that actually bought the customer.

    2 h
  4. Lesson 4: Retries, backoff and dead letters

    Give the queue a retry budget with exponential backoff and jitter, somewhere for hopeless messages to go, and a way to bring them back - then cause the failure and watch all three work.

    1 h 40 min
  5. Lesson 5: Caching, and what a hit actually saves

    Add cache-aside with a TTL, invalidation on write and one flight per key, then prove what it bought in database work as well as in latency - and name the staleness you are accepting.

    1 h 40 min
  6. Lesson 6: Backpressure, load shedding and circuit breaking

    Overload the service on purpose and watch it absorb the damage, then add admission control, deadlines, a bounded retry budget, a circuit breaker and a degraded answer - and cause the failure again.

    2 h 10 min
  7. Lesson 7: Observability, and alerts that have been made to fire

    Make the logs machine-readable, carry one request id across the queue, expose the queue depth, and write alert rules - then cause each condition on purpose and watch the rule fire and clear.

    1 h 50 min
  8. Lesson 8: Consistency, migrations and recovery

    Lose an update on purpose and then make it impossible, migrate a live schema by expand and contract with a client reading throughout, find out what a lock timeout is for, and restore a backup rather than believing in it.

    2 h 20 min
  9. Lesson 9: Scaling out - instances, replicas and shards

    Put two instances behind a proxy and stop one during a polling test, add a streaming read replica and pause its replay so the stale read is a fact, and shard a table by key - including the query that sharding makes impossible.

    2 h 20 min
  10. Lesson 10: Design review, and the decision you can defend

    Review a partner's proposal so that every finding is reproducible, turn your own design's claims into an acceptance suite, and run that suite against a deliberately broken build to find out whether it means anything.

    2 h 10 min

Where it leads

Part of these career paths