System Design Foundations · Lesson 1 of 10

Free preview

Requirements, objectives and a baseline you measured

Turn a business brief into testable objectives and a sizing estimate, then build the smallest working prototype and measure it, so every later decision has something to argue against.

time
2 h
on completion
+120 XP

Why this matters

You have been asked to design the order platform for Northreach Provisions, a wholesale grocer with about five hundred trade customers. The brief is three sentences long, which is normal: "orders come in bursts twice a day, every order has to be confirmed through our shipping partner, and the account managers keep the catalogue summary open all day".

Most system design goes wrong in the next hour. Somebody draws boxes. Somebody says the catalogue summary "needs a cache". Somebody else says the confirmations "should go on a queue". Both may be right, and neither of them knows yet, because nobody has asked what this system actually has to do or measured what it currently does.

This course takes the other route. In this lesson you will turn the brief into requirements that can be tested, make a sizing estimate that is explicitly an estimate, build the smallest prototype that does one real thing, and measure it. Everything after this is an argument with that measurement. When you add an index in lesson 2 or a cache in lesson 5, you will be able to say what it bought, on this machine, in numbers you can recompute from raw data. That is the difference between a design and a preference.

Concepts

Functional and non-functional requirements

A functional requirement says what the system does: a customer may place an order and receive a reference. A non-functional requirement says how well: the order is accepted within some time, nothing is lost, a retry does not create two orders. Functional requirements decide what you build. Non-functional requirements decide how you build it, and they are where system design lives.

Indicators, objectives and windows

An objective written as "the API should be fast" cannot be met or missed, so it will be argued about forever. A useful objective has three parts:

  • an indicator: something measurable, such as the 95th percentile latency of

GET /v1/catalogue/summary, measured at the service, over the requests it answered;

  • a target: 250 ms;
  • a window: over one hour.

Write the indicator precisely enough that two people measuring it get the same number. "Latency" measured at the client includes the network and the client's own overhead; measured at the service it does not. Neither is wrong, but a target attached to one of them is meaningless against the other.

One more decision belongs in the objectives, and it is easy to get backwards. When the service is overloaded and deliberately refuses a request, is that an error? In this design it is not: refusing work under load is the system behaving correctly, and an error budget that punishes it pushes the design towards queueing everybody into a timeout instead. So the error objective counts 5xx failures, and the refusals get an objective of their own that says they must always carry a Retry-After header.

Back-of-envelope sizing

Before measuring anything, estimate. The estimate is not a prediction, it is a way of finding out which numbers matter. For Northreach:

text
500 trade customers x 3 orders per working day   = 1,500 orders/day
1,500 orders spread over a 6-hour ordering window = 0.07 orders/second average
bursts at 09:00 and 16:00, assume 40x the average = ~3 orders/second peak
20 account managers polling the summary every 10s = 2 reads/second, 10 at busy times
an order row is ~120 bytes with indexes
1,500/day x 365 x 3 years                         = 1.6 million rows, ~400 MB

Two useful things fall out immediately. Storage is irrelevant here: 400 MB is nothing, and any design decision justified by "the data will get huge" is not justified. And the peak rate is small in absolute terms, which means the interesting risks are latency and failure, not volume. That is worth knowing before you spend a week on sharding.

Every one of those numbers is a guess derived from the business. Label them as guesses. The measurement you take in a moment is not a guess, and the difference must stay visible.

Percentiles, and why the method matters

An average latency hides everything you care about. If 99 requests take 10 ms and one takes 4 seconds, the average is 50 ms and one customer is furious. Percentiles answer the question "how bad is it for the unlucky ones": the p95 is the value that 95% of requests came in under.

There is more than one way to compute a percentile, and different tools choose differently. This course uses nearest rank: sort the latencies, and the p-th percentile is the value at index ceil(p / 100 * n) - 1, counting from zero. With 200 samples, p95 is the 190th value in sorted order. State the method wherever you publish a percentile. Numbers computed by different methods do not compare, and "our p99 is 80 ms" from two tools can differ by more than the improvement you are claiming.

The prototype

The prototype is a small HTTP service in Python, PostgreSQL in a container, and a load generator you write yourself. It uses only what the lab machine already has: the python:3.12-slim and postgres:16 images are baked into your lab image, and the one Python library the service needs, psycopg, is installed from an offline wheelhouse that ships with the machine. Nothing in this course downloads anything.

The service exposes /health (liveness, touches nothing), /ready (readiness, does touch the database), /metrics, and GET /v1/catalogue/summary, which aggregates 150,000 catalogue rows by category. That aggregate is deliberately the expensive thing: it is what the account managers poll, and it is what lesson 5 will cache.

Guided exercise

Work on linux01 as ubuntu. The full reference solution for every step is in the instructor material; the steps below are what you type.

  1. Create the project and install the one dependency offline.
bash
mkdir -p ~/northreach && cd ~/northreach
mkdir -p app ops sql logs results backups design/adr design/measurements design/changes
python3 -m pip install -q --no-index --find-links /opt/ultiblob/wheelhouse \
  --target vendor "psycopg[binary]==3.3.5"

--target vendor puts the library in a directory you can mount read-only into a container, which is why nothing needs to be installed inside the images.

  1. Write `.env.template` and generate `.env` from it. The template lists *every* switch this design will ever have, all of them off. A lesson turns one on only after measuring why it should be on, and a reviewer can read one file to see what the system can be asked to do. Never commit .env itself: it holds the database password.
  1. Write the service. app/config.py reads the switches; app/metrics.py keeps counters, gauges and a latency histogram and renders them in the Prometheus text exposition format; app/db.py is a small bounded connection pool; app/router.py, app/web.py and app/service.py are the HTTP machinery; app/routes_health.py and app/routes_catalogue.py are the two route modules.

The connection pool deserves a sentence. The database accepts a fixed number of connections, so the queue for one exists whether you write it or not. Writing it puts the queue somewhere you can measure it and give it a deadline, instead of leaving it inside the kernel where you cannot.

  1. Write `compose.yaml` and the schema, then start it.
bash
docker compose up -d
bash ops/wait-ready.sh http://127.0.0.1:8080/health 180
bash ops/migrate.sh
docker compose exec -T db psql -v ON_ERROR_STOP=1 -X -q -U "$DB_USER" -d "$DB_NAME" -f - < ops/seed.sql

Three small operations scripts are written here and used for the rest of the course. ops/wait-ready.sh polls until a URL answers and ops/wait-db.sh until a database accepts connections - never sleep 10 and hope, because a fixed sleep is a race that passes on your machine and fails on a loaded one. ops/deploy.sh reconciles the stack *and restarts the services you name*, because the application code is bind-mounted and docker compose up -d only replaces a container when its configuration changed.

  1. Write `loadgen.py`. It runs a fixed number of worker threads for a fixed number of seconds, each holding one keep-alive connection, and it writes two files: a raw CSV with one row per request (start time, latency, status, request id) and a summary JSON computed from that CSV. The keep-alive connection is not an optimisation. A generator that opens a new TCP connection per request spends more time in its own socket setup than the service spends answering, and then reports that time as the service's latency.
  1. Take the baseline.
bash
python3 loadgen.py --name 01-baseline --url http://127.0.0.1:8080/v1/catalogue/summary \
  --concurrency 6 --seconds 20

Observed on the authoring machine; yours will differ, and that is fine, because nothing in this course compares your numbers to mine:

text
"requests": 905, "seconds_measured": 20.101, "throughput_rps": 45.022,
"latency_ms": { "p50": 115.849, "p95": 224.135, "p99": 293.311 }
  1. Write the measurement record. design/measurements/01-baseline.json holds the figures you will quote, and it names the raw file they came from. This is the habit the whole course rests on: a number in a document that cannot be traced to raw data is an opinion with decimal places.
  1. Write `design/requirements.md` with the functional requirements, the objective table, the sizing estimate and an explicit out-of-scope list. Then write design/adr/0001-record-architecture-decisions.md, which commits the project to citing a measurement behind any performance claim.
  1. Commit. git init, then commit everything except .env, vendor/, logs/ and backups/.

Troubleshooting

The API container restarts in a loop and `docker compose logs api1` shows an ImportError → the vendor directory is empty or was not mounted. Re-run the pip install --target vendor command and check that compose.yaml mounts ./vendor at /vendor with PYTHONPATH=/vendor set.

`ops/migrate.sh` cannot connect → it reads .env for DB_USER and DB_NAME, and runs psql inside the db container. If .env does not exist yet, or was generated before compose.yaml, the variables are empty and psql fails with a confusing error about a database called "root".

Your p50 is suspiciously flat at around 40 ms regardless of load → the response leaves the handler as two writes, the headers and the body, and without TCP_NODELAY the kernel holds the second one waiting for an acknowledgement of the first. Setting disable_nagle_algorithm = True on the handler removes it. This is worth remembering: the first time you measure a service, some of what you measure is not the service.

Throughput is much lower than you expected and the load generator's own CPU is high → the generator is running on the same machine as the service and shares its CPU, and a Python generator at high concurrency becomes the bottleneck before the service does. The comparisons in this course are all between two runs of the same generator on the same machine, which is why that is survivable.

The seed takes a long time → it inserts 150,000 catalogue rows and 120,000 orders with generate_series. Around ten to twenty seconds is normal on this machine.

Check your understanding

Take the chapter 1 quiz when you have finished lesson 2. Before you move on, be able to answer these in a sentence each:

Why does the objective for shed load live separately from the error objective?

Because shedding is the system working. If a deliberate 503 with Retry-After counted against the error budget, the cheapest way to protect the budget would be to stop shedding and let everybody time out instead, which is worse for every customer.

Your summary says p95 is 224 ms. What must a reader be able to do with that?

Recompute it. The record names the raw file; the raw file has one row per request; the percentile method is stated. Anything less is a number they have to take on trust.

Summary and next step

  • Requirements are worth writing only if they can be tested: an indicator, a target and a

window, measured at a stated place.

  • The sizing estimate tells you which risks are real. For Northreach the volume is small and

the risks are latency and failure.

  • A baseline is the thing every later decision argues against, and it is only useful if the

raw results survive alongside the summary.

Next: the order API itself, where a retried request must not create a second order and a page of order history must not repeat or skip rows - and where an index is added because the measurement says so, not because indexes are generally good.

References

  • Python 3.12 documentation: http.server — https://docs.python.org/3.12/library/http.server.html (accessed 2026-09-20)
  • Python 3.12 documentation: http.client — https://docs.python.org/3.12/library/http.client.html (accessed 2026-09-20)
  • psycopg 3 documentation — https://www.psycopg.org/psycopg3/docs/ (accessed 2026-09-20)
  • Docker Compose specification — https://docs.docker.com/reference/compose-file/ (accessed 2026-09-20)
  • Google SRE Book, "Service Level Objectives" — https://sre.google/sre-book/service-level-objectives/ (accessed 2026-09-20)
  • Prometheus, exposition formats — https://prometheus.io/docs/instrumenting/exposition_formats/ (accessed 2026-09-20)
  • Architecture decision records — https://adr.github.io/ (accessed 2026-09-20)

Sign in to record your progress

Signing in saves eligible progress. It does not enroll you or include a lab; review the career path for access terms.

Sign in