Ai · Intermediate

Evaluating AI Applications

Build the evaluation programme that decides whether a change to an AI application ships: golden sets, graders, metrics, regression gates, cost and safety.

About this course

"It seems better" is not a result. This course builds the measurement programme that turns a question about an AI application — did this change make it better or worse? — into evidence a colleague can check and a manager can act on. You start from a retrieval question-answering application you inherit, with two deliberately degraded twins kept beside it, because a metric that gives a good build and a broken build the same number is measuring nothing. You then build the programme piece by piece: a versioned golden dataset with provenance, tags, behaviour classes and a hash-based dev/held-out split that cannot be gamed; deterministic graders for citation, schema, abstention, refusal, pattern and length, with unit tests that fail in both directions; a rubric, a judge stand-in behind the provider interface, and the agreement arithmetic — percent agreement and Cohen's kappa — that tells you whether the rubric works at all; retrieval metrics (recall@k, precision@k, MRR) measured separately from answer quality, with sweeps over k and chunk size; a variant comparison with bootstrap confidence intervals and a per-case flip table; a fail-closed CI gate with per-tag floors and an explicit known-gap policy; a SQLite results database, run diffs and a triage workflow; token, cost and latency reporting driven from configuration rather than literals; and a safety scorecard built from an adversarial corpus. Everything runs offline against a local deterministic stub, with no API key anywhere and no network call to any model provider, which is what makes every number in the course reproducible to the digit. The course is explicit about what such a stub can and cannot tell you, and what would change with a real model. This is a learning pathway toward an AI application developer role. It does not promise employment, seniority, salary or any vendor certification; completing it earns an Ultiblob Certificate of Completion.

Content time
17 h
Lessons
10
Certificate
Yes
on completion
Choose a career path

Lesson 1 is free. Enroll in a career path to access its full courses.

Lesson 1 is a free preview — read it without an account.

Ai — the kind of infrastructure this course is practised on

Outline

Lessons

10 lessons · 17 h
  1. Lesson 1: Evaluation as engineeringFree preview

    Build the system you will spend this course measuring, build two deliberately broken twins beside it, and discover that your first grader cannot tell them apart.

    1 h 30 min
  2. Lesson 2: Golden datasets

    Author a versioned case set whose expectations come from the documents, whose split comes from a hash, and whose validator you have watched reject a leak.

    1 h 40 min
  3. Lesson 3: Deterministic graders

    Build a registry of six graders that can tell a cited answer from an uncited one, and unit tests that go red when every grader is forced to pass.

    1 h 50 min
  4. Lesson 4: Rubrics, judges and agreement

    Write a rubric two people can apply the same way, implement a judge behind the model interface, and measure agreement with percent agreement and Cohen's kappa.

    1 h 50 min
  5. Lesson 5: Retrieval metrics

    Measure the search half of the pipeline on its own with recall@k, precision@k and MRR, and find out whether retrieval is the bottleneck before you tune it.

    1 h 40 min
  6. Lesson 6: Comparing variants

    Put two prompt templates head to head with per-tag pass rates, bootstrap confidence intervals and a per-case flip table — and be clear about what the interval covers.

    1 h 40 min
  7. Lesson 7: Evals as regression tests

    Turn the suite into a fail-closed gate with per-tag floors, a strict known-gap policy and prompt snapshots — then seed a fault and watch it go red.

    1 h 50 min
  8. Lesson 8: Results over time and triage

    Keep every run in a database with a stated grain, diff two of them, work a regression the gate never noticed, and add a case from a user report through a change record.

    1 h 50 min
  9. Lesson 9: Cost and latency

    Report tokens, estimated cost and latency percentiles per case, drive the prices from configuration, and measure what a cache actually saves.

    1 h 30 min
  10. Lesson 10: Safety and robustness evals

    Build an adversarial corpus, add a leak detector, prove the safety suite against a build with its guard removed, and publish a scorecard that reports the weaknesses too.

    1 h 40 min

Where it leads

Part of these career paths