Ai · Advanced

Secure AI Deployment and Monitoring

Deploy a model-backed service the way it has to be run: threat-modelled, authenticated, rate-limited, logged, alerted, gated and rollable.

About this course

A retrieval question-answering application becomes a different kind of system the day it gets a URL, a customer and a login. This course takes one and deploys it properly, control by control, and proves every control against the running service rather than against a configuration file. You start by mapping the trust boundaries and writing a threat model that names a control for each of the ten OWASP risks for LLM applications. Then you build: FastAPI in a non-root, read-only container on a network with no route out, behind nginx terminating TLS; instruction and data separated in the template, with a request guard for direct injection and an ingest sanitiser for the instructions people write inside documents; an output filter that closes the exfiltration channels an answer can carry; per-client keys hashed at rest and bound to one workspace and one scope; a token bucket, a per-client daily token budget in shared state, a concurrency ceiling and size caps at two layers; one structured log line per request carrying digests instead of text; metrics scraped per instance, alert rules with an evaluator, and an eval canary that runs against the deployment; and a release pipeline that builds once, gates the artefact against a candidate stack, deploys the image id it gated and rolls back to one that passed. Two faults are seeded on purpose so you can see where each stage of a gate earns its place, and a third is seeded on a disposable copy of the stack so you can watch an alert fire and then put it away. You finish with a kill switch that works without a deploy, a runbook a stranger could follow, and a final project that onboards a third workspace, brings the whole stack up from nothing and hands it over. Everything runs offline against a local deterministic stub. No provider SDK is installed, no provider host is reachable from the application network, and at no point does the course ask you for an API key. The course is explicit about what that does and does not prove. This is a learning pathway toward an AI application developer role. It does not promise employment, seniority, salary or any vendor certification; completing it earns an Ultiblob Certificate of Completion.

Content time
21 h 20 min
Lessons
10
Certificate
Yes
on completion
Choose a career path

Lesson 1 is free. Enroll in a career path to access its full courses.

Lesson 1 is a free preview — read it without an account.

Ai — the kind of infrastructure this course is practised on

Outline

Lessons

10 lessons · 21 h 20 min
  1. Lesson 1: What changes when a model is in the loopFree preview

    Map the trust boundaries of a model-backed service, build the engine you will deploy, and write the threat model that names a control for every OWASP LLM risk.

    1 h 40 min
  2. Lesson 2: Serving the engine securely

    Put the engine behind FastAPI in a non-root, read-only container on a network with no way out, terminate TLS at nginx, and publish nothing else.

    2 h 10 min
  3. Lesson 3: Prompt injection, direct and indirect

    Build the two different controls that deal with text aimed at the assistant — a narrow request guard and an ingest sanitiser — and prove both against a corpus with a benign half.

    2 h 20 min
  4. Lesson 4: Exfiltration and tenant isolation

    Close the channels an answer can carry data out through, redact before anything is written down, keep one workspace out of another's answers, and write the retention policy.

    2 h
  5. Lesson 5: Clients, identity and secrets

    Give every client a key that is hashed at rest and bound to one workspace and one scope, and prove that no raw credential exists in the image, the environment, the logs or the history.

    2 h 10 min
  6. Lesson 6: Rate limits, budgets and caps

    Refuse too-fast clients with a Retry-After that is true, cap what each client may spend with a budget that survives a restart, and answer every refusal in a shape a client can act on.

    2 h 10 min
  7. Lesson 7: Logging for an AI service

    Write one structured line per request carrying digests instead of text, keep a correlation id from the caller to the log store, and review a synthetic log set that names two clients worth a conversation.

    2 h
  8. Lesson 8: Metrics, alerts and quality canaries

    Expose the series that change when behaviour changes, write alert rules with an evaluator, run the eval suite against the deployed service, and prove the abstention alert fires by seeding a fault.

    2 h 20 min
  9. Lesson 9: Safe releases and rollback

    Build once, record the image id, gate that exact artefact, deploy the id you gated, and roll back to one that passed — with two seeded faults showing where each stage of the gate earns its place.

    2 h 20 min
  10. Lesson 10: Incidents and the kill switch

    Build a switch you can pull before you understand the problem, block one client without touching the others, rotate an exposed key, and write the runbook and the post-incident review.

    2 h 10 min

Where it leads

Part of these career paths