Why this matters
A security operations centre is not a room full of screens. It is a queue, a set of decisions, and a record of who decided what and why. On a normal shift nobody is chasing an intruder through the network. Somebody is looking at the ninth "multiple failed logons" alert of the morning and deciding whether it is a locked-out warehouse account, a misconfigured backup job, or the first three minutes of something that will matter tomorrow.
That decision is the job. Tools change every few years; the decision does not. An analyst who can say "here is what I observed, here is what it supports, here is what it does not support, and here is what I need next" is useful on any platform. An analyst who can only read one dashboard is useful until that dashboard is replaced.
This course teaches the decision, using data you generate yourself. Everything you investigate is synthetic: produced on your lab machine soc01 from addresses reserved for documentation. You will never touch a production log, a customer system, or anything outside your own lab. That restriction is not a limitation of the course — it is how a beginner should be introduced to security work at all.
Concepts
A signal is not an alert, an alert is not a case, and a case is not an incident. These words get used loosely; keeping them separate keeps you honest.
- A signal is a raw observation: a log line, a telemetry event, a phone call from a user.
- An alert is a signal that a rule decided is worth a human's attention.
- A case (or ticket) is the record you open when you start working an alert. It holds the
evidence, the reasoning and the outcome.
- An incident is a determination — a statement that something bad actually happened. It is made
from evidence, and declaring one usually triggers obligations: notification, containment authority, sometimes legal or regulatory clocks.
Most alerts never become incidents. That is not failure; it is the system working. The failure mode to avoid is the reverse: quietly closing something that was an incident, or declaring an incident because a word in a log line looked frightening.
Tiers are a division of labour, not a ranking of people. In a tiered SOC, tier 1 triages the queue: confirm the alert is real, gather the obvious evidence, close what is explainable, escalate what is not. Tier 2 investigates escalations in depth. Tier 3 (often incident response or detection engineering) handles confirmed incidents and improves the rules so the same alert does not consume tier 1 again next week. Small teams collapse all three into one person, and that person still moves through the same three modes of work.
Severity and confidence are different numbers and must be recorded separately. Severity estimates how much it would matter *if true*. Confidence estimates how strongly the evidence supports the claim. "A domain administrator account was used from an unknown address" is high severity and, at first, low confidence. Collapsing the two into one field is how a SOC ends up either crying wolf or sitting on something serious. Write both down, and write down what would raise or lower each.
The workflow is a loop, not a line. Alert arrives → you *authorise and scope* (what am I allowed to touch, what is in scope) → you *classify* (what could explain this) → you *collect* evidence → you *analyse* → you *decide*: close, escalate, or contain → you *record* → the outcome feeds back into tuning so the queue gets better. This course walks the loop once, in order, over ten lessons.
Shift handover is part of the work, not an afterthought. A case that changes hands loses everything that lived only in the previous analyst's head. A usable handover says: what is open, what was already checked, what the current best explanation is, what the next action is, and what question you could not answer. Three lines of that beat a page of narrative.
What "good" looks like. Good tier-1 work is boring, fast and reproducible. Someone else can read your case note and reach the same conclusion from the same evidence. You state what you checked *and* what you could not check. You never write "no malicious activity found" when what you mean is "the two sources I can see show nothing". Absence of evidence is a statement about your visibility.
Guided exercise
Your lab machine is soc01, an Ubuntu 24.04 server. You sign in as the analyst account and you have sudo. Three services matter for this course: auditd records kernel-level audit events, fail2ban watches authentication logs and bans repeat offenders, and nginx serves a small website so there is a web log to read.
- Open the lab console for `soc01` and sign in as `analyst`. The lab platform confirms the machine is up and the three services are running before it lets a check pass.
- Confirm who you are and what time the machine thinks it is. Every observation you record needs a time you can defend.
id
date -u +"%Y-%m-%dT%H:%M:%SZ"
date +"%Z %z"Observed on the drafting host:
uid=1001(analyst) gid=1001(analyst) groups=1001(analyst)
2026-09-13T05:18:58Z
UTC +0000Your uid number and the current time will differ; the shape will not. Write the UTC time down. From here on, every timestamp you record in a case note is UTC, whatever the log file used.
- Create your working directories. Keep the evidence and your own work separate — you never edit evidence.
mkdir -p ~/soc/case
ls ~/soc- Generate your first bundle. The
baselinescenario is a quiet day with nothing hostile in it. Reading a normal day first is what makes an abnormal one visible later.
python3 ~/soc/generate_events.py --scenario baseline --seed 4213 --out ~/soc/baselineObserved output (abridged — the tool prints one line per file):
scenario : baseline seed 4213 host soc01
window : starts 2026-09-08T02:55:00Z
output : /home/analyst/soc/baseline
5 lines 526 bytes logs/auth.log
211 lines 40576 bytes logs/nginx/access.log
10 lines 921 bytes snapshot/ps.txt
manifest.json lists a sha256 for each file above- Read the quiet day. Five lines of authentication log is the whole day's story.
cat ~/soc/baseline/logs/auth.logSep 8 03:01:00 soc01 sshd[5895]: Failed password for ubuntu from 192.0.2.10 port 37235 ssh2
Sep 8 03:01:09 soc01 sshd[5895]: Failed password for ubuntu from 192.0.2.10 port 44960 ssh2
Sep 8 03:01:31 soc01 sshd[2749]: Accepted publickey for ubuntu from 192.0.2.10 port 60020 ssh2: ED25519 SHA256:brWSVgN7yCkTtwRgjPhJpozzZvCrYExtoKZvjx6AXHc
Sep 8 03:01:31 soc01 sshd[2749]: pam_unix(sshd:session): session opened for user ubuntu(uid=1000) by (uid=0)
Sep 8 03:01:32 soc01 systemd-logind[721]: New session 14 of user ubuntu.Two failures then a key-based success from the same address, thirty-one seconds apart. Most SOCs would not alert on this at all. Notice what it teaches you: on this host, failures happen, and two of them from a known workstation are normal.
- Write the case note. In
~/soc/case/notes.md, record: the UTC time you started, the machine, the account you are using, what you generated, and one sentence saying what a normal day looks like on this host. You will extend this file in every lesson.
Troubleshooting
Symptom → python3: can't open file '/home/analyst/soc/generate_events.py'. The generator lives in the analyst home directory as ~/soc/generate_events.py. If you are in a sudo -i root shell, ~ means /root, not your home. Exit back to your own shell and try again.
Symptom → refusing to overwrite non-empty /home/analyst/soc/baseline (use --force). The generator will not silently destroy a bundle you may have annotated. Either choose a new --out directory or add --force when you genuinely want to regenerate.
Symptom → the check for a service fails even though the service looks fine to you. The platform evaluates service state from outside the guest; it does not read a file you can edit. If it reports auditd inactive, the service really is not running — restart it with sudo systemctl restart auditd and re-run the check.
Symptom → your timestamps do not match the ones printed in this course. They should not. The lesson output above was captured on the authoring host on 2026-09-13. Your date output is your own; the *bundle* timestamps, however, are deterministic and will match exactly for the same seed.
Check your understanding
Take the chapter quiz after lesson 2. Before you do, answer these two for yourself.
An alert says "brute force detected". Your manager asks whether you have an incident. What do you say?
That you have an alert and an open case, and that "incident" is a determination you have not made yet. Then say what would settle it: evidence of a successful authentication from the same source, and any activity after it. Give a time by which you will know.
Why record severity and confidence separately when the queue only shows one colour?
Because they move independently. Evidence changes confidence, not severity. If you merge them, a high-severity possibility with weak evidence either gets over-escalated on day one or gets under-recorded and forgotten — and the case note no longer shows which of the two happened.
Summary and next step
- A SOC is a queue of decisions; the artefact you produce is a defensible record, not a verdict.
- Signal, alert, case and incident are four different things, and only evidence moves you along that
chain.
- Severity is "how bad if true", confidence is "how sure am I" — record both, separately.
Next, lesson 2 inventories the log sources you will be reading and states, for each one, what it can prove and what it cannot.
References
- MITRE ATT&CK — https://attack.mitre.org/ (accessed 2026-09-13)
- Ubuntu Server documentation — https://documentation.ubuntu.com/server/ (accessed 2026-09-13)
- RFC 5737, IPv4 address blocks reserved for documentation — https://www.rfc-editor.org/rfc/rfc5737 (accessed 2026-09-13)
Sign in to record your progress
Signing in saves eligible progress. It does not enroll you or include a lab; review the career path for access terms.