Detection Engineering · Lesson 1 of 10

Free preview

From analyst to detection engineer

What detection engineering is, the lifecycle every detection moves through, and the telemetry you will build detections from.

time
1 h 15 min
on completion
+120 XP

Why this matters

In SOC Analyst Foundations you worked an incident: a brute force, a successful login, keys written, a cron job, a systemd unit, a beacon. You found all of it *after the fact*, from artefacts someone collected. The obvious question at the end of that case was why nobody was told at the time. The answer is almost always the same: no rule was watching the thing that mattered, or a rule was watching but nobody had ever tested whether it worked, so it drowned in false alarms and got muted.

Detection engineering is the discipline that fixes that. It is the difference between "I noticed this in the logs" and "there is a tested, reviewed, version-controlled rule that notices this every time, that we know the false-positive rate of, and that we can change safely." A detection engineer treats a detection the way a developer treats code: it has a source file, a test, a review, a coverage record and a change history. That is the whole course in one sentence, and this lesson sets up the two things everything else rests on — the lifecycle a detection moves through, and the telemetry you build detections from.

Nothing here touches anything but your own lab machine soc01. Every log you analyse is synthetic, generated on the box from addresses reserved for documentation. You never write a real rule against a real system in this course; you learn the craft on data that is safe by construction.

Concepts

The detection lifecycle. A detection is not a thing you write once. It moves through stages, and every artefact you produce in this course belongs to one of them:

text
requirement  ->  research   ->  build   ->  test        ->  deploy    ->  tune        ->  retire
(what must    (what data     (write     (positive and   (put it on   (measure and   (remove it when
 we catch      can see it,     the rule)  negative        the         adjust with     the threat or the
 and why)      what it looks             fixtures)        platform)    a rationale)    data source is gone)

Read it as a loop, not a line: tuning sends you back to build, a new requirement restarts it, and retirement is a deliberate step, not neglect. The failures that hurt a SOC live in the gaps between these stages — a rule built and deployed but never tested (unknown false-positive rate), or tuned but never documented (a silent blind spot the next analyst discovers during the next incident). Each chapter of this course is one part of the loop: you research and normalise data (this chapter), build and convert rules (chapter 2), test and tune (chapter 3), keep them as code with review and CI (chapter 4), and deploy and run them (chapter 5).

Telemetry is where a detection starts, not the rule. A rule that watches for a behaviour a source cannot see is dead on arrival. Before writing anything, you take inventory: what sources do I have, what can each prove, and what can it *not* prove? On soc01 you have four synthetic families, and they disagree with each other on purpose:

| Family | What it proves | What it cannot prove | Time format | |---|---|---|---| | SSH auth (auth.log) | login attempts, source, success/failure | what happened in the shell afterwards | syslog: local time, no year, no offset | | Web access (nginx) | requested paths, status, client, probing | the server-side effect of a request | clock time with an explicit -0500 offset | | Endpoint / EDR (JSON) | process exec, file writes, network connects | anything the sensor did not emit | ISO 8601 UTC (...Z) | | Windows security (samples) | the shape of event IDs 4624/4625/4688/7045 | Linux host behaviour | ISO 8601 UTC |

The single most important consequence: the same event can appear in more than one source with different field names and different clocks. A login failure is a line in auth.log *and* a JSON record in the EDR feed. If you count both, you double-count. Chapter 1's second lesson turns all of this into one stream with one clock and one set of field names so a rule written once is correct across sources — the first real job of the trade.

ATT&CK is the shared vocabulary — and it is versioned. MITRE ATT&CK gives every technique a stable id (T1098.004) and a name (SSH Authorized Keys) so a rule, a report and a coverage matrix can all refer to the same behaviour. You will tag every rule with an attack. tag, and a validator will check the tag is real. Two cautions that matter in practice, both true today:

  • A technique's tactic and even its id can change between ATT&CK versions. In ATT&CK v19 (April 2026)

the *Defense Evasion* tactic was split into Stealth and Defense Impairment, and T1070.002 Clear Linux or Mac System Logs was revoked and moved to T1685.006. A rule tagged with an old id will validate against an old data set and fail against a new one.

  • Your tools may carry different ATT&CK versions. This lab's Sigma validator uses ATT&CK v19.2; the Wazuh

manager ships an older bundled copy. You will see this disagreement in chapter 5, and it is a real detection-lifecycle fact, not a bug: keeping your platform's ATT&CK data current is maintenance work.

Always cite the technique you verified today at attack.mitre.org, never one you remember.

Guided exercise

Sign in to soc01 as the analyst account. Your whole workspace for this course is ~/det, already seeded with the tools you will use (run ls ~/det; read ~/det/README.md).

  1. Generate the working bundles. You build detections against data you control. Generate the three scenarios you will use throughout, at two different seeds, so no rule can accidentally depend on one seed's coincidences:
bash
cd ~/det
mkdir -p data
for seed in 4213 7788; do
  for s in baseline bruteforce persistence; do
    python3 generate_events.py --scenario "$s" --seed "$seed" --out "data/$s-$seed" --quiet --force
  done
done
ls data
text
baseline-4213  baseline-7788  bruteforce-4213  bruteforce-7788  persistence-4213  persistence-7788

Each bundle carries a manifest.json listing a sha256 for every file. Confirm one is intact — this is the same integrity habit ULC-006 taught, and a check grades it:

bash
cd ~/det/data/persistence-4213
jq -r '.files[] | "\(.sha256)  \(.path)"' manifest.json | sha256sum -c --quiet && echo BUNDLE_OK
text
BUNDLE_OK
  1. Write the telemetry inventory. Create ~/det/telemetry-inventory.md. For each of the four families, write one row: what it can prove, what it cannot, and its time semantics. This is not busywork — it is the document you will consult every time you ask "can a rule even see this behaviour?" A worked version is in the model solution, but write yours from what you observe in the bundle:
bash
head -3 data/persistence-4213/logs/auth.log
jq '.events[0]' data/persistence-4213/telemetry/edr.json
  1. Read one behaviour across two sources. Find the write to /root/.ssh/authorized_keys in the EDR feed, then confirm the SSH login that preceded it in auth.log. Notice they use different clocks and different field names for "who":
bash
jq -c '.events[] | select(.path=="/root/.ssh/authorized_keys")' data/persistence-4213/telemetry/edr.json
grep "Accepted password" data/persistence-4213/logs/auth.log

You have just done, by hand, the thing the next lesson automates: correlating one story across sources that do not agree on how to say anything.

Troubleshooting

Symptom → generate_events.py refuses with "refusing to overwrite non-empty". You already generated that bundle. Add --force, or choose a new --out directory. The generator never overwrites without being told.

Symptom → sha256sum -c prints FAILED for a file. You edited a bundle file, or a copy truncated. Never edit a bundle by hand; regenerate it. The manifest is the source of truth for what the data should be.

Symptom → two sources disagree on the time of "the same" event. That is by design — auth.log has no time zone and the nginx log carries -0500. Do not "fix" one to match the other; the next lesson converts both to UTC so the comparison is correct.

Symptom → you tag a rule with an ATT&CK id you remember and a later check rejects it. The id was renamed or revoked between versions. Look it up at attack.mitre.org today and use the current id.

Check your understanding

Why take a telemetry inventory before writing any rule?

Because a rule that watches for a behaviour no source records can never fire, and a rule that reads a source that cannot prove what it claims will mislead. The inventory tells you, up front, which detections are even possible with the data you have — and which require a new data source before they can exist.

You find a login failure in both auth.log and the EDR JSON. Is that two events?

No — it is one event observed by two sensors. If a detection counts both, its numbers are wrong. Normalising to one stream (next lesson) and knowing which source is authoritative for a given field is how you avoid double-counting.

A rule you wrote last year is tagged attack.t1070.002. Why might it fail validation now?

Because ATT&CK v19 revoked T1070.002 and moved it to T1685.006 when it split Defense Evasion into Stealth and Defense Impairment. The behaviour still exists; the identifier changed. Keeping tags current is part of maintaining detections.

Summary and next step

  • Detection engineering turns what an investigation learned into tested, versioned, maintainable detections

that move through a lifecycle: requirement, research, build, test, deploy, tune, retire.

  • A detection starts from telemetry: know what each source can and cannot prove before you write a rule.
  • ATT&CK is the shared vocabulary, and it is versioned — cite the id you verified today.

Next, lesson 2 turns the three disagreeing log families into one UTC event stream in a SQLite database, so a rule written once queries every source correctly.

References

  • Sigma detection format — https://github.com/SigmaHQ/sigma-specification (accessed 2026-09-19)
  • MITRE ATT&CK (Enterprise), technique and tactic reference, incl. the v19 Stealth / Defense Impairment split — https://attack.mitre.org/ (accessed 2026-09-19)
  • MITRE ATT&CK v19 release notes (Defense Evasion split; T1070.002 → T1685.006) — https://attack.mitre.org/resources/updates/updates-april-2026/ (accessed 2026-09-19)
  • RFC 5737 — IPv4 documentation address blocks — https://www.rfc-editor.org/rfc/rfc5737 (accessed 2026-09-19)

Sign in to record your progress

Signing in saves eligible progress. It does not enroll you or include a lab; review the career path for access terms.

Sign in