Infrastructure Automation with Ansible · Lesson 1 of 10

Free preview

Why automation, and how Ansible reaches a node

What Ansible actually does when it configures a server, how it gets in, and why the account it uses is a security decision you make on day one.

time
1 h
on completion
+120 XP

Why this matters

One server configured by hand is a craft. Ten servers configured by hand are ten slightly different servers, and the differences surface at the worst moment: the one node that still allows password logins, the one that missed a patch, the one somebody edited at 2 a.m. and never wrote down.

Automation is not about typing less. It is about being able to say, with evidence, what state your servers are in — and to put them back in that state on demand. Ansible does that by letting you *declare* that state in files you keep in version control, then converging every machine on it.

You work from ctl01, your control node, against two managed machines, node01 and node02. Nothing is installed on those nodes to make this work, and that shapes everything else in this lesson.

Concepts

Agentless: what actually happens during a task

Ansible has no daemon on the managed node. For each task the control node opens an SSH connection as a specific account, sends the module (a small Python program) and its arguments, runs it with the node's own python3, reads back one JSON result — what it found and whether it changed anything — and leaves nothing behind.

So a managed node needs only SSH, Python 3 and a way to become root, and there is nothing to patch or attack between runs. The price is an SSH connection per run.

Control node, managed node, and the two packages

The control node is where Ansible runs; the managed nodes are what it configures. Two packages matter: ansible-core is the engine (the ansible, ansible-playbook, ansible-galaxy and ansible-vault commands plus the ansible.builtin modules), and ansible is the community package that bundles it with several hundred collections — modules grouped by namespace, such as ansible.posix and community.postgresql. Lesson 5 shows how to pin the collections a repository depends on instead of relying on whatever is installed.

Inventory: the list of what you manage

An inventory names your hosts and arranges them in groups, so a playbook can say "web servers" instead of naming machines; groups are also where most variables live. A host can be in many groups, and all contains every host. Reading back the inventory Ansible actually parsed, with ansible-inventory, is the cheapest bug-avoidance habit in this course.

Host keys: trust the node before you configure it

The first time you connect to a machine over SSH it offers its host key. Accepting whatever arrives means accepting whatever answered, which is not the same as the machine you meant. Ansible refuses to connect to an unknown host, and that refusal is a feature: confirm the fingerprint out of band (here, at the node's console), record it in ~/.ssh/known_hosts, and a changed key becomes an alarm rather than a shrug. Turning host key checking off is a decision to trust the network.

The account Ansible uses is a security decision

Your pod was built with a dedicated service account on each managed node, called automation:

  • it is a system account with no password — nobody logs into it interactively;
  • it accepts one SSH key, and only from the control node's address (from="..." in authorized_keys);
  • it has a sudo rule of its own in /etc/sudoers.d/90-automation, and every elevation it performs is written to the node's journal.

Two honest caveats. That sudo rule lets automation become root and cannot be narrowed to a command list: become runs module code through /bin/sh, so a list either blocks everything or allows everything. The control is the *identity* — one account, one key, one source address, no password, full audit trail. And the key the platform issued is a bootstrap credential; a credential issued by a provisioning system should be replaced by one you control, which you will do at the end of this lesson.

Guided exercise

Work on ctl01 as ubuntu. Open the console for node01 and node02 in separate tabs — you need them once, for the host keys.

  1. Confirm the toolchain and see what the community package brings with it:
bash
ansible --version | head -n 1
dpkg-query -W -f='${Package} ${Version}\n' ansible ansible-core ansible-lint yamllint
ansible-galaxy collection list 2>/dev/null | grep -E "^(ansible.posix|community.general|community.postgresql) "
text
ansible [core 2.16.3]
ansible 9.2.0+dfsg-0ubuntu5
ansible-core 2.16.3-0ubuntu2
ansible-lint 6.17.2-1
yamllint 1.33.0-1
ansible.posix                            1.5.4
community.general                        8.3.0
community.postgresql                     3.3.0
  1. Look at what the platform gave you, and at nothing else:
bash
ls -l ~/.ssh
ssh-keygen -lf ~/.ssh/ub_bootstrap_ed25519.pub
text
-rw------- 1 ubuntu ubuntu 419 Sep 20 01:22 ub_bootstrap_ed25519
-rw-r--r-- 1 ubuntu ubuntu 108 Sep 20 01:22 ub_bootstrap_ed25519.pub
256 SHA256:f/G7rqUfkos7tgWdyhfYgv+b4+9hbw03jSJ0ulGChtg automation-bootstrap@ctl01 (ED25519)
  1. Create the repository that will hold everything you build in this course, and write your first inventory:
bash
mkdir -p ~/infra && cd ~/infra
git config --global user.name "Your Name"
git config --global user.email "you@lab.example"
git config --global init.defaultBranch main
git init

inventory.yml:

yaml
---
all:
  children:
    web:
      hosts:
        node01:
        node02:
    db:
      hosts:
        node02:

ansible.cfg — settings for this repository, read automatically when you run Ansible from inside it:

ini
[defaults]
inventory = inventory.yml
remote_user = automation
private_key_file = ~/.ssh/ub_bootstrap_ed25519
interpreter_python = /usr/bin/python3

[ssh_connection]
pipelining = True

pipelining = True sends the module over the open SSH session instead of copying a temporary file first: fewer round trips, and no temporary files on the node.

  1. Read back what Ansible parsed, not what you think you wrote:
bash
ansible-inventory --graph
text
@all:
  |--@ungrouped:
  |--@web:
  |  |--node01
  |  |--node02
  |--@db:
  |  |--node02
  1. Try to reach the nodes before you know their host keys, and read the error:
bash
ansible all -m ansible.builtin.ping
text
node02 | UNREACHABLE! => {
    "changed": false,
    "msg": "Data could not be sent to remote host \"node02\". Make sure this host can be reached over ssh: Host key verification failed.\r\n",
    "unreachable": true
}
  1. Verify the fingerprints and record them. On each node's console run sudo ssh-keygen -lf /etc/ssh/ssh_host_ed25519_key.pub, then on ctl01 compare:
bash
ssh-keyscan -t ed25519 node01 node02 2>/dev/null | ssh-keygen -lf -
text
256 SHA256:w/yk38vvOXePx3K4RWq53mYhwm8M9xhbtef5WFbrOSk node02 (ED25519)
256 SHA256:QCmfJC/jwOtFFioT69CK+dWkkor2qmpOUE4SFAMFTrQ node01 (ED25519)

The two fingerprints must match what the consoles showed. Only then record them:

bash
ssh-keyscan -t ed25519 node01 node02 2>/dev/null >> ~/.ssh/known_hosts
ansible all -m ansible.builtin.ping
text
node01 | SUCCESS => {
    "changed": false,
    "ping": "pong"
}
  1. See the agentless model and the privileges you are using:
bash
ansible node01 -m ansible.builtin.ping -vvv 2>&1 | grep -E "ESTABLISH SSH|SSH: EXEC ssh" | cut -d\> -f2- | head -n 2 | cut -c1-110
ansible node01 -a "id"
ansible node01 -a "sudo -n -l"
ansible node01 -b -a "journalctl -t sudo -n 2 --no-pager -o cat"
ansible node01 -a "dpkg-query -W ansible-core"
text
 ESTABLISH SSH CONNECTION FOR USER: automation
 SSH: EXEC ssh -C -o ControlMaster=auto -o ControlPersist=60s -o 'IdentityFile="/home/ubuntu/.ssh/ub_bootstrap
node01 | CHANGED | rc=0 >>
uid=107(automation) gid=108(automation) groups=108(automation)
node01 | CHANGED | rc=0 >>
...
User automation may run the following commands on node01:
    (root, postgres) NOPASSWD: ALL
node01 | CHANGED | rc=0 >>
automation : PWD=/home/automation ; USER=root ; COMMAND=/bin/sh -c 'echo BECOME-SUCCESS-jotzxvjdgkdzuefhhopecetrozqtvzdc ; /usr/bin/python3'
pam_unix(sudo:session): session opened for user root(uid=0) by automation(uid=107)
node01 | FAILED | rc=1 >>
dpkg-query: no packages found matching ansible-corenon-zero return code

The final line runs two streams together — dpkg-query's message on stdout has no trailing newline, so Ansible's non-zero return code from stderr continues it. Reading output exactly as the tool printed it is part of the job. The last two results are the lesson in one screen. The sudo line shows *why* a command list cannot scope Ansible: every task arrives as /bin/sh -c ... python3. And node01 has no Ansible installed at all — agentless is not a slogan.

  1. Look at how the key is restricted, then replace it with one of your own:
bash
ansible node01 -b -a "cat /home/automation/.ssh/authorized_keys"
ssh-keygen -t ed25519 -N "" -C automation@ctl01 -f ~/.ssh/automation_ed25519
text
from="10.254.149.2",no-agent-forwarding,no-port-forwarding,no-X11-forwarding ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAINWRJrESipuWo8PiORZGPbYd5NCbc5eVUSRGiNVFNEGh automation-bootstrap@ctl01

Your lab's addresses will differ. The new key has no passphrase on purpose: scheduled automation (lesson 9) cannot answer a prompt. It is protected by file permissions, by the from= restriction, and by being replaceable in a minute.

playbooks/rotate-automation-key.yml:

yaml
---
- name: Rotate the automation account's SSH key
  hosts: all
  gather_facts: true
  vars:
    new_public_key: "{{ lookup('ansible.builtin.file', '~/.ssh/automation_ed25519.pub') }}"
    control_node_address: "{{ ansible_facts['env']['SSH_CLIENT'].split() | first }}"
  tasks:
    - name: Authorise the new key, usable only from the control node
      ansible.posix.authorized_key:
        user: automation
        key: "{{ new_public_key }}"
        key_options: 'from="{{ control_node_address }}",no-agent-forwarding,no-port-forwarding,no-X11-forwarding'
        exclusive: "{{ revoke_old_keys | default(false) | bool }}"

ansible_facts['env']['SSH_CLIENT'] is the address the *node* sees you connecting from — exactly what from= must contain. The play does not use become: the account owns its own authorized_keys, so root is not needed.

Rotate in the order every credential change follows — add, verify, switch, revoke, prove:

bash
ansible-playbook playbooks/rotate-automation-key.yml --check --diff     # what would change
ansible-playbook playbooks/rotate-automation-key.yml                    # add the new key
ansible all -m ansible.builtin.ping --private-key ~/.ssh/automation_ed25519   # verify it works
sed -i "s#ub_bootstrap_ed25519#automation_ed25519#" ansible.cfg         # switch to it
ansible-playbook playbooks/rotate-automation-key.yml -e revoke_old_keys=true  # remove every other key
ssh -i ~/.ssh/ub_bootstrap_ed25519 -o IdentitiesOnly=yes -o BatchMode=yes automation@node01 true
text
automation@node01: Permission denied (publickey,password).

That refusal is the proof. Now delete the revoked key, confirm normal operation, and run the playbook once more to see idempotence:

bash
rm ~/.ssh/ub_bootstrap_ed25519 ~/.ssh/ub_bootstrap_ed25519.pub
ansible all -m ansible.builtin.ping -o
ansible-playbook playbooks/rotate-automation-key.yml
text
node01 | SUCCESS => {"changed": false,"ping": "pong"}
node02 | SUCCESS => {"changed": false,"ping": "pong"}
node01                     : ok=2    changed=0    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0

changed=0 on the second run is the first idempotence you will see in this course. The task did not "do nothing"; it checked reality and found it already correct.

  1. Commit, and save the evidence into ~/notes/lesson01.md: the two host-key fingerprints, the ping result and the refusal of the old key.
bash
git add -A && git commit -m "Inventory, Ansible configuration and key rotation playbook"

Troubleshooting

`Host key verification failed` → the node's key is not in ~/.ssh/known_hosts. Verify the fingerprint at the console and add it with ssh-keyscan. Never "fix" this by disabling host key checking: you would be trusting whatever answers to the name.

`Permission denied (publickey)` → the key you offered is not authorised for that account, or the from= option does not allow your address. ansible all -m ping -vvv shows which key file was offered.

`[Errno 2] No such file or directory: b'command'` → you passed a shell builtin (command -v, cd, export) to the command module, which does not use a shell. Use a real executable, or ansible.builtin.shell when you genuinely need shell features.

Every ad-hoc command reports `CHANGED` → the command module cannot know whether it changed anything, so it assumes it did; lesson 4 shows changed_when.

`ansible` works from your home directory but uses the wrong settings → ansible.cfg is read from the current directory. Run Ansible from inside the repository, or set ANSIBLE_CONFIG.

Check your understanding

Question 1: Your colleague suggests putting host_key_checking = False in ansible.cfg because a rebuilt node keeps failing. What do you say?

That the failure is doing its job: the node's identity changed, and Ansible noticed. The fix is to verify the new host key out of band (the node's console) and update known_hosts — for one rebuilt node, ssh-keygen -R node02 then ssh-keyscan. Disabling the check hides every future impersonation as well as this one.

Question 2: Why can the automation account's sudo rule not be narrowed to "only run nginx and systemctl"?

Because Ansible does not run those commands. It runs a Python module through /bin/sh -c, which a command list must either allow (and then everything is allowed) or block (and then nothing works). The scope comes from the account itself: no password, one key, accepted only from the control node, and every elevation logged.

Summary and next step

  • Ansible is agentless: it opens SSH, runs a Python module with the node's own interpreter, reads one JSON result and leaves nothing behind.
  • The inventory is data. Read it back with ansible-inventory before you run anything against it.
  • Verify a node's host key before you trust it, and treat the credential the platform issued as a bootstrap credential: add your own key, verify it, switch, revoke, and prove the old one fails.

Next: playbooks — writing the state you want down, applying it, and proving that a second run changes nothing.

References

  • Ansible documentation — Getting started: how Ansible works — https://docs.ansible.com/ansible/latest/getting_started/index.html (ansible-core 2.16 documentation)
  • Ansible documentation — How to build your inventory — https://docs.ansible.com/ansible/latest/inventory_guide/intro_inventory.html (ansible-core 2.16 documentation)
  • Ansible documentation — Connection methods and details (host key checking, SSH plugins) — https://docs.ansible.com/ansible/latest/inventory_guide/connection_details.html (ansible-core 2.16 documentation)
  • Ansible documentation — ansible.posix.authorized_key module — https://docs.ansible.com/ansible/latest/collections/ansible/posix/authorized_key_module.html (collection documentation)
  • OpenSSH manual — sshd(8), AUTHORIZED_KEYS FILE FORMAT (from=, no-port-forwarding) — https://man.openbsd.org/sshd (accessed 2026-09-19)
  • Ubuntu Server documentation — OpenSSH server — https://documentation.ubuntu.com/server/how-to/security/openssh-server/ (accessed 2026-09-19)

Ansible's own documentation site answered HTTP 429 to every request from the authoring workstation on 2026-09-19, so the pages above were not retrieved that day; their paths were confirmed in the stable-2.16 documentation source, and every behaviour cited from them was executed in this lab.

Sign in to record your progress

Signing in saves eligible progress. It does not enroll you or include a lab; review the career path for access terms.

Sign in