Step 5: Upgrade and roll back by moving a tag¶
This page ships a configuration change to a running node without touching the node: you promote a new release in the registry, the node upgrades to it, and then you roll it back.
Make targets: make vm-upgrade and make vm-rollback, which run
deploy/vm-upgrade.sh and
deploy/vm-rollback.sh. make promote REL=4 moves the tag
on its own.
Promote and upgrade¶
Upgrading the fleet does not touch the nodes. The second site configuration release from
Step 3, edge-site-config release 4, has a new MOTD
and a new label on the demo Deployment, and it is already built and pushed as
rocky-edge:10.2-4. You promote it by copying that tag onto the floating
rocky-edge:10 inside the registry. The copy only touches the manifest and takes about a
second, because every blob is already in oci-bootc. Then the node runs bootc upgrade:
$ skopeo copy --src-tls-verify=false --dest-tls-verify=false \
docker://localhost:30080/oci-bootc/rocky-edge:10.2-4 \
docker://localhost:30080/oci-bootc/rocky-edge:10
[root@edge-node-01 ~]# bootc upgrade
layers already present: 73; layers needed: 3 (7.8 MB)
Deploying...done (2 seconds)
Queued for next boot: 10.0.2.2:30080/oci-bootc/rocky-edge:10
Version: 10.2-4
Digest: sha256:c98876108f864a75e7a7f958fda049d72673652e8eae5849c63dac893bb38f16
[root@edge-node-01 ~]# reboot
The node downloads 7.8 MB, stages the new deployment, and reboots into it. After the reboot,
bootc status shows 10.2-4 booted with 10.2-3 as the rollback, the MOTD reads
edge-site-config 1.0-4.el10: day-2 update via bootc upgrade (release 4), and RKE2 has
re-applied the re-seeded manifests, so the Deployment rolled to the release-4 label.
Roll back¶
bootc rollback and a reboot put everything back, manifests included, because the seeding
service runs from whichever image is booted. The proof of concept runs it in both directions
to be sure. The day-2 flow is diagrammed on the Architecture page.
Image-mode lesson: the missing bubblewrap¶
This step produced the failure we learned the most from. On an earlier build, the upgrade
staged cleanly, the node rebooted, and it came back on the old image. The staged deployment
was gone, and there was no error in the current journal. The clue was in the next boot's
ostree-boot-complete output:
ostree-finalize-staged.service failed on previous boot: Finalizing deployment:
Finalizing SELinux policy: Failed to execute child process "/usr/bin/bwrap"
(No such file or directory)
Here is what is going on. At shutdown, ostree finalizes the staged deployment. Because
rke2-selinux adds policy modules, /etc/selinux differs from the image's copy, so ostree
rebuilds the SELinux policy by running semodule inside bubblewrap (a small sandboxing tool).
The RESF minimal base does not include the bubblewrap package, and no package declares a
dependency on it, because libostree calls /usr/bin/bwrap by its hard-coded path.
Why nothing caught it: the base build never noticed, because the builder container had
bubblewrap as a dependency of rpm-ostree, and the community image happens to ship it. We proved
the diagnosis on the live node by temporarily overlaying /usr with bootc usr-overlay and
installing the package; the next upgrade went through. The fix is the bubblewrap package in
the first RUN of the edge Containerfile,
plus a tmpfiles.d entry for /var/log/journal so that the journal survives reboots and the next
person does not have to find this in a serial log. vm-upgrade.sh now prints the
ostree-boot-complete journal itself whenever the booted digest is not the staged one.
The lesson generalizes: neither this bug nor the read-only /opt in Step 3 showed up in
podman build, podman run or bootc container lint. If you adopt image mode, your CI
needs a stage that boots the image and upgrades it, not just builds it. The VM harness in this
walkthrough's code is that stage. See
Troubleshooting.