Self-hosting

Installing a system

How the development machine installs a new kernel next to the old one, boots it once, and keeps it only if you say so.

When you build SlopOS inside the development machine, you install the result next to the system you are running, and the running system stays untouched. The new system boots once. If it works, one command makes it the default. If it doesn't, the next boot brings back the old one. With the panic=reboot boot option that fallback is automatic: a kernel that crashes resets the machine, and the machine comes back on the old system.

The idea is the A/B update. Android and ChromeOS update this way: the device keeps two copies of the system, called slots A and B, runs from one, writes the update into the other, and can go back to the old slot if the new one fails. The "boot it once, then decide" step is the same as Linux's grub-reboot and systemd-boot's automatic boot assessment. SlopOS uses the boot loader variables that systemd defined for the latter.

Why a kernel needs a safe way to be replaced

Most programs can be replaced while the machine runs: if the new version is broken, you run the old one. A kernel can't be treated that way. If you overwrite the only kernel and the new one crashes during boot, there is no running system left to fix it from. The same goes for an update interrupted halfway: a machine that loses power while its kernel file is half written may not boot at all.

So an update needs three things:

  • The old system stays intact on disk until the new one has proven itself.
  • Trying the new system doesn't commit you to it.
  • Writing the new system can't leave a half-written file behind.

What a system is

On SlopOS a system is a kernel plus the base, a single archive holding the programs, the C library and the data they read (everything under /bin, /sbin, /lib, /usr/bin, /usr/share and /etc/ssl). The two are built and installed together. The kernel mounts its base read-only over the root disk, so the programs that boot are always the ones built with that kernel, and your files on the root disk are never part of an install. Filesystem explains how the base is mounted.

The boot disk

The firmware starts the machine from a small FAT32 partition called the EFI system partition (ESP). A computer that already runs another OS has one, and other boot loaders (GRUB, Windows) keep their files on it too. So SlopOS keeps its share of the ESP to one directory, and puts its systems on a partition of its own. A SlopOS disk has these partitions:

PartitionFormatWhat it holds
ESPFAT32The boot loader, Limine, and its menu, both in \EFI\SlopOS\
SlopOS bootFAT32, 1 GiBThe two slots: /boot/a/kernel.elf, /boot/a/base.img, /boot/b/kernel.elf, /boot/b/base.img
SlopOS rootext4/, unless the root is a disk of its own
SlopOS crashraw, 4 MiBReserved for the record of the last crash (not written yet)

The three SlopOS partitions are marked with partition type IDs that belong to SlopOS, not the standard Linux ones, so a Linux installed on the same disk never mistakes them for its own and mounts them. Limine reads the menu stored next to itself before any other, so it finds SlopOS's menu even if another copy of Limine shares the ESP. On an ESP that SlopOS created, it also puts a copy of Limine at the fallback path \EFI\BOOT\BOOTX64.EFI, which firmware boots when it has no other entry. On a shared ESP that path belongs to whoever was there first, and SlopOS leaves it alone.

The boot menu has one entry per slot, slopos-a and slopos-b. Each loads that slot's kernel and base from the boot partition and adds slot=a or slot=b to the kernel's command line. The menu doesn't say which slot is the default; a firmware variable does (see below), and with none set Limine boots the first entry, slot a. When the system is up, the boot log names the base it booted with, in a line starting BOOT: base.

On the development machine the root stays a separate disk image, because the host grows and refreshes it between sessions; the boot disk carries the ESP, the boot partition and the crash partition. Every just boot rebuilds the boot disk from the host's build, with the same system in both slots and slot a as the default. Reboots inside the guest keep it, so what you install survives until you quit QEMU.

Installing, trying and committing

In the guest, from the clone:

cd /src/slopos
scripts/selfhost.sh install

This builds the kernel and the base, then uses bootctl, the program that manages the boot partition:

bootctl spare                     # prints the slot that isn't the default, here b
bootctl install b builddir/kernel-release.elf builddir/initramfs.cpio
bootctl oneshot slopos-b

install copies the kernel and base into the spare slot. oneshot asks the boot loader to boot that slot on the next boot only. Now try it:

bootctl reboot

The machine boots the new system. Look around, run your tests. If you are happy, make it the default:

bootctl commit

commit asks the boot loader which entry it booted and makes that entry the default, by setting a firmware variable. It doesn't write the ESP at all. From now on slot b boots every time, and the next install goes into slot a.

bootctl status shows where things stand: the entry that booted, the default, any boot that is armed, and the size of each slot's kernel and base.

What happens if the new system fails

If you never run bootctl commit, the new system never becomes the default. The boot loader cleared the one-time request as soon as it acted on it, so the next boot, for whatever reason, is the old default.

That covers a system that boots but misbehaves: reboot and you are back. A kernel that crashes needs one more thing. By default a SlopOS kernel that panics prints its report and halts. With panic=reboot on the kernel command line it resets the machine instead, which lands on the old default with no one at the keyboard. The development machine's slots don't set that option, so on a plain just boot you restart QEMU yourself after a crash; Boot options describes panic=reboot.

Rollback doesn't undo anything outside the boot disk. Files the new system wrote to the root disk stay written. Since the root disk holds no system programs, that is usually what you want.

How the boot loader is told what to boot

The boot loader and the running system talk through UEFI variables: small named values the firmware keeps in non-volatile memory, which both the boot loader and the OS can read and write. systemd's Boot Loader Interface defines a set of them, and bootctl uses three that Limine implements:

VariableWritten byMeaning
LoaderEntryOneShotbootctl oneshotBoot this entry on the next boot only. The boot loader deletes it when it acts on it
LoaderEntrySelectedThe boot loaderThe entry that booted this time. bootctl commit reads it
LoaderEntryDefaultbootctl commit, bootctl set-defaultThe entry to boot when nothing else is asked for

The boot loader also reports which partition it was started from (LoaderDevicePartUUID), which is how bootctl finds the right disk and its boot partition. oneshot stores its variable as non-volatile, so it survives the reset.

Choosing the default through a variable rather than by editing the menu file means a commit never touches the ESP, so the menu stays byte for byte as it was installed. That matters later for Secure Boot, where the boot loader checks the menu against a hash recorded when it was signed.

These variables are shared by every boot loader on the machine that follows the same interface. If another loader left a default that names no SlopOS entry, bootctl install refuses to go on until bootctl set-default names one, because it can't tell which slot is safe to overwrite.

Writing the boot disk safely

bootctl never overwrites a file in place (the technique is called copy-on-write). It writes the new contents into free space, then switches the file's directory entry to point at them with a single small write, then frees the old contents. A crash before the switch leaves the old file, a crash after it leaves the new one, and either way the only damage is some unused space that fsck.fat reclaims.

It also refuses to install into the slot that is currently the default, so the system that boots by default is never the one being written. It writes the base before the kernel and reads each file back to compare it. And if a one-time boot was already armed for the slot, it clears that first, so a failed install can't leave a half-written slot armed to boot.

Who may change what boots

A program that can rewrite the boot partition decides what the machine runs next, so the kernel gives that ability to /bin/bootctl and not to programs in general. SlopOS grants privileges to programs by their path, and bootctl gets two (Permissions explains how that works):

  • Mount lets it write the raw boot partition. Ordinary programs can't write a raw disk at all, and nobody can while the partition is mounted.
  • Power lets it reboot and read and write the boot loader's variables. Even then, the kernel only allows the Boot Loader Interface's variables and SlopOS's own. Without that limit, a program allowed to reboot could also rewrite the firmware's settings or its Secure Boot keys.

A third privilege, BootEntry, reaches the firmware's own boot list (Boot####, BootOrder, BootNext, BootCurrent) so that SlopOS can add an entry for itself, and nothing else in the firmware's settings. The kernel checks every write against the format the firmware expects, because the firmware reads these variables at every power on, before anything could repair a bad one. It is meant for the installer; today only the install test holds it.

The grants are safe to give by path because /bin/bootctl comes from the read-only base, which no program can replace.

How it is tested

  • just test-install builds one disk in the full layout (ESP, boot, root and crash partitions) and, across the reboots of one QEMU, registers a firmware entry for SlopOS, copies slot a into b, boots b once, commits it, then boots once into a slot whose kernel deliberately panics under panic=reboot and checks that the reset lands on the committed default. Afterwards the host checks that the ESP is unchanged byte for byte and that the root partition passes e2fsck.
  • just test-install-guest does the same with a system the guest built itself: it runs scripts/selfhost.sh install tests as a person would, boots the result, checks that it is the build it just made, has it push a commit the host fetches, and compares the installed slot byte for byte with what the guest built. It needs a built toolchain, a clean working tree and DEV_QEMU_MEM=8G.

What isn't done

All of this runs in QEMU. The disk layout is the one meant for real computers, but installing SlopOS on one needs more:

  • An installer. Today the host builds every disk, and the guest has no tools to format or partition one, nor can its FAT writer create the long file names under \EFI\SlopOS\.
  • A record of a crash that survives the reset. The crash partition exists, but nothing writes to it yet, so on a machine with no serial console panic=reboot erases the only report.
  • Secure Boot signing. The layout allows it; the signing isn't planned.

See Known limitations and Roadmap.

For contributors

bootctl writes FAT32 through fat-core, a no_std crate with no unsafe code. A replaced file's new clusters and their FAT chain are written first; the switch is one 32-byte directory entry, inside one sector, so the disk can't tear it. Keep that order if you change the writer.

User space reaches UEFI variables through two system calls, efivar_get and efivar_set. Both need the Power privilege and reach only the Boot Loader Interface's GUID (4a67b082-0a4c-41cf-b6c7-440b29bb8c4f) and SlopOS's (5a1b0b05-5105-4e57-a11e-0000000000a1); with BootEntry (TASK_FLAG_INSTALL) they also reach Boot####, BootOrder, BootNext and BootCurrent under the EFI global GUID, each write held to non-volatile, boot-and-runtime access and the variable's format. Firmware runtime services are mapped only in the kernel's own address space and are not reentrant, so every call runs on one dedicated kernel thread, one request at a time, while the calling task sleeps. On a BIOS boot both calls return ENODEV.

boot-core describes the disk once, as data: the partition type GUIDs and sizes, the GPT reader the kernel's partition probe uses, load options and device paths, the boot variables' formats and the Limine menu. The kernel, bootctl and the host's disk builder (through tools/bootdisk) all read it from there, so change the layout in boot-core and nowhere else.

Further reading

  • A/B system updates, Android's documentation. The clearest description of two slots and falling back to the old one, at the scale of a phone.
  • Automatic Boot Assessment, systemd. A more general version of "try once, then commit", which counts boot attempts instead of allowing exactly one.
  • The Boot Loader Interface, systemd. The specification of LoaderEntryOneShot, LoaderEntrySelected and the other variables.
  • Disk format, ChromeOS. How ChromeOS lays out two kernel and root partitions and marks one as successfully booted.

In the source

WhereWhat
userland/src/apps/bootctl.rsbootctl: slots, one-time boots, commit
userland/src/boot_disk.rsFinding the boot disk; registering the firmware entry
boot-core/The disk layout, GPT, boot variables and Limine menu, host-tested
scripts/selfhost.shThe guest's build and install command
scripts/build_bootdisk.sh, tools/bootdisk/Building the boot disk on the host
core/src/efivar.rsThe kernel side of the variable system calls and what each may reach
userland/src/bin/tests/install_test.rsThe install, commit and rollback test

On this page