Architecture

Overview

What the parts of SlopOS are, what each one is responsible for, and how they work together when a program reads a file.

SlopOS gives you a working Unix-like system on a 64-bit x86 PC or in QEMU: a shell, the core utilities, an editor, network tools and a desktop, all running as ordinary programs that can't read each other's memory or take down the machine. If you have written programs for Linux, the interface will be familiar, because SlopOS uses Linux's system call numbers, signals and error codes. What is unusual is inside the kernel, which is written in Rust: all of its unsafe code sits in one small trusted core, and the compiler checks the rest. This page names each part of the system, says what it is responsible for, and then follows one ordinary action, a program reading a file, down through all of them and back.

Almost nothing here was invented for SlopOS. The way the kernel is split, with one small trusted part underneath everything else, comes from the Asterinas project, which calls the design a framekernel. The interface programs use is Linux's, down to the numbering.

The parts of SlopOS

The shell, the terminal, the desktop, ls and cat are all ordinary user programs, under the same rules as the programs you write. The few that need more, such as the compositor taking over the screen or bootctl writing the boot slots, are granted a specific permission by the kernel when they start (Permissions). The kernel itself is split in two: a small trusted core at the bottom, and everything else on top of it.

Programs ask the kernel for everything through system calls. Inside the kernel, only the trusted core may use unsafe Rust; drivers reach their devices through its safe interfaces.

The trusted core

At the bottom of the kernel sits a small library that does the things the Rust compiler can't check: setting up the CPU, building the tables that decide what memory each program can see, switching between programs, copying bytes to and from a program's memory, and receiving interrupts from devices. All of these need unsafe Rust, which is Rust's way of saying "the compiler can't prove this is correct; a person has to". SlopOS puts every line of the kernel's own unsafe code in this one library, called OSTD (the trusted core), and the build refuses unsafe anywhere else in the kernel.

That changes what a bug can do. A mistake in the rest of the kernel can still return a wrong answer, crash a program or even allow something it shouldn't, but the compiler rules out memory corruption there: reading freed memory, writing past the end of a buffer, the kind of bug that lets one program read another's data. That kind of bug can only start in the trusted core, so that is where the proofs, the reviews and the extra checking go. The framekernel explains the idea and where it comes from, and The trusted core describes what OSTD offers the rest of the kernel.

The kernel

On top of the trusted core, the rest of the kernel is written in safe Rust. It is split by responsibility:

  • Processes and scheduling. Processes, threads, signals and job control behave as on Linux. The scheduler preempts on a timer and keeps one run queue per CPU, and almost every wait in the kernel ends at once when the task is killed. See Processes and signals, Scheduling and waiting and Task lifetimes.
  • Memory. Each process gets its own address space, filled in on first use, and fork shares pages copy-on-write. Unlike Linux, the kernel refuses memory when a program asks for it rather than when it first touches it, so running out is usually an error the program can handle instead of a kill. See Memory and Resource limits.
  • Permissions. Some system calls, such as rebooting the machine or taking over the screen, are only for programs that have been given the right to make them. See Permissions.
  • Filesystem. Disks, the read-only base image and the special files under /dev form one directory tree. Disks use ext4 with a journal, the same format Linux uses. See Filesystem, Disks and Crash recovery.
  • Network stack. TCP, UDP, IPv4, DHCP and DNS run inside the kernel, as they do on Linux. See Networking.
  • Drivers. NVMe and virtio disks, virtio network and graphics, the Intel Xe display, the PS/2 keyboard and an I²C touchpad, each matched to its device in one declarative step. See Drivers.

Boot describes how all of this is started, in order, when the machine powers on.

Userland

Above the system call line is userland:

  • The C library, slibc, is written in Rust. Every program links against it, Rust programs included through the standard library, and it doubles as the dynamic loader.
  • SlopRing is a queue of requests a program shares with the kernel, so it can have many operations in flight without one system call each. A small async runtime for Rust programs is built on it. See SlopRing.
  • The programs: init, which the kernel starts first and which starts the desktop; the shell; the core utilities; an editor; and network tools such as curl, ping and nc.
  • The desktop: a compositor, which owns the screen and combines the other programs' windows into one picture, and the apps that draw into those windows, the terminal among them.

Userland covers the C library and the programs, and Desktop and windowing covers the compositor and how windows reach the screen.

Following a read through the system

Here is what happens when you type cat notes.txt in the shell and cat reads the file from the disk. Every layer from the diagram takes part.

  1. The program asks. cat has already opened the file and holds a file descriptor for it. It calls read() in the C library with a buffer in its own memory.
  2. The C library makes the system call. It puts the number of the read call (0, the same as on Linux), the file descriptor, the buffer's address and its size into CPU registers, and executes the syscall instruction. The CPU switches into kernel mode and jumps to the entry point the trusted core registered at boot. The trusted core saves cat's registers so it can resume it later.
  3. The kernel finds the handler. It looks up call 0 in its table of system calls and finds the read handler. Reading needs no special permission: holding the file descriptor is the permission, because cat could only have got one by opening the file itself or by being handed it by another program.
  4. The filesystem works out where the bytes are. The handler looks up the descriptor in the process's table of open files and asks the filesystem for the bytes. The ext4 code works out which blocks of the disk hold that part of the file. If those blocks are already in the kernel's block cache, the bytes are in memory and the next two steps are skipped.
  5. The driver asks the disk. Otherwise the block layer hands a read request to the disk's driver, which tells the disk (an NVMe drive under QEMU and on most laptops) where in memory to put the data. Unless the answer comes back almost at once, cat sleeps and the scheduler runs something else on that CPU.
  6. The disk interrupts. When the data has arrived, the disk raises an interrupt. The driver's handler marks the request done and wakes cat.
  7. The kernel copies the bytes out. Back in the read handler, the kernel copies the bytes into cat's buffer. It does this only through the trusted core's copy routine, which first checks that the buffer is memory cat owns and may write to.
  8. The kernel returns. The handler's result, the number of bytes read, goes into a register. On the way out the kernel checks whether a signal is waiting for cat (if you pressed Ctrl-C, for example) and then returns to user mode, where read() hands cat the count.

Then cat makes a write system call to send the same bytes to its standard output, and a similar trip happens through the terminal instead of the disk.

How the parts are checked

The design only helps if its rules hold, so the build checks them rather than trusting people to remember:

  • A source check fails the build if unsafe appears in any kernel crate other than the trusted core, and another expands every macro first, so unsafe can't hide inside one.
  • Another check fails the build if any kernel code uses async. The kernel blocks a task while it waits; asynchronous code is a userland feature, built on SlopRing.
  • Every system call states, where it is defined, which permission it needs, and the build fails if any entry in the table lacks one.
  • The trusted core's most dangerous invariants are proved with the Verus proof checker, the core runs under Miri, and the kernel's test suite boots it under QEMU on every change. See Proofs, Miri and Safety gates.

For contributors

The source tree is a Cargo workspace with one crate per subsystem, and Cargo.toml lists every one. Crate names are directory names. These are the ones you will open most often:

CrateResponsibility
slopos-ostdThe trusted core, and the only kernel crate allowed unsafe
kernel, bootThe kernel's entry point and the boot sequence
abiEverything shared with userland: syscall numbers, errno values, flags, structs
coreSystem call dispatch and handlers, exec, the OOM killer
schedTask lifecycle, scheduling, sleeps, futexes
mmPhysical and virtual memory, address spaces, ELF loading
fs, drivers, netThe filesystem layer, device drivers, the network stack
*-core (ext4-core, shell-core, ...)Pure logic with no system calls, so its tests also run on the development host
slibc, userlandThe C library and the programs
verificationVerus proofs for the trusted core, checked by just verify

Two pinned third-party crates, vendor/unwinding and vendor/gimli, are part of the trusted code base as named exceptions. They serve OSTD's panic unwinding and nothing else.

Rules a change must keep:

  • No unsafe outside slopos-ostd. If a subsystem needs something only unsafe can do, add a safe API to OSTD. See Unsafe code and FFI.
  • No async in the kernel. Block on a wait queue; see Scheduling and waiting.
  • Linux numbering. A Linux x86-64 system call number means the same call in SlopOS or nothing at all. SlopOS-only calls live in a private range from 1024. See System call ABI.
  • The ABI lives in code. Numbers, structs and flags are defined once, in the abi crate; reference pages describe them but the source is the authority.

Further reading

In the source

WhereWhat
Cargo.tomlThe workspace: every crate
slopos-ostd/The trusted core
kernel/src/main.rsThe kernel's entry point
core/src/syscall/handlers.rsThe system call tables
scripts/check_unsafe_outside_ostd.sh, scripts/check_no_kernel_async.shTwo of the build's rule checks

Safety gates lists every check the build runs.

On this page