User mode and kernel mode
How SlopOS uses the CPU's user and kernel modes, and how control passes between a program and the kernel.
Programs on SlopOS run in the CPU's user mode and the kernel in kernel mode, as on Linux. What is specific to SlopOS is where the boundary code lives and how it is shaped: every entry into the kernel and every exit from it goes through the trusted core, the kernel always returns to a program with the same instruction, and it touches a program's memory only through one routine that checks the program owns that memory.
OSTEP calls the general technique limited direct execution: run the program directly on the CPU, after setting up the hardware to limit what it can do. The next two sections recap how x86-64 enforces the limits and how control crosses them; the SlopOS specifics come after, including where it differs from Linux.
How the CPU enforces the limits
Two privilege levels
An x86 CPU always runs at a privilege level. The hardware offers four, numbered 0 to 3 and called rings, but SlopOS, like Linux and Windows, uses only two: ring 0 for the kernel (kernel mode) and ring 3 for programs (user mode). In user mode the CPU refuses the instructions that would let a program take over the machine: turning interrupts off, switching to a different set of memory mappings, talking to devices through I/O ports, or halting the CPU. Trying one causes a fault: the CPU stops the program at that instruction and runs a kernel handler instead.
Each program sees only its own memory
Programs don't use physical memory addresses. Every address a program uses goes through a page table, a structure the kernel builds for each process that maps the program's addresses, in chunks of 4 KiB called pages, to physical memory. A page that isn't in the table doesn't exist as far as the program is concerned, so a program has no way even to name another program's memory.
Each entry in the table also carries permission bits: whether user mode may use the page at all, whether it may be written, and whether its contents may be run as code. The CPU checks them on every access. SlopOS maps the kernel into the top half of every process's address space, so that entering the kernel doesn't require switching tables, but marks all of it as kernel-only. A program that reads a kernel address gets a fault, the same as one that reads an address nothing is mapped at. Memory explains how the kernel builds and changes these tables.
Protecting the kernel from its own mistakes
Two further CPU features turn kernel bugs that would be exploitable into crashes. With SMEP (supervisor mode execution prevention) the CPU refuses to run code from a user page while in kernel mode, so a bug that sends the kernel to a program-supplied address can't run the program's code with kernel privileges. With SMAP (supervisor mode access prevention) the CPU refuses kernel reads and writes of user pages unless the kernel has opened an explicit window first, so a kernel bug can't read or write program memory by accident. SlopOS turns both on for every CPU that has them.
How control passes between a program and the kernel
A program in user mode can't call into the kernel like a function, because then it could jump into the middle of the kernel, past whatever checks the kernel was about to make. So the CPU only lets control into the kernel at entry points the kernel registered at boot. There are three ways in:
- A system call. The program asks for something on purpose, with the
syscallinstruction. System calls follows one from start to finish. - An interrupt. A device needs attention, or the timer fires. The timer is what guarantees the kernel gets the CPU back even from a program stuck in an infinite loop: SlopOS sets it to fire 100 times a second, and each time the scheduler can decide to run something else.
- A fault. The program did something it isn't allowed to, or touched a page the kernel hasn't filled in yet.
In each case the CPU switches to kernel mode, switches to a stack that belongs to the kernel (the program's own stack can't be trusted), saves where the program was, and jumps to the entry point. When the kernel is done, it restores the program's registers and executes a return instruction that drops back to user mode and resumes the program exactly where it stopped. The program never notices it was interrupted, except that time passed.
What happens when a program faults
Not every fault is a mistake. The kernel often gives a program memory without filling it in yet, and fills in each page the first time the program touches it; it also shares pages between a parent and child process until one of them writes, and only then makes a private copy. Both start as a fault, the kernel resolves it, and the program carries on without knowing.
A fault the kernel can't resolve, such as a read from an address with nothing
mapped, becomes a signal, usually SIGSEGV. If the program has installed a
handler for it, the kernel runs the handler; otherwise the program ends,
with the signal as its exit status. Either way, the rest of the system is
unaffected. Processes and signals explains
signals and how a program ends.
How SlopOS crosses the boundary
The machinery that crosses between user mode and kernel mode is some of the most delicate code in any kernel: a mistake in it can hand a program the kernel's privileges. In SlopOS all of it lives in the trusted core (OSTD, described in The trusted core): the CPU's descriptor tables, the entry and exit code, the switch between page tables, and the routine that copies program memory. The rest of the kernel reaches it only through safe Rust interfaces.
One way into user mode. The kernel runs each program as a loop: enter user mode, wait for something to bring control back, deal with it, and enter user mode again. Only one function in the trusted core enters user mode, so there is one place that loads a program's registers and one place that makes sure nothing from the kernel leaks into them.
System calls come back to that loop. The syscall instruction lands in a
short piece of assembly in the trusted core. It saves the program's registers
into the task's saved register record and returns to the loop, which runs the
system call and then enters user mode again.
Interrupts and faults are handled where they land. The trusted core's interrupt entry code saves the program's registers on the task's kernel stack and calls the kernel's handler. If the handler decides another task should run, the switch happens there, and the original task resumes later from the same point.
One way out. SlopOS always returns to user mode with IRETQ, the
instruction that returns from an interrupt, even after a system call. x86-64
also offers a faster return made for system calls (SYSRET), which SlopOS
doesn't use.
One way to touch program memory. When the kernel needs to read a system
call's arguments from program memory or write results back, it goes through
the trusted core's copy routine. The routine checks that the whole range is
in the program's half of the address space, then checks the program's page
table to see that every page is present and accessible to the program (and
writable, for a copy into it). Only then does it open the SMAP window, copy,
and close it again; no other code in the kernel can open that window. If a
page isn't filled in yet, because the program allocated memory but hasn't
touched it, the kernel fills it in the way a fault would have and tries once
more. If a fault interrupts the copy halfway, the routine stops, reports how
far it got, and the copy is retried once. A buffer that still can't be
copied makes the system call fail with EFAULT; it never crashes the kernel.
The legacy interface is closed. Old 32-bit Linux programs make system
calls with the int 0x80 instruction. SlopOS doesn't support that; the
instruction is answered with ENOSYS ("function not implemented") rather
than killing the program.
Switching between programs cheaply
Every switch from one process to another loads a different page table. The CPU caches recent address translations in a TLB (translation lookaside buffer), and loading a new table normally empties that cache, so the next program starts with every memory access slow until the cache refills.
Where the CPU supports process-context identifiers (PCIDs), SlopOS tags each
cached translation with the address space it belongs to, so switching tables
doesn't have to empty the cache. Each CPU keeps 16 tags for programs (plus
one for the kernel alone) and remembers which address space last held each
one, so switching back to a recently run program reuses its cached
translations. The kernel's own mappings are marked global, which keeps them
in the cache across every switch. The scheme follows Linux's
(arch/x86/mm/tlb.c). On CPUs without PCIDs, or with a known erratum that
disables them, every switch empties the program part of the cache, which is
slower but correct.
What SlopOS doesn't do
No kernel page-table isolation. In 2018 the Meltdown flaw showed that on many Intel CPUs a program could read memory that is mapped but marked kernel-only, by exploiting speculative execution. Linux's defence, KPTI, gives each process a second page table with the kernel almost entirely removed while user code runs. SlopOS doesn't do this: the kernel stays mapped, kernel-only, in every process's table, so on a CPU with that flaw a program could read kernel memory. Known limitations lists this.
How it is tested
Tests check the boundary from both sides. Kernel tests check the CPU's descriptor tables and the system call registers after boot, that a deep chain of traps from user mode can't overwrite the frames the kernel returns into, and that a user copy which faults partway through is retried. Host tests of the trusted core check that the copy routine refuses addresses outside the program's half, pages the program can't access and read-only pages. The invariant that user mappings can't reach kernel memory is one of the properties proved with Verus (see Proofs).
For contributors
The rules any change to this boundary must keep:
- Only OSTD crosses. Descriptor tables, MSR access, the naked entry and
exit assembly, CR3 writes, user copies and task switching belong to
slopos-ostd. Every other kernel crate is#![forbid(unsafe_code)], and Safety gates lists the checks that keep it that way. UserMode::executeis the only way into ring 3. It returns aReturnReasonread from the task's ownUserContext, never from a per-CPU slot, so a preemption or migration after entry can't mix up two tasks.- The
SYSCALLtrampoline (__ostd_user_return) never pushes onto the kernel stack. The top of that stack is where the next interrupt'sIRETQframe goes, so user state is spilled through per-CPU scratch slots instead. The kernel's user loop handles onlyReturnReason::Syscall; interrupts and exceptions from ring 3 enter through the OSTD IDT stubs and return with their ownIRETQ. SFMASKclears the interrupt, trap, direction, alignment-check, I/O privilege and nested-task flags onSYSCALL, and the trampoline re-enables interrupts.STARfollows the AMD64 selector order, with user data before user code in the GDT.- Every context switch points
TSS.RSP0and the per-CPUSYSCALLstack at the incoming task's kernel stack, so no entry can land on another task's stack. - IST stacks. Double faults, stack faults and general-protection faults
run on per-CPU IST stacks with guard pages, as do the keyboard and mouse
interrupts. The NMI deliberately has none: any
IRETQunblocks NMIs, and a nested NMI would reset to the top of the same stack. The page fault has none so that a user fault can block and push a signal frame. The kernel is also built with SafeStack, which keeps address-taken locals on a separate stack. - User copies return values, not references.
copy_from_user::<T>returnsTby value, so a user pointer never becomes a kernel reference, and system call arguments that name user memory arrive as typed wrappers (UserPtr,UserSlice,UserPath). No public API exposesSTACorCLAC; the only window is between__ostd_usercopy_startand__ostd_usercopy_end, which the page-fault handler recognises. - Address spaces. Every process table shares the kernel half (PML4
entries 256 to 511) with the kernel master table, supervisor-only and
global. The user half is changed only through VM-space cursor operations.
PCID 0 is kernel-only loads; PCIDs 1 to 16 are per-CPU slots.
INVPCIDis used where the CPU has it. The dual-table KPTI scaffolding inmm/src/mmu/kpti.rsandmm/src/mmu/trampoline.sis not linked in.
Further reading
- Mechanism: Limited Direct Execution, chapter 6 of Operating Systems: Three Easy Pieces by Remzi and Andrea Arpaci-Dusseau. Start here. It builds the whole idea from the problem up: user and kernel mode, traps, the trap table, and the timer interrupt.
- The Abstraction: Address Spaces and Paging: Faster Translations (TLBs), chapters 13 and 19 of the same book. Why each program gets its own view of memory, and why switching between those views is expensive, which is what PCIDs fix.
- Supervisor mode access prevention, LWN (2012). What SMAP is for and how Linux opens and closes the window around its own user copies.
- The current state of kernel page-table isolation, LWN (2017). How KPTI works and what it costs, written as Linux adopted it; the background to the one thing SlopOS doesn't do.
- Kernel entries, from the Linux kernel documentation. A short, dense list of every way into the Linux kernel on x86-64, for comparison with the three described here.
In the source
| Where | What |
|---|---|
slopos-ostd/src/user/mode.rs, slopos-ostd/src/user/asm/user_return.s | Entering user mode and the SYSCALL return trampoline |
slopos-ostd/src/irq/asm/handlers.s, slopos-ostd/src/irq/idt.rs | Interrupt and exception entry |
slopos-ostd/src/user/copy.rs | The user-memory copy routine |
slopos-ostd/src/cpu/x86_64/security.rs | Turning on SMEP, SMAP and global pages |
core/src/syscall/user_loop.rs | The loop that runs each user task |
boot/src/gdt.rs, boot/src/ist_stacks.rs | System call registers and IST stacks |
mm/src/mmu/asid.rs | PCID slots |
boot/src/exception.rs | User faults becoming signals |