Reference

SlopRing

The shared-memory queues a program uses to hand the kernel many I/O operations at once and collect the results.

A normal system call does one thing and waits for it: one read, one accept, one send. A server handling hundreds of connections pays for hundreds of calls, and any one of them can stop the whole program while it waits. With SlopRing the program writes a batch of requests into memory it shares with the kernel, makes one call, and later reads the results back from the same memory. Each request carries a number the program chose (user_data), so results can arrive in any order and still be matched to their requests.

The design is Linux's io_uring (Jens Axboe, 2019), cut down to the operations SlopOS needs: one queue of requests (submissions) and one of results (completions), both in shared memory. SlopRing is not binary-compatible with io_uring. This page lists the calls, memory layouts, opcodes and rules. The layouts, opcodes, flags and limits are defined in abi/src/ring.rs; the kernel side is the ring crate and core/src/syscall/ring_handlers.rs. Userland describes the runtime that drives the ring.

System calls

All three are SlopOS-only calls (see System call ABI).

CallArgumentsReturns
ring_setupentries, pointer to a RingParamsA ring descriptor. Maps the shared region into the caller and copies the layout to the RingParams
ring_enterring_fd, to_submit, min_complete, flagsThe number of submissions consumed. Consumes up to to_submit submissions, then, if min_complete > 0, waits until that many completions are ready
ring_registerring_fd, op, arg, nr_args0. Registers or unregisters buffers (see Registering buffers)
  • entries is the submission queue size: a power of two from 1 to 4096 (SLOPRING_MAX_ENTRIES). The completion queue is twice as large (SLOPRING_SQ_TO_CQ).
  • ring_enter returns -EINTR only if a signal interrupts the wait and it submitted nothing; otherwise it returns the submission count.
  • ring_enter's flags are currently ignored. SLOPRING_ENTER_GETEVENTS is reserved.
  • Every call on a ring from a process other than the one that created it fails with -EBADF.

Shared region

ring_setup allocates one region, maps it into the caller and writes a RingParams at its start. All offsets in RingParams are bytes from the start of the region.

RingParams
SQ control words   head (kernel writes), tail (program writes), mask, dropped
CQ control words   head (program writes), tail (kernel writes), mask, overflow, flags
SQE array          submission i is slot i; there is no indirection array
CQE array
RingParams fieldMeaning
sq_entries, cq_entriesQueue sizes
flagsFeature bits the kernel supports
sq_off_head, sq_off_tail, sq_off_mask, sq_off_dropped, sq_off_arrayOffsets of the submission queue's control words and array
cq_off_head, cq_off_tail, cq_off_mask, cq_off_overflow, cq_off_array, cq_off_flagsOffsets of the completion queue's control words and array
region_addr, region_bytesWhere the region is mapped, and its size
Feature bitMeaning
SLOPRING_FEAT_SINGLE_MMAPThe whole ring is one mapping (the only mode)
SLOPRING_FEAT_MULTISHOTSLOPRING_SQE_MULTISHOT is honoured
SLOPRING_FEAT_REG_BUFFERSFixed and provided buffers are supported

Each control word has one writer. The kernel keeps its own copy of the indices it owns and trusts that copy, not the shared one. It reads and writes the region with bounded copies and acquire/release loads and stores, and never holds a Rust reference into memory the program can change.

Submission entry (Sqe, 64 bytes)

FieldTypeMeaning
opcodeu8Operation
flagsu8SLOPRING_SQE_BUFFER_SELECT (bit 0) or SLOPRING_SQE_FIXED_BUFFER (bit 1); not both
fdi32Target descriptor; -1 for OP_NOP and OP_TIMEOUT
offu64File offset for OP_READ/OP_WRITE; nanoseconds for OP_TIMEOUT
addru64Buffer, msghdr or socket address; for OP_CANCEL, the user_data to cancel
lenu32Length of the buffer at addr
op_flagsu32Per-operation flags: poll mask, send/receive flags, open flags, cancel flags
user_datau64Copied unchanged into the completion
addr2u64Second pointer (the address-length out-pointer for OP_ACCEPT, the source address for OP_RECVFROM)
sqe_flags2u16SLOPRING_SQE_MULTISHOT (bit 0)
buf_groupu16Provided-buffer group, with SLOPRING_SQE_BUFFER_SELECT; 0 otherwise
buf_indexu16Fixed-buffer index, with SLOPRING_SQE_FIXED_BUFFER
reservedu16, u64Must be zero

Completion entry (Cqe, 16 bytes)

FieldTypeMeaning
user_datau64From the submission
resi32Result, or a negated errno
flagsu32Completion flags below; with SLOPRING_CQE_F_BUFFER, bits 16 to 31 hold the provided-buffer id
Completion flagBitMeaning
SLOPRING_CQE_F_MORE0More completions follow for this operation
SLOPRING_CQE_F_BUFFER1The upper 16 bits carry a provided-buffer id
SLOPRING_CQE_F_BUF_MORE2Reserved; not emitted
SLOPRING_CQE_F_NOTIF3Final completion of OP_SEND_ZC: the buffer may be reused

Opcodes

ValueOpcodeOperationNotes
0OP_NOPNothingCompletes with 0
1OP_READread/pread on a file or socket
2OP_WRITEwrite/pwrite on a file or socket
3OP_RECVMSGrecvmsg, including Unix sockets and passed descriptors (SCM_RIGHTS)addr is the msghdr; multishot
4OP_SENDSend a buffer
5OP_ACCEPTAccept a connection and install its descriptorMultishot
6OP_POLL_ADDWait until a descriptor is readyop_flags is the poll mask; multishot
7OP_TIMEOUTDeadline for the current waitoff is nanoseconds; completes with -ETIME when it passes
8OP_CANCELCancel in-flight operationsSee below
9OP_RECVFROMReceive a datagram and its source addressSource address written to addr2
10OP_OPENATOpen a path and install a descriptoraddr is the path, len its length, op_flags the open flags; same path rules as open
11OP_CLOSEClose a descriptorCompletes at once
12OP_SEND_ZCSend from a fixed buffer without copyingRequires SLOPRING_SQE_FIXED_BUFFER; otherwise -EINVAL
13OP_CONNECTConnect a socketaddr is a SockAddrIn or SockAddrUn; one completion, 0 or a negated errno
  • An opcode above 13 (OP_MAX), or one outside the ring's allowed set, completes with -EPERM.
  • Each ring has a set of allowed opcodes fixed at ring_setup; nothing can widen it later. ring_setup currently allows all fourteen.
  • Every opcode has a capability class fixed at compile time, like a system call, and a compile-time check requires that none needs a capability beyond the descriptors it names. A ring therefore carries no copy of its creator's credentials. See Permissions.
  • Descriptors named in a submission are looked up in the creating process's descriptor table.

OP_CANCEL removes the in-flight operation whose user_data equals the cancel entry's addr, posts -ECANCELED for it, then posts the cancel entry's own result: 0 if it found one, -ENOENT if not. With SLOPRING_ASYNC_CANCEL_ALL in op_flags it removes every match.

Completion rules

  • ring_enter first tries each submission without blocking. If it finishes, the completion is posted at once.
  • If it would block, the operation stays in flight. The kernel resolves its descriptor once, at submission, and holds the open file until the operation ends, so closing or reusing the descriptor number does not affect it.
  • With min_complete > 0, the caller sleeps until the files its in-flight operations wait on become ready, retries them, and returns once min_complete completions are ready, a signal arrives, or the nearest OP_TIMEOUT passes.
  • Completions are posted only inside ring_enter. No kernel thread completes operations in the background, and the ring descriptor never becomes readable.
  • If the completion queue is full, a completion that can be dropped is counted in the overflow word and sets SLOPRING_CQ_OVERFLOW (bit 1) in the completion flags word. The flag stays set until the ring is set up again.
  • OP_ACCEPT and OP_OPENAT (which create a descriptor) and OP_READ, OP_RECVMSG and OP_RECVFROM (which consume data) reserve a completion slot before they run, so their completion is never dropped.
  • A multishot operation (SLOPRING_SQE_MULTISHOT, honoured for OP_ACCEPT, OP_RECVMSG and OP_POLL_ADD) stays in flight and posts one completion per event with SLOPRING_CQE_F_MORE set. The last completion has the flag clear. Cancelling it posts one -ECANCELED with the flag clear.

Registering buffers

opValueargnr_args
RING_REGISTER_PBUF_RING1Pointer to a RegisterBufRingCmdunused
RING_REGISTER_BUFFERS2Pointer to an array of BufIovecArray length
RING_UNREGISTER_PBUF_RING3Group idunused
RING_UNREGISTER_BUFFERS4unusedunused

Any other op returns -ENOSYS.

Fixed buffers. RING_REGISTER_BUFFERS pins up to 1024 buffers (SLOPRING_MAX_FIXED_BUFFERS), each {addr: u64, len: u32, _pad: u32} and at most 1 GiB (SLOPRING_MAX_REG_BUF_BYTES). A submission uses one by setting SLOPRING_SQE_FIXED_BUFFER and buf_index. The buffer is reserved from submission until the operation's final completion or cancellation, and RING_UNREGISTER_BUFFERS returns -EBUSY while any reservation is held.

Provided buffers. RING_REGISTER_PBUF_RING pins a ring of buffers the program refills, for one buffer group:

RegisterBufRingCmd fieldTypeMeaning
ring_addru64Address of the buffer ring
ring_entriesu32Power of two, at most 32768 (SLOPRING_PBUF_RING_MAX_ENTRIES)
buf_groupu16Group id, 1 to 64 (SLOPRING_MAX_BUF_GROUPS); -EEXIST if already registered
flagsu16SLOPRING_PBUF_RING_INC (bit 0) reserved for incremental consumption

Each ring entry is an IouringBuf {addr: u64, len: u32, bid: u16, resv: u16}. The program publishes buffers by advancing a u16 tail at byte offset 14 of the ring, which overlaps bufs[0].resv as in Linux's io_uring_buf_ring. A receive submission sets SLOPRING_SQE_BUFFER_SELECT and buf_group; the kernel takes the next published buffer, fills it, advances the ring head and reports the buffer id in the completion. RING_UNREGISTER_PBUF_RING returns -EBUSY while any in-flight operation selects that group.

Zero-copy send

OP_SEND_ZC posts a result completion with SLOPRING_CQE_F_MORE, then a final completion with SLOPRING_CQE_F_NOTIF once the kernel has released the buffer.

SocketBehaviour
IPv4 datagramThe network card reads directly from the pinned buffer; the final completion follows once the card has finished with it
TCPThe pinned pages stay queued, including for retransmission; the final completion follows once the data is acknowledged and the card has finished with it
Anything else, or when zero copy is not possible (no checksum offload, address not yet resolved, Unix sockets)The kernel copies the data once and posts both completions at once

Ring descriptor

  • The descriptor can be closed and duplicated within the process. It is not inherited across fork or exec.
  • It cannot be polled; read and write on it fail.
  • Rings live in a table of handles with generation numbers, so a closed and reused descriptor, a descriptor of another kind, or a stale handle gets an error rather than someone else's ring.

See also

  • Efficient IO with io_uring by Jens Axboe: the design SlopRing follows, with the reasoning behind shared queues, registered buffers and polling.
  • io_uring(7): the Linux interface, for comparison with the fields above.

On this page