Skip to content

About

A bare-metal microkernel written from scratch in no_std Rust, targeting x86_64-unknown-none

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

76 Commits

Folders and files

Repository files navigation

Kernel Panda

Kernel Panda

A bare-metal microkernel written from scratch in no_std Rust, targeting x86_64-unknown-none.

It boots on bare metal or under QEMU, brings up every processor, and runs preemptively scheduled threads in their own address spaces. Drivers and the display server live in Ring 3 and talk through capability-mediated IPC. It has persistent storage: SATA, NVMe and virtio-blk drivers, GPT partitioning, and a copy-on-write filesystem that survives a power cut. It is on the network, with the protocol stack running as an unprivileged process.

Memory Bitmap frame allocator, four-level paging, per-process address spaces, kernel heap
Protection NX, SMEP, SMAP, W^X, guard-paged kernel stacks, per-process quotas, per-process users with no superuser, password logins, a CSPRNG seeded from RDSEED/RDRAND
Scheduling Preemptive, three priorities, per-CPU run queues with work stealing, sleep and join
Multiprocessing Every core started and scheduling, ticket locks, acknowledged TLB shootdown
User space Ring 3, a trap-gate syscall surface, ELF loading, preemptible system calls, granted I/O ports and interrupt lines
IPC Bounded endpoints, unforgeable sender identity, SEND/RECEIVE/GRANT capabilities, and shared message rings that cost no system call per message
Devices Local APIC and I/O APIC, PCIe with ECAM and MSI-X, AHCI, NVMe and virtio-blk storage, virtio-net, framebuffer, 16550 serial, PS/2 keyboard and mouse (driven from Ring 3)
Networking ARP, IPv4, IPv6 with NDP and SLAAC, ICMP echo, UDP, TCP, DHCP and DNS in a Ring 3 daemon; the kernel only moves Ethernet frames
Storage Block layer, GPT and MBR, a copy-on-write filesystem with atomic commits, owners and permission bits
Graphics Shared buffers with capability-checked handles, a Ring 3 compositor with z-order, damage tracking, a pointer and click-to-focus

Testing: 230 cases across 29 boot-and-assert test kernels, run on four cores under QEMU with SMEP and SMAP enabled.

Kernel Panda v0.1.0
  serial console : COM1 @ 38400 8N1
  framebuffer    : online
  descriptor tbls: GDT + TSS + IDT loaded
  timer          : Local APIC, periodic

physical memory map:
  0x00000000000..0x0000009fc00    639 KiB  Usable
  0x0000009fc00..0x000000a0000      1 KiB  UnknownBios(2)
  ...
  usable: 246 MiB across 11 regions

frames: 65504 total (255 MiB), 62830 free (245 MiB), 2674 in use
heap:   1048576 bytes total, 0 allocated, 0 peak, 0 live allocations

alloc smoke test: [1, 4, 9, 16, 25, 36, 49, 64]

timer at 100 Hz, waiting for ticks:
  uptime   270 ms  (27 ticks)
  uptime   470 ms  (47 ticks)
  uptime   670 ms  (67 ticks)
  uptime   870 ms  (87 ticks)
  uptime  1070 ms  (107 ticks)

scheduler: spawning two workers that never yield
  [worker-b] step 1 of 3
  [worker-a] step 1 of 3
  [worker-b] step 2 of 3
  [worker-a] step 2 of 3
  [worker-b] step 3 of 3
  [worker-a] step 3 of 3
  both workers finished; 3 threads live, running as 'boot'

ring 3: loading a user program and dropping privilege
  [ring 3] hello from user space
  [ring 3] still running after a yield
  user program exited after writing 72 bytes through syscalls

ipc: a blocking logger thread fed over a capability
  [logger] tag 0x0101 word0  10 from thread 0
  [logger] tag 0x0102 word0  20 from thread 0
  [logger] tag 0x0103 word0  30 from thread 0
  [logger] tag 0xcafe word0 48879 from thread 6
  logger exited; endpoint drained to 0

compositor: a ring 3 display server
  compositor mapped the scanout buffer and is waiting for surfaces
  presented a blue surface at (120, 260)
  presented a green surface at (260, 260)
  presented a red surface at (400, 260)
  input daemon sent shutdown; both daemons exited

net: a ring 3 stack pings the gateway
  reply from 10.0.2.2 in 20 ms, parsed entirely in Ring 3

shell: a ring 3 daemon reading the serial port
panda> help
commands: help version hello exit
panda> version
Kernel Panda, ring 3 shell
panda> exit
shell exiting

pci: 6 devices
  00:00.0  8086:1237  class 06.00  host bridge
  00:02.0  1234:1111  class 03.00  display controller
  ...
    bar0: Memory { address: fd000000, size: 1000000, prefetchable: true }

Neither worker yields — the interleaving is entirely the timer taking the CPU away from them. Thread 6 is a Ring 3 process sending through a SEND-only capability; the kernel stamped its identity into the message, overwriting the value the program had put there.

Building and running

Requires a Windows host with QEMU installed (winget install SoftwareFreedomConservancy.QEMU). The Rust toolchain pins itself via rust-toolchain.toml; rustup will fetch it on first build. QEMU does not need to be on PATH — xtask looks in C:\Program Files\qemu, or honours the QEMU / QEMU_DIR environment variables.

cargo xtask build                     # compile the kernel, emit BIOS + UEFI disk images
cargo xtask run                       # boot in QEMU with a window
cargo xtask run --headless --timeout=10   # boot, capture the serial log, exit
cargo xtask run --uefi                # boot via OVMF instead of BIOS
cargo xtask run --verbose-boot        # re-enable the bootloader's own logging
cargo xtask test                      # boot every test kernel and assert on the result

BIOS is the default because it needs no firmware blob and starts faster. The kernel is boot-mode agnostic: bootloader_api normalises the firmware memory map either way, so everything here behaves identically under --uefi.

Layout

kernel-panda/
├── xtask/           host-side build driver: images, QEMU, test runner
├── userland/        Ring 3 programs in Rust (its own cargo workspace)
│   ├── src/lib.rs   syscall wrappers, entry macro, panic handler
│   └── src/bin/     shell, compositor, input daemon (the PS/2 driver), network daemon, client, test probe
└── kernel/          the kernel itself (its own cargo workspace)
    ├── src/
    │   ├── console/   16550 UART, framebuffer text console, 8x8 font
    │   ├── arch/x86_64/   GDT + TSS, IDT, Local APIC + timer, 8259 masking
    │   ├── memory/    memory map, frame allocator, page tables, heap, kernel stacks, DMA regions
    │   ├── allocator/ bump and linked-list `GlobalAlloc` implementations
    │   ├── sched/     threads, context switch, priorities, per-CPU run queues
    │   ├── block/     block layer, AHCI, NVMe and virtio-blk drivers, GPT and MBR partitioning
    │   ├── fs/        copy-on-write filesystem and its formatter
    │   ├── crash.rs   the panic handler: stop, report, save, find next boot
    │   ├── device.rs  I/O port and interrupt-line grants for Ring 3 drivers
    │   ├── smp.rs     starting the other processors, per-CPU identity
    │   ├── acpi.rs    MADT, MCFG and DMAR: processors, I/O APICs, the PCIe window, VT-d
    │   ├── iommu/     VT-d remapping-unit inventory; no domain yet confines a device to it
    │   ├── quota.rs   per-process resource limits
    │   ├── userspace.rs  user regions, program loading, the drop to Ring 3
    │   ├── users.rs      which user each thread runs as, accounts, logging in
    │   ├── sha256.rs     SHA-256, HMAC and PBKDF2, for password hashes
    │   ├── syscall.rs    the entire Ring 3 surface
    │   ├── ipc.rs        endpoints, capabilities, blocking receive
    │   ├── ring.rs       message rings two processes share
    │   ├── pci.rs        bus enumeration, BAR decoding, ECAM, MSI-X
    │   ├── net.rs        the virtio-net driver and the frame-moving syscalls
    │   ├── timer.rs      one-shot timers, delivered as messages
    │   ├── random.rs     random bytes, seeded from the processor's generator
    │   ├── virtio.rs     virtio's legacy PCI interface and virtqueues, shared
    │   └── gbm.rs        shared graphics buffers and the scanout
    └── tests/       one standalone boot-and-assert kernel per file

kernel/ is deliberately a separate cargo workspace, excluded from the root. [unstable] build-std in a .cargo/config.toml applies to the whole cargo invocation and cannot be scoped per-target; sharing a workspace would make cargo try to build std from source for the Windows host. Two workspaces also avoid depending on the unstable -Z bindeps feature.

For the same reason, the cargo runner key points at the compiled target/release/xtask.exe rather than at cargo run -p xtask: a nested cargo launched from kernel/ would inherit that build-std block and try to build xtask for bare metal.

Dependency policy

Everything that ships inside the kernel image has to earn its place. This is a kernel meant to be read and audited, and a dependency is code nobody here has read. Two crates earn it.

Crate Why
bootloader_api Required by the chosen boot path.
x86_64 IDT/GDT/page-table structures and privileged instructions. Pure Rust; reimplementing is weeks of work for no safety gain.

spin was the third. It supplied the lock until the in-house ticket lock replaced it, then Once and Lazy until those were written too. Both swaps touched one file, kernel/src/sync.rs, because every call site had gone through it from the start.

Written in-house rather than pulled in: the 16550 UART driver, the framebuffer console and its font, the physical frame allocator, both heap allocators, the locks and one-time initialisation, and SHA-256 for password hashes. linked_list_allocator is deliberately unused — ours is ~250 lines and the audit goal is the point.

xtask is host-side tooling and is not subject to this budget.

Design notes

Bitmap frame allocator, not a bump. A microkernel returns physical memory every time a Ring 3 process exits, so an allocate-only design would have had to be thrown away as soon as user space existed. The bitmap is sized from the highest usable address rather than the highest address in the memory map — firmware puts MMIO windows near the top of the address space (QEMU's sits at 0xfd_0000_0000), and covering up to there would mean a 32 MiB bitmap carved out of 246 MiB of real RAM. Device memory is mapped through paging::map_to_frame, which takes an explicit frame and never consults the bitmap.

Coalescing heap. The free list is address-sorted so a returned block only has to inspect its two immediate neighbours to merge. Without this, alternating allocations and frees shatter the heap into unusable fragments and the kernel dies of exhaustion with most of its memory nominally free. heap_alloc.rs asserts that emptying the heap leaves exactly one free block of HEAP_SIZE.

GDT before IDT. The double-fault handler runs on a dedicated 20 KiB stack via the TSS Interrupt Stack Table. Without it, a stack overflow faults, then faults again trying to push the exception frame, then triple-faults and resets the machine. stack_overflow.rs exists solely to prove the difference — both outcomes otherwise look identical from outside.

Two heap allocators. The bump allocator is kept as a diagnostic: if Box::new faults under --features bump-allocator, the fault is in the page mapping, not the allocator. Both pass the full suite.

Local APIC, not the 8259 PIC. Legacy hardware support is out of scope, and per-CPU delivery is a prerequisite for SMP later, so PIC-based delivery would be throwaway work. The PIC is still remapped clear of the exception vectors and then fully masked — masking alone is not enough, because a spurious IRQ 7 can be delivered on a masked controller, and at power-on those land on vectors that look exactly like CPU exceptions.

The timer is calibrated, not assumed. The APIC counts at the core crystal frequency, which varies by machine and is not dependably reported by CPUID. It is measured at boot against PIT channel 2 — channel 2 because its output is readable from a port, so calibration needs no interrupt, which matters when interrupts are still masked. The poll is bounded so a machine whose PIT never asserts fails cleanly instead of hanging the boot.

The context switch exchanges only callee-saved registers. context_switch is called like an ordinary C function, so the compiler has already spilled anything caller-saved at the call site; saving it again would be wasted work. Between the two halves the whole register file is covered.

The boot thread cannot double as the idle thread. Idle is by definition the last choice, so any CPU-bound worker would starve it permanently — and with it, whatever the kernel booted into. There is a separate idle thread that only runs when the ready queue is empty.

The scheduler lock is released before the switch. Holding a spinlock across a context switch leaves it locked by a thread that is no longer running. This is safe only because the whole scheduler runs with interrupts disabled on a single core, so nothing can observe the gap. It is the first thing that will need rethinking for SMP.

Entry to the kernel is an interrupt gate, not SYSCALL. An interrupt gate switches to the stack in TSS.privilege_stack_table[0] automatically, where SYSCALL does not switch stacks at all and needs swapgs plus a per-CPU block to find one. The gate costs more cycles and buys a great deal less that can go quietly wrong. The ABI does not depend on the mechanism, so it can be swapped later.

A user fault kills the thread, not the kernel. A fault in unprivileged code never take the system down, so the page-fault and GP handlers check the saved CS and, if the fault came from Ring 3, destroy that thread and carry on. ring3.rs proves it by running a program that dereferences the kernel heap: the fault reports PROTECTION_VIOLATION | USER_MODE, meaning the page is mapped and the USER_ACCESSIBLE bit is what stopped it — not a lucky unmapped address.

Every pointer from Ring 3 is walked before it is believed. A user pointer is an attacker-controlled integer. Validation checks both that the range lies inside the user region and that every page it spans is present with the permissions the access needs. Checking only the base is the classic confused-deputy hole, so there are tests for a range that runs off the end and for a length that wraps the address space.

The kernel runs on every core. Application processors come out of reset in 16-bit real mode, so they are started through a trampoline that walks them back up to long mode — copied into low memory and identity mapped, because the instant it enables paging it keeps executing at the address it is already at. Every address inside that page is a compile-time constant offset, so nothing is patched at runtime.

Each CPU has its own GDT, TSS and double-fault stack. privilege_stack_table[0] names the stack that processor traps onto; sharing one would have two cores landing on the same stack and destroying each other's frames — corruption rather than a fault.

No thread is runnable, or freeable, while a processor is still standing on its stack. Releasing the scheduler lock before a context switch was safe with one core because nothing could observe the gap. With more, the window between "stops being this CPU's current" and "the mov rsp inside context_switch" is a window in which the thread looks idle and is not. Another core resuming it there loads a saved stack pointer that has not been written yet; another core freeing it there unmaps a live stack.

Each thread therefore carries an on_cpu flag, set when a CPU commits to switching to it and cleared by the incoming context once the switch has actually happened. The ready queue, unblock and reap all respect it. unblock is the one that is easy to miss — a wake arrives from another core at a moment of the waker's choosing, including while the thread it is waking is halfway off its processor. It sets the state and leaves the enqueue to the handshake.

Symptom when this was wrong: an intermittent double fault with rsp of zero, one run in ten, from a blocking IPC receive.

The filesystem never overwrites live data. The obvious design writes a file's blocks, then updates the pointer to them. A power cut between the two leaves a directory naming a block that holds something else, and nothing on the disk records that it happened — so the next mount reads corruption and believes it.

Instead, every block a change touches is written to free space, and the path from it up to the root is rewritten the same way. The last step is a single sector: the superblock naming the new root. Until that sector lands the old tree is complete and the disk mounts exactly as it was; after it lands the new tree is live. There is no in-between a reader can observe.

Two superblocks alternate, and the higher generation wins. Both carry a checksum, so a superblock torn mid-write fails it and loses to its sibling rather than being believed. That is the entire crash-recovery story: no journal, no replay, and therefore no replay bugs. The guarantee is exactly one commit deep — once a second commit lands, the blocks of the version before it are free and may be reused.

Blocks a transaction stops referring to are released only after its commit has landed. Releasing them earlier would let that same transaction allocate one and write over a block the old tree still needs, which is the one way copy-on-write can still corrupt itself.

The first version of this claimed copy-on-write and updated directory blocks and inodes in place; only file data actually had it. a_crash_before_the_commit_ leaves_the_previous_state is what caught that, and it is why the machinery is now in one place rather than repeated per operation.

The disk driver is in the kernel, and that contradicts the Ring 3 rule. A disk controller is a DMA engine: it writes wherever its command tables point, and those are physical addresses the device does not check against anyone's page tables. A Ring 3 driver handed that controller can write to any physical page in the machine by asking the hardware to do it. So a Ring 3 disk driver without an IOMMU is not isolated — it merely looks isolated, which is worse than an honest kernel driver because it invites trust it has not earned. Making it real needs VT-d with a per-device domain. The layer above it — partitions, filesystem, policy — has no such excuse and is meant to move out.

Resource limits are per process, and they only ever narrow. They were four constants identical for every thread, which bounds the damage one process can do but says a throwaway test program deserves the same share as the compositor, and gives whoever launches a process no way to say otherwise. A thread's limits are now set by its spawner, capped by what the spawner itself holds — the rule the capability system already uses. Without that cap the limits mean nothing: a process at its ceiling spawns a helper, grants it a larger quota, and asks the helper for memory.

The APIC has no cacheable alias. Its registers are mapped uncached at their own address, and the bootloader's physical-memory window maps every physical address including that one — cacheable, in a 2 MiB page. Two views of one device with different memory types is architecturally undefined, and in practice means a speculative read through the cacheable one can hold a value the device has since changed.

Fixing it needed the huge page split, since changing the memory type of the whole 2 MiB would take neighbouring device registers with it. Splitting moves the permissions down a level rather than duplicating them: read, write and Ring 3 access are the AND of every level, so the new parent is made permissive and the leaves carry the restrictions — but NO_EXECUTE is ORed down the path instead, so it has to be cleared on the parent or the split would silently make an executable range non-executable.

Timer calibration takes five samples and uses the median. One 10 ms sample is at the mercy of whatever else the machine was doing during those 10 ms, and under a hypervisor — the only place this has run — the host can deschedule the guest mid-measurement. The APIC count then reflects the pause, and the timer runs at the wrong rate for the life of the boot with nothing to say so. The median rather than the mean because the errors are one-sided: a stolen slice makes a sample far too large and nothing makes one too small.

The compositor keeps a surface table, composes back to front, and only redraws what changed. It was a blitter: each message painted straight into the scanout as it arrived, so what ended up on top was decided by which client sent last. Surfaces now carry a depth and are composed in that order, into an off-screen buffer that reaches the display in one copy per frame — drawing directly into the scanout lets the display controller read a half-composed frame, which with overlapping surfaces is a visible flicker of whatever was underneath.

Composition has to clear before it draws, or a surface that moved leaves its old pixels behind; and having cleared, it has to redraw every surface intersecting the cleared region, not just the newest. Both directions are tested, because each one alone passes a plausible-looking wrong implementation.

Damage is a grid of 32-pixel tiles, not one bounding box. A surface that moves damages where it was and where it went, and one rectangle covering both recomposes everything in between — for a surface crossing the screen, the screen. It was a list of eight regions before, merging the least wasteful pair when full; a bitmap cannot fill, so there is no degraded mode to reason about, at the cost of rounding each update out to whole tiles. Each row's adjacent damaged tiles are composed as one run.

Each pixel reaches the display as one four-byte store, so the display can never latch half of one — a row memcpy copies in eight-byte strides, which split three-byte pixels. The fourth byte belongs to the next pixel and carries the value already on the display, not the back buffer's, since outside the damage the two need not agree; the last pixel of the display ends its store at itself instead. Under QEMU no watcher caught a torn pixel even from a byte-at-a-time flush, so this is argued rather than observed; only the corner case is tested.

Tearing is observed, not argued: a thread on another core samples one pixel while that area is recomposed repeatedly, and checks it never catches the cleared-to-black intermediate state. Around 25,000 samples per run see it zero times; composing straight into the scanout instead makes it about 16.

Writing that test found something the design does not cover. The final copy to the scanout is a memcpy, and a 24-bit pixel is three bytes written non-atomically — so a reader can catch one byte updated and the next not. The first version of the test alternated blue (FF,00,00) and green (00,FF,00), and a half-written pixel between them reads as (00,00,00): indistinguishable from the cleared state it was hunting. Double buffering prevents a frame from being seen half-composed; it does not make the copy atomic, and was never going to. The test now alternates two colours differing in one byte, so it measures the property it names.

Each of these was checked by breaking it. depth_decides_what_is_on_top_not_ arrival_order fails against the previous arrival-order compositor; a_surface_that_moves_far_does_not_repaint_everything_between fails against a single bounding box; the_display_never_shows_a_half_composed_frame fails against composition without the back buffer.

PCI configuration space is reached by memory when the firmware describes a window. The port mechanism latches an address in 0xCF8 and reads 0xCFC, which works everywhere and reaches only the first 256 bytes of each function — its selector has nowhere to put a wider offset. Everything PCI Express added lives above that: MSI-X, AER, link control. ECAM makes bus, device and function into address bits of a window named by the ACPI MCFG table, so there is no latch, no pair of accesses to keep together, and 4 KiB per function.

The window is mapped a bus at a time, on first use past offset 0xFF, and at most eight buses stay mapped; beyond that the least recently used is unmapped. That is the one lock ECAM needs. Each bus counts the accesses in flight and only an idle one is evicted, because a caller reads through a raw address and unmapping underneath it is a page fault. The unmap itself — 256 pages and one shootdown — happens after the lock is released, since waiting for other processors to acknowledge while holding a lock with interrupts masked waits on the very processors spinning for it.

The two are views of the same registers, so they must agree about the low 256 bytes. both_views_of_configuration_space_agree checks that rather than assuming it — if MCFG describes a window somewhere other than where the firmware actually put it, every extended read afterwards lands on unrelated physical memory, which is far worse than having no ECAM at all.

The test harness runs QEMU as -machine q35. The default i440FX is a 1996 chipset with no PCI Express, so it publishes no MCFG and every extended-config path would go untested.

The IOMMU inventory is read before anything trusts it. VT-d's DMAR table names remapping hardware the way MCFG names the ECAM window, so it is parsed the same way: acpi::dmar walks the structures, iommu::init maps each unit's register page and reads VER/CAP/ECAP back, and a unit that answers 0 or all-ones — what a misaddressed MMIO page reads back as — is treated as absent rather than believed, the same rule both_views_of_configuration_space_agree already applies to ECAM. No domain exists yet, and no device is any more confined than before; this stage is only the inventory the rest is built on. A machine with no DMAR table, or no unit that answers, boots on without it — every driver reaches all of physical memory exactly as it always has.

The test harness runs QEMU with -device intel-iommu,intremap=off,aw-bits=48. caching-mode=on was tried and reverted: the specification frames it as costing an extra invalidation on a not-present-to-present mapping change, but on this QEMU (11.1.0) it made every AHCI DMA transaction dramatically slower — fs::the_allocator_never_hands_out_metadata, which writes on the order of a hundred files to a small disk in well under a second, timed out at 90 seconds with caching-mode=on and nothing else different, three runs running, and passed immediately with it off. The kernel does not yet touch a single IOMMU register beyond the three read here, so the cost was QEMU's own emulation, not anything this driver does — measured, not argued, in keeping with how everything else in this file is written. Handling CAP.CM moves to whichever stage tests it deliberately, in isolation, rather than being paid on every one of 29 test-kernel boots. PANDA_IOMMU=off bisects a future regression to this device in one command.

Serial input arrives by interrupt, not by polling. It was drained from the timer handler before, which capped throughput at the tick rate and made a keystroke wait up to a full quantum to be noticed. That was not laziness: routing IRQ 4 needs an I/O APIC, and finding one needs ACPI, so it had to wait for both.

The MADT's interrupt source overrides are honoured rather than assumed away. Firmware is allowed to wire an ISA IRQ to a pin other than the one its number suggests — the timer usually arrives on pin 2, not pin 0 — and programming the obvious pin instead is the classic way to configure a redirection entry nothing is connected to, then wait forever.

Order matters at both ends. The redirection entry's high half is written before its low half, because the low half carries the mask bit and the entry must never be briefly live with no destination set — on a multiprocessor that is an interrupt delivered to whichever core happens to be APIC id zero. And the UART is told to raise the line only after something is listening: a level-triggered input asserted with the entry still masked stays asserted.

The timer handler still polls the console when routing did not happen — no ACPI, no chip, a firmware layout this does not understand. A machine whose only interface is the serial port should be slow rather than deaf.

There is no priority inheritance, and the reason is structural. Unbounded priority inversion needs a lock holder that is not running: a Low thread takes a lock, is preempted, and a High thread waits for as long as the scheduler keeps choosing something in between.

Acquiring a lock masks interrupts on the holder's processor, and every lock in the kernel is that one lock. A thread that cannot be interrupted cannot be preempted, so a holder always runs its critical section to completion and a waiter waits for that section rather than for a scheduling decision. Boosting a holder that is already running and cannot be descheduled would change nothing.

The masking used to live in a separate IrqMutex wrapper, which left every call site to pick the right type — a decision nobody should have to get right repeatedly, and the I/O APIC's register lock got it wrong. Folding it into the one lock makes the invariant hold by construction, and a_lock_holder_cannot_be_preempted checks the property the argument rests on rather than the types.

The lock is a ticket lock, so waiters are served in the order they arrived. A test-and-set spinlock has no queue: every waiter races for the same word on release, and the core whose cache already holds the line tends to win again. Under sustained contention on the heap or the frame allocator — which every core touches — that leaves a processor waiting for reasons nothing in the code explains. Each caller now takes a number and waits for it, so the longest waiter is always next. The cost is one extra atomic per acquisition, against a single contended cache line bouncing between four cores.

A TLB shootdown names one page and waits for an answer. It used to flush everything and return immediately. The old comment argued the gap was harmless because the sender had already finished its unmap — true only if nothing reuses the frame in the meantime, and freeing it is exactly what happens next.

Waiting introduces a deadlock that has to be designed out rather than hoped away: two processors can each be waiting on the other, and callers arrive with interrupts already masked, so neither would ever take the other's IPI. The request lives in a per-CPU slot, and a processor waiting for acknowledgements services its own slot inline on every pass of the wait loop. That breaks the cycle without relying on interrupt delivery at all. Requests merge rather than overwrite — two different pages become "everything", because dropping one would leave a stale translation alive with nothing left to report it.

Freeing an intermediate page table still flushes wholesale: what the processors cached is the structure, not one leaf, and a single-address invalidation does not reach it.

cpu_index takes no lock. It is called from the timer handler, from every scheduler operation and from every wake, and it used to lock a Vec and search it — a contended shared lock on the hottest path in the kernel, which is what the scheduler had just been restructured to avoid. The IPI fan-out cloned that Vec, so a page unmap could end up inside the heap allocator. Both now read a fixed table of atomics written once during boot.

System calls are preemptible. The gate is a trap gate, not an interrupt gate, so IF survives the transition. As an interrupt gate a syscall ran to completion however long it took, the calling thread's quantum meant nothing, and every future call had to stay short — a constraint a microkernel cannot keep, since the point is that calls do real work. What it demands in return is that the dispatcher tolerate preemption: every lock it touches masks interrupts for the window it holds them, the user-memory windows are bracketed by a guard that does the same, and the frame it edits lives on the calling thread's own kernel stack.

Run queues are per-CPU, and an idle core steals rather than idling. A thread goes back on the queue of the processor that last ran it, so it tends to return to a core whose caches still know about it. Left there, that would be four independent schedulers with wildly different amounts of work — a core that emptied its own queue would run its idle thread next to another core's backlog — so a core with nothing of its own takes from the busiest queue instead.

The timer tick does not take the scheduler lock. It runs on every core on every tick and asks two questions, both of which are almost always answered "no": has this slice expired, and is a sleeper due. Both now live in atomics outside the lock, so four cores no longer queue on one spinlock a hundred times a second for the privilege of subtracting one. The slice is advisory — the authoritative reset happens under the lock at the switch — so a lost race costs at most one early or late preemption, which round-robin cannot distinguish from a normal one.

Reaping used to walk the entire thread table on every context switch, looking for the rare thread that had died. A finished thread is now set aside by the processor that switched away from it, at the moment it leaves the stack.

A context switch takes one processor's lock, not everyone's. Each processor has its own lock over its ready queues and the thread it is running, and every thread is owned by exactly one processor: its scheduling state is read and changed only under that processor's lock. Ownership moves only while the thread is on no CPU — when it is woken, to the waker's processor, or when an idle processor steals it — and only under both processors' locks, taken in index order. Anything that finds a thread by id locks the owner and then checks the owner has not changed, following the thread if it has. The id-to-thread table is a separate lock that the switch path never touches, and current_id, which every system call asks, is a lock-free per-CPU read.

It was one lock before, taken by every switch on every core. Under emulation a virtual CPU spinning for it can be descheduled by the host while another holds it, and that queue was where a loaded host made the whole machine crawl.

Three priorities, and a guard against the obvious consequence. Strict priority starves: a High thread that never blocks means nothing below it runs again. Every eighth switch is therefore taken from somewhere other than the top. Three levels rather than thirty-two, because scheduling policy belongs in Ring 3; what the kernel owes is enough separation for an input daemon to preempt a compute loop.

The first version of that guard served the lowest occupied queue, which reads as reasonable and quietly skips the middle: with High and Low both busy, every guarded turn went to Low and a Normal thread never ran again. The boot thread is Normal, so the kernel hung outright the moment anything saturated the other two levels — it woke from a sleep and was simply never picked. The guard alternates between Low and Normal instead, which bounds any level's wait at two guard intervals. the_middle_priority_is_not_squeezed_out saturates the outer two deliberately and requires an ordinary thread to still get in.

The priority test spawns six threads at each level rather than one. With four cores and one thread of each, both simply get a core and the choice never happens — the queues have to be contended for the result to mean anything.

A thread can sleep, and a thread can be waited for. Both had been missing, so polling with yield_now was the only way to await anything that was not an IPC message. Sleepers sit in an unsorted list next to the earliest deadline in it, so the timer handler — which runs on every core on every tick — compares two integers in the common case and walks the list only when something is actually due.

join registers the waiter and blocks under the lock of the processor that owns the thread being waited for, and exit_current takes the waiter list under that same lock. That is what makes the race unrepresentable: a join either gets in before the thread finishes and is woken, or sees Finished and does not park at all. The list lives on the thread being waited for, so finishing is one look-up rather than a scan.

One processor keeps the clock. Every core has its own APIC timer and all of them reach the same handler, so counting the clock on each made uptime run at four times real speed on a four-core machine — and any duration measured in ticks come out short by the same factor. Invisible from inside, because everything was measured against the same wrong clock. Per-CPU interrupt counts are kept separately, which is both the honest diagnostic and what makes the property testable.

An unmap gives back the page tables it emptied, except at level 4. Removing a mapping clears one level 1 entry; the P1 that held it and the P2 above it used to stay allocated forever, so a range that is mapped and released repeatedly leaked a frame per level per region. Each unmap now walks back up and frees whatever it left empty.

Level 4 is deliberately excluded. Every entry there outside the user slot is shared by pointer with every process's cloned table, so clearing one would unmap a whole kernel region from every address space at once and hand a live table to the allocator. The user slot has its own path — AddressSpace::release frees that subtree wholesale when the process exits. Everything below level 4 is the same physical table in every space, so freeing one there is right rather than merely tolerable: the mapping really has gone everywhere.

Freeing a table means invalidating more than one address. Processors cache paging structures as well as translations, so a P2 that still remembers a P1 just handed back would walk into whatever gets allocated next — the local TLB is flushed wholesale and the other cores are told to do the same.

The kernel has to ask before it may touch user memory. With SMAP on, every supervisor read or write of a user-accessible page faults unless EFLAGS.AC is set. The handful of places that legitimately do it — copying a syscall's buffer, filling a program's image before it starts — hold a UserAccess guard, which sets AC and clears it again on drop. Everywhere else, a stray dereference of an attacker-supplied pointer now faults instead of quietly succeeding.

The guard also masks interrupts. Nothing clears AC on the way into a handler, so an interrupt landing inside the window would run the entire handler with SMAP disabled, and a context switch there would carry the relaxation into an unrelated thread. Every window is a bounded copy, so the cost is small and the alternative is a protection that lapses at moments an attacker can choose.

ring0_cannot_touch_user_memory_without_asking is the proof, and it needed a small piece of machinery to write: a kernel-mode page fault normally panics, so the test arms a one-shot flag that turns the expected fault into a thread death instead. Its partner case performs the same access through the guard and requires it to succeed, so the pair cannot both pass for a trivial reason.

The test harness runs QEMU as -cpu qemu64,+smep,+smap. The default model advertises neither, the kernel would detect them as absent and skip them, and a missing stac would read as working code.

Kernel stacks are mapped, not allocated, with an unmapped guard page beneath each. A stack on the heap has nothing below it but more heap, so overflowing it writes into another allocation and surfaces later as corruption somewhere unrelated. An unmapped page turns that into a fault on the first byte past the end — and on a kernel thread that fault escalates to a double fault, which lands on the IST stack and prints. Slots are spaced twice the stack size apart, so an overflow cannot skip the guard and land in the neighbour.

Each process has its own page tables. A new address space is a clone of the kernel's level 4 table with one slot — the 512 GiB entry covering the whole user region — replaced by a private subtree. Kernel mappings are therefore shared by pointer, so a later kernel mapping is visible everywhere at once, and only the user region diverges. That holds while no kernel mapping needs a brand-new level 4 entry after the first process exists; every kernel region has its slot populated during boot.

The page-table mapper is rebuilt from CR3 on every call rather than cached. A cached mapper always describes the boot tables, so once processes have spaces of their own it answers questions about the wrong one — silently, and only for user addresses.

one_process_cannot_read_another_address_space is the proof: a process handed an address that is mapped only in another space faults with USER_MODE and no PROTECTION_VIOLATION, meaning the page is not present rather than merely forbidden. The kernel-trespass test shows the opposite pair.

Authority narrows, never widens. A grant is intersected with what the granter already holds, and requires the GRANT right to perform at all. Naming an endpoint conveys nothing on its own.

The display's pixel format is read, not assumed. QEMU's framebuffer is 24-bit, and a client rendering 32-bit pixels into it produces an image sheared a little further right on every row. Buffers take their depth from the display so the two always agree.

The compositor never touches the hardware. It reaches the screen only through a shared buffer handle and learns what to draw only through IPC. The scanout buffer is a singleton — two handles to one screen would let two compositors fight over the same pixels with neither aware of the other — and it refuses to be destroyed, because returning MMIO to the frame allocator would be a catastrophic double free.

The compositor tests assert on real pixels. They read the display's memory back and check the colour at the requested coordinates, and that just past the surface's edge nothing was touched. A blit that runs one row long, or lands at the wrong offset, passes every structural check and fails these.

Any lock the timer handler can reach masks interrupts. The console, the heap, the scheduler, the page tables and the frame allocator all do. The rule is not about multi-core exclusion — the spin still provides that — it is about a CPU deadlocking against itself: a tick that lands on the very core holding a lock, and then needs it, spins forever, because the holder cannot run again to release it.

Which locks qualify changes as the kernel grows, and that is the trap. The frame allocator did not qualify until kernel stacks moved off the heap: after that, a tick could schedule, scheduling could drop a finished thread, and dropping one unmaps a stack and returns its frames. It hung about one run in thirty, in whichever test allocated the most physical memory.

The console disables interrupts while it holds its lock. Without this the kernel deadlocks the first time a handler prints: it spins on a lock held by the code it interrupted, which cannot run again to release it. The window is small, which only means the hang would be intermittent.

Random bytes come from the processor, stretched by a hash. RDSEED, or RDRAND where there is no RDSEED, seeds a SHA-256 construction and is mixed in again on every request; each request ends by replacing the key with a hash of itself, so the state a request leaves behind cannot recompute what it handed out. Salts, TCP initial sequence numbers, DHCP transaction ids, DNS query ids and the port lookups go out from all come from it — Ring 3 by a system call. A processor with neither instruction gets a seed hashed from timing jitter and a line at boot saying it is guessable; xtask asks QEMU for both so the hardware path is the one tested.

A panic stops the world, then reports without waiting for anything. Other processors are halted with an NMI, not an ordinary IPI: the processor most likely to be in the way is one spinning on a lock with interrupts masked, and only a non-maskable interrupt reaches it. After that every lock on the path is tried rather than taken — the console, the scheduler, the page tables, the disk — because the holder may be a processor that was just stopped, or the code that panicked. The backtrace follows frame pointers (forced on in kernel/.cargo/config.toml) and checks each frame is mapped and kernel-only before reading it; a panic inside a system call has the user's stack at the top of the chain, and SMAP would fault on it. Each return address is named from the kernel's own symbol table, read out of the ELF the bootloader leaves in memory: a linear walk, with no allocation and no locks. Symbols use legacy mangling so the names demangle in a few lines instead of a v0 demangler in the panic path.

The report goes to a GPT partition whose type GUID is the ASCII KernelPandaCrash, text first and header last, so an interrupted save leaves the old state rather than a header describing half-written text. If the disk lock is held the record is not written at all: going ahead would overwrite the bounce buffer a command already in flight is about to be DMA'd from. The next boot prints the record and clears it.

xtask attaches a 2 MiB disk holding that partition to every launch. cargo xtask run keeps it in target/images/crash.img between boots, so a panic is there to read the next time. Test kernels get a fresh one, and a test kernel may exit asking to be booted again on the same disk: the crash test panics on its first boot, and on its second checks the record it finds names the function that panicked.

The PS/2 driver is a Ring 3 process, and is actually contained by it. The disk driver cannot be, because a DMA engine ignores page tables; a keyboard controller does no DMA, so a process that can reach exactly ports 0x60 and 0x64 and hear exactly lines 1 and 12 can do nothing worse than lie about keys. Those grants are given by whoever spawns the driver, never requested.

Ports are reached through a system call rather than the TSS I/O bitmap. The bitmap would let the driver use in and out directly, but it is per processor, and every switch between two processes with different grants would have to rewrite it on whichever processor the switch landed. The call checks the grant in one place, for the price of a trap per byte — nothing, at the rate a keyboard produces them.

An interrupt becomes a message stamped with a sender id no thread can have. The kernel never touches the device; it notifies the endpoint the driver bound and acknowledges the APIC. A notification dropped on a full queue loses nothing: a full queue means notifications are already waiting, and each one sends the driver to drain the controller completely.

The tests put bytes in through the controller itself, with its "write output buffer" commands, so every case goes through the real interrupt line, the grant checks and the decoder.

The compositor believes input from one thread. Every client holds SEND on the compositor's endpoint, so a compositor that acted on any key event would let one client type into another's window. Key and pointer events count only from the kernel, or from the thread the kernel has named as the input daemon — and naming it is a message only the kernel can send.

The network stack is a process; the kernel moves frames. Everything that parses bytes from the network — ARP, IPv4 headers, ICMP, UDP, TCP, DHCP, DNS — runs in a Ring 3 daemon, so a malformed packet that finds a bug there kills an unprivileged process. The kernel's part is four system calls: the card's MAC, send a frame, take a frame, and have arrivals announced on an endpoint. Only the one thread the kernel designates may make them, the same way only the designated display server may have the screen.

The virtio-net driver is in the kernel for the same reason the disk driver is: the card is a DMA engine, and without an IOMMU a Ring 3 driver holding one is isolated in appearance only. It speaks virtio's legacy interface, a block of I/O-port registers, because QEMU offers it and it is a fraction of the modern one's machinery.

Arrivals are signalled by MSI-X. A device's legacy interrupt pin reaches the I/O APIC through wiring only the firmware's AML describes, and there is no AML interpreter; an MSI-X entry is an address and a value the device writes, and the address is simply the Local APIC's. The daemon's clients reach it over IPC, with datagrams in buffers they share, since a message carries four words.

The tests run against QEMU's user-mode network: the gateway answers ARP and pings, and its built-in TFTP server hands out a file xtask writes, so the UDP path is checked byte for byte with nothing outside the emulator. The first driver offered its buffers in the descriptor table rather than the ring after it; QEMU's trace of a queue notified and never popped is what found it.

TCP holds one segment in flight each way, and waits on the client for both. A send is up to one segment from a buffer the client shared, and is answered once the peer acknowledges it; the next send waits for that. Arriving data goes into the client's buffer and the window closes until the client says it has read it, so nothing arrives that there is no room for. What is not acknowledged is sent again from the client's buffer — the daemon keeps no copy, so a connection costs it a few dozen bytes — after a second, then two, and so on to eight, and after six tries the connection is reset and reported as timed out. Only the thread that opened or accepted a connection may use it; the daemon checks the sender the kernel stamped on each request.

Resending needs time, and a Ring 3 process had no sense of it. A timer is now a system call: after a delay the kernel sends a message to an endpoint the caller can receive on, so a daemon waiting for a packet and waiting for a deadline is waiting in one place. It runs off the tick that wakes sleeping threads.

Started without an address, the daemon asks DHCP for one. Clients asking for the configuration before it arrives are answered when it does, so nothing has to guess when the network is up. It resolves names with the DNS server DHCP offered, or one it was given at start-up; a lookup is one question for an IPv4 or an IPv6 address, asked three times two seconds apart before it is reported as timed out.

IPv6 and IPv4 share everything above the network layer. An address is 16 bytes throughout, an IPv4 one held IPv4-mapped, so UDP, TCP, DNS and the neighbour cache are one piece of code for both, and only the header, the checksum's pseudo-header and the way a neighbour is found differ. The daemon makes its link-local address from the MAC, asks for a router, and makes a global address from the prefix advertised, taking a DNS server from the advertisement if it names one. An IPv6 address does not fit in a message word, so a client names one by putting it at the start of the buffer it shares, and is told of one the same way.

A client's buffers are mapped into the daemon when it first names them, and a full table used to be the end of new clients: nothing could take a buffer back out of a process. There is now a system call that gives a buffer up, unmapping it and dropping the access sharing granted; a buffer whose owner has exited is freed when the last holder does. The daemon gives up the buffers nothing refers to any more when it needs room.

The tests reach services xtask runs on the host for as long as a test kernel does: a server the guest connects out to, which greets, answers a line and closes; a client that keeps connecting in, through a port QEMU forwards, until the guest answers it; and a DNS server that knows one name. QEMU's own DNS server forwards to whatever the host uses, and a test that passes only on a machine with a working resolver is not a test of this code. The servers listen on both loopbacks, since QEMU carries the guest's IPv4 to 127.0.0.1 and its IPv6 to ::1. PANDA_PCAP=<file> makes xtask record every frame, for Wireshark.

A disk behind SATA, NVMe or virtio looks the same from above. The block layer asks for sectors by number, and each driver answers through the same interface, including the write the panic path needs. The three share a shape on purpose. A disk has request slots — eight, or as many as the controller has — each with a bounce buffer of two pages, which is also as far as NVMe's second PRP reaches without a list. A request takes a slot, hands the device its command, and waits for the answer, so requests on different slots are with the device together.

How it waits depends on where it is. Where the caller may sleep, it sleeps, and the device's interrupt — MSI-X for NVMe and virtio-blk, MSI for AHCI — wakes it. Where it may not, it polls the device itself: during boot, on the panic path, and with a spinning lock held. Every spinning lock masks interrupts, so "interrupts are on" is the same test as "no spinning lock is held", and it is the one the driver makes. A caller waiting for a slot sleeps or polls on the same rule.

That rule put the filesystem in the way: it held a spinning lock across every disk I/O, so every file read polled. Its lock now sleeps. SleepMutex parks a caller that finds it taken, and falls back to spinning where parking is not allowed, so it is usable anywhere a Mutex is, only slower there.

AHCI queues through NCQ, where both the controller and the drive offer it: a slot is a tag, and the drive answers tags in whatever order it likes. Flush and identify are not queued commands, so they take every slot and run alone, and a caller waiting to do that holds back new single-slot requests so a steady stream cannot keep it waiting forever. A drive without NCQ gets one command at a time.

NVMe is two pairs of rings: the admin queue, through which the driver learns the namespace's size and creates the other, and one I/O queue. virtio-blk is one virtqueue carrying three-descriptor requests — header, data, one byte of outcome — on the same virtqueue code the network driver uses.

The test harness attaches a disk of each kind at a distinct size, and tests/storage.rs runs the same cases on both new ones, down to formatting a filesystem and reading a file back after a remount. On each disk four threads write and read back patterns of their own at once, and the case checks that the data survived, that requests really were in flight together, and that answers really came by interrupt. The first NVMe driver put a submission queue's completion-queue id where its flags go; QEMU's trace, reading "invalid cqid=0", found it faster than the specification did.

A ring is a queue two processes share, and a busy one never enters the kernel. An endpoint moves each message through two system calls, and a context switch whenever the receiver was waiting. A ring is an endpoint with pages attached: the sender writes a 64-byte slot and advances its tail, the receiver reads it and advances its head, and both do it in memory they map. Fifty thousand messages through one cost six system calls in the test — two to map, two to report, two to exit.

Capabilities carry over unchanged, because the ring is an endpoint: SEND lets a thread map the sending side, RECEIVE the receiving side, and GRANT hands either on. Each side is claimed by one thread, since two writers on one unlocked counter would corrupt it.

The pages are split so each is writable by exactly one party. The sender's page holds the tail and the slots; the receiver's holds the head; the slot count is on a page neither can write, because both size their reads by it and either could otherwise point the other past the end. A receiver that writes a message, or a sender that moves the receiver's head, takes a page fault — both tested. A side can still lie in its own counter; the other then sees a broken channel, and nothing outside the ring's own memory.

A side that finds nothing to do spins briefly, then sets a flag and sleeps; the other side, having published its counter, checks the flag and wakes it through the kernel. Both writes are exchanges, which on x86 are locked instructions and so committed before the read that follows — at least one side always sees the other, and a receiver cannot park beside a message. The spin matters: without it a receiver that kept up slept after every message, and the first run spent 38,000 system calls on 50,000 of them. The wake path is tested by making each side in turn wait on a deliberately slow partner.

Capabilities decide what a process can reach; users decide who it is. Every thread runs as a user, a number. It starts as its spawner's, and only kernel code says otherwise — sched::spawn_as sets it before the thread can be picked up by any processor, so there is no moment in which a thread runs as someone it is not. The one system call that changes a thread's user is logging in, and it takes the account's password: a process cannot promote itself, or pass itself off as another without knowing what that user knows.

Accounts live in /users on the root filesystem, a line each: name, user, a salt, and PBKDF2-HMAC-SHA-256 of the password at 10,000 rounds. The file is owned by a user no thread runs as and no account may be, with every permission bit clear, so no process can read the hashes or open the file up — the system's included; the kernel reads it on its own account. A login costs the same rounds whether or not the name exists, a wrong name and a wrong password are the same refusal, and a refusal waits half a second before returning. The shell's login reads the password without echoing it. The boot demo formats a blank disk if it found no filesystem, adds panda with password panda, and logs in.

Every message carries its sender's user next to its thread, both stamped by the kernel on the way through, so a server can decide by user without trusting what a client wrote.

Every file and directory has an owner and four bits: read and write, for the owner and for everyone else. Reading or writing a file needs that bit on it; adding or removing a name needs write on the directory; looking a name up needs read on every directory along the path, so a private directory hides what is in it even from someone who knows the name. Only an owner changes the bits, and nothing changes an owner.

There is no superuser. User 0 is the system's own user — it owns the root, which everyone may read and only it may write — and it is refused a private file like anyone else, which a test checks. What the system can do beyond its own files comes from capabilities it holds, the same as every other process; one number that bypasses every check is the kind of ambient authority the rest of this kernel exists to avoid. Kernel code acting on its own account is not a user at all and is not checked.

The owner and bits live in bytes the inode format had left zero, which made this format version 2; a version 1 disk is refused rather than read as everything owned by the system and closed to all.

Known limits

  • Only ever run under QEMU. Firmware variance in ACPI layout and AP start-up timing is exactly where this class of code breaks.
  • The ticket lock is fair but not priority-aware: a High thread queues behind a Low one that asked first. That is bounded by a critical section rather than by a scheduling decision — see the note on priority inversion below — so it is a fairness cost, not a liveness one.
  • One user program is still hand-written assembly: the W^X test, which plants two bytes of machine code on its own stack and jumps to them. That is not something Rust will express, and it is the right tool for that one job.
  • The keyboard layout is US, caps lock is not tracked, and the Pause key's E1 sequence is not decoded. An interrupt line stays routed after its driver exits; every ISA line is edge-triggered, so the cost is one ignored interrupt per event, not a storm.
  • Permissions are read and write for an owner and for everyone else: no groups, no access lists, and no sticky directories, so whoever may write a directory may remove anything in it.
  • A ring has exactly one sender and one receiver, fixed 64-byte slots, and at most 4,096 of them. A side spins for 2,000 attempts before sleeping, which is a guess tuned under emulation rather than a measurement on hardware.
  • Every transfer is copied through a slot's bounce buffer rather than built as a scatter-gather list over the caller's memory. A request waiting on an interrupt that never comes waits forever: only polled waits time out. An AHCI error fails every queued command on the port and restarts the port, without resetting the drive. NVMe namespaces must use 512-byte blocks; one formatted with 4 KiB blocks is refused rather than misaddressed. virtio devices are driven through the legacy interface only.
  • The network card has one transmit buffer, so frames go out one at a time. The daemon holds two frames while an address resolves and drops the rest, keeps one datagram per bound port, does not reassemble fragments, and checks UDP checksums only on IPv6, where they are mandatory.
  • IPv6 follows no extension headers and does no duplicate address detection. A prefix is taken from the first advertisement and kept, whatever its lifetime. Accepting a connection over IPv6 is not tested: QEMU forwards host ports to the guest's IPv4 address only.
  • Giving a buffer up does not return the address range it was mapped at; a process's shared-mapping area is handed out once.
  • A DHCP lease is never renewed, and the offer's lease time is ignored. DNS lookups go over UDP to one server, with no cache, and answers are trusted from whichever server was asked, by port and query id alone.
  • TCP drops out-of-order segments and waits for them to be sent again, ignores the peer's window and options, and has no congestion control and no TIME-WAIT. Resending after a timeout is not tested: QEMU's network loses nothing, and a connection QEMU cannot complete to the host is never refused, only left to time out, which takes longer than a test kernel may run.
  • Accounts are added only by kernel code; there is no system call to add one or to change a password, and adding two accounts at once can lose one.
  • A crash record holds the first 4 KiB of a report. Backtraces name functions, not lines: file and line need the DWARF data, which is far larger than the symbol table and slower to search.

Not built yet

The gaps that matter, so nobody has to discover them by trying:

  • A libc or a toolchain for third-party software. Programs are built in this repository's userland workspace against its own syscall wrappers.
  • An IOMMU. Without one, a Ring 3 driver handed a DMA-capable device is not isolated from the rest of physical memory, which is why the disk driver is in the kernel.

About

A bare-metal microkernel written from scratch in no_std Rust, targeting x86_64-unknown-none

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages