Repository navigation
Measurements: --state-in-memory beats the C backend on every title and platform tested; without it the LLVM backend loses to C #29
Description
Activity
Correction. One claim in the original post is wrong, and the reason turns out to be more useful than the claim.
I re-ran everything on validated scene sets — 6 scenes per title instead of 2–3 — after scanning every available savestate for boot failures, zero frame progress, and SMC hash mismatches. 38 of 105 states (36%) were rejected, including one (
course-peach-beach.sav) that does not boot at all and had been used in an earlier run of mine.What changed
title scenes C speed range llvmllvm --state-in-memoryLuigi's Mansion (GLME01) 6 0.86–8.04 1.194 1.468 Skyward Sword (SOUE01) 6 1.23–3.51 1.259 1.578 Mario Kart DD (GM4E01) 6 0.61–2.05 0.882 1.220 Pokémon Colosseum (GC6E01) 6 0.68–1.30 0.877 1.117 The headline conclusion is unchanged and slightly stronger:
--state-in-memorywins on all four titles, now +11.7% to +57.8% rather than +4.0% to +47.2%.But the original post said plain
llvmloses to C on three of four titles. That is wrong — it is two of four. Skyward Sword flipped sign, 0.896 → 1.259. My earlier figure came from two scenes; with six it reverses.Why it flipped: the ratio depends on scene weight
Pooling all 48 measurements and splitting by how fast the C backend runs each scene (low C speed = heavy scene):
llvm heavy scenes 0.960 light scenes 1.147 r=0.73 llvm --state-in-memory heavy scenes 1.227 light scenes 1.464 r=0.68Per title:
SOUE01 llvm heavy 1.084 light 1.435 GLME01 llvm heavy 1.085 light 1.303 GM4E01 llvm heavy 0.905 light 0.859 <- flat GC6E01 llvm heavy 0.881 light 0.874 <- flatLLVM gains on light scenes and gives ground on heavy ones, and the effect is concentrated in the two titles whose scene sets span a wide load range. So neither of my Skyward Sword numbers is "the" answer: the old one sampled two heavy scenes, the new one has four near-identical light scenes (C ≈ 3.4–3.5) dominating the mean.
A per-title mean is only meaningful alongside the scene weight distribution it was drawn from. I have quoted such means twice now and they moved both times.
What is solid
--state-in-memorybeats C on every title and every individual scene measured — 24 of 24 scene-level comparisons.- Colosseum is flat across scene weight (0.881 heavy / 0.874 light), so its numbers do not depend on scene choice. It is also the tightest:
--state-in-memory1.1165, SE 0.0135, n=6. - Measured noise floor is unchanged at +/-1.7% (13 null comparisons, same module as both control and arm).
What I would not claim
- Any single per-title ratio for LLVM without
--state-in-memory. It is load-dependent, and the sign can change with scene selection. - Anything about aarch64 or macOS. Still x86-64 Windows only.
Apple Silicon data, which the original post said it lacked.
Same DOL (
md5 c55df119fad43d7d4ee976c2242211cb, 3,779,808 bytes) and the same six validated Colosseum scenes as the x86-64 numbers, so the two platforms are directly comparable rather than approximately so.Host: MacBook M5, 10 cores, Apple clang 17, LLVM 20 (Homebrew
llvm@20, matching the 20.1.8 used on Windows). Metal backend. 3 reps per cell, arms interleaved per scene so thermal drift hits all three equally.scene C llvmllvm --state-in-memoryllvm/C sim/C clean-demo 1046 708 1188 0.677 1.135 clean-field2 1298 992 1315 0.764 1.013 clean-mid 1902 1596 2004 0.839 1.054 play-Colo-Fresh-002 1661 1404 1787 0.845 1.076 play-Colo-Fresh-009 1221 949 1233 0.777 1.010 play-Colo-Fresh-011 969 855 1080 0.882 1.115 llvm/C mean 0.7976 SE 0.0301 95% CI 0.737-0.858 SLOWER than C sim/C mean 1.0673 SE 0.0210 95% CI 1.025-1.109 FASTER than C within-cell repeatability: median 1.014x, worst 1.090xCross-platform agreement
x86-64 Windows arm64 macOS llvm / C 0.8775 0.7976 llvm --state-in-memory / C 1.1165 1.0673Same direction and similar magnitude on both.
--state-in-memoryturns the LLVM backend from a regression against C into a win on AArch64 as well, which is the part I was least sure would carry over: AArch64 has 31 GPRs against x86-64's 16, so if the benefit were purely about register pressure it could plausibly have vanished here. It does not.Two things worth flagging for anyone reproducing on macOS
The LLVM backend cannot build modules on macOS without #26. Every chunk aborts:
LLVM ERROR: Global variable 'fix_pair_nan_f64' has an invalid section specifier '.text.unlikely.fix_pair_nan_f64': mach-o section specifier requires a segment and section separated by a comma.27 of these, then failure. With #26 applied: clean build, zero errors, 145.9 MB dylib. I had assumed that PR was cosmetic; it is not — it gates the whole backend on Darwin.
Partial output from a failed run is reused. After applying #26 the errors persisted until I deleted the 665 MB of partial module output the failed attempts had left behind. Worth clearing the output directory rather than assuming a fix did not work.
Caveat
The macOS figures are frames-advanced over a fixed 20s window (3 reps, interleaved). The x86-64 figures came from the interleaved forward/reversed pair harness with medians. Both control for drift, but they are not the identical statistic, so I would treat the cross-platform agreement as directional rather than as a like-for-like ratio comparison.
- changed the title
[-]Measurements: --state-in-memory beats the C backend on all four titles tested; without it the LLVM backend loses to C on three[/-][+]Measurements: --state-in-memory beats the C backend on every title and platform tested; without it the LLVM backend loses to C[/+]on Sep 1, 2026 Ran the same comparison on Linux and got a materially different answer for MKDD, which I think sharpens the conclusion rather than undermining it: the "without
--state-in-memory, LLVM loses to C" half appears to be host-dependent.Numbers
MKDD
race.sav, the interleaved harness ported to Linux (6 forward + 6 reversed pairs, medians decide), headless, idle box:host --backend llvm--backend llvm+ SIMthis issue (MKDD, 3 scenes) 0.906 1.145 Linux, Ryzen 9 5900X 1.2002 1.5621 Both Linux comparisons were 12/12 pairs favouring the arm, with tight deltas: +17.3%..+22.1% for
llvm, +49.4%..+60.9% for SIM. Absolute: C 32.97 → LLVM 39.57 → SIM 52.03 fps.So on this host plain LLVM does not lose to C — it wins by 20%, and SIM adds another 30 points on top.
Why the two disagree, most likely
The hosts are two generations and 2× L3 apart, and the thing SIM changes is footprint:
arm module C 75 MB LLVM 404 MB LLVM+SIM 166 MB perf statacross the three arms on the 5900X, normalised per frame (the arms run at different fps, so per-instruction normalisation flatters the slow one):per frame C LLVM LLVM+SIM cycles 175.3 M 154.2 M 122.4 M cache misses (L3 proxy) 4.27 M 4.01 M 2.41 M L1i misses 469 K 310 K 235 K iTLB misses 112 K 136 K 68 K branch misses 820 K 360 K 356 K IPC 1.58 1.65 1.78 Two separable effects fall out:
- LLVM's gain is codegen — branch misses drop 56% against C. But it costs memory-system pressure: iTLB misses rise 22%, which is what a 404 MB code module does to a TLB. LLC barely moves.
- SIM's gain is footprint. It keeps the branch improvement (−57%, flat against plain LLVM) and additionally cuts LLC misses 40%, iTLB 50%, L1i 24% versus plain LLVM, with branch misses unchanged. That residual is a pure working-set effect and it tracks 404 MB → 166 MB.
That predicts exactly the split observed: on a part with enough L3 to absorb plain LLVM's footprint, the codegen win is cancelled by the footprint cost and
llvmlands at or below C — which is what the table in this issue shows. On a cache-starved part the footprint cost dominates first, so both arms win and SIM wins hugely.If that reading is right, the practical consequence cuts the way this issue argues, only harder: SIM is worth more the less cache the host has, so benchmarking it exclusively on a large-L3 part systematically understates it.
Caveats, so these are read at the right precision
- The counters were multiplexed at ~35% enable time (more events than PMU slots), so they are scaled estimates. The 40–50% deltas are far outside multiplexing error; I would not quote single-digit differences from that run.
LLC-*events are unsupported on this Zen 3 part, socache-missesis the last-level proxy.- Linux ran LLVM 20.1.2 against 20.1.8 elsewhere — a patch-level codegen difference.
- Linux was headless under Xvfb; a windowed run carries present/GX cost that a headless one does not, which shifts how CPU-bound the workload is. That is a real confound against the numbers in this issue and I would not rule it out as part of the gap.
- Different scene selection: one savestate here versus three.
Host: Ubuntu, Ryzen 9 5900X (12C/24T, 64 MB L3 as 2×32 MB), RTX 3090,
performancegovernor, headless, nothing else running.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsTodo
Not a bug report — a measurement writeup, since
--state-in-memoryis opt-in and off by default and the numbers suggest that default is worth revisiting.Everything below uses code already in
main; nothing here needs a patch.Result
Control is the C backend. Arms are LLVM variants from the same recompiler binary against the same DOL; only backend and state mode differ. Ratios are arm/control, so >1.0 is faster than C.
--backend llvm--backend llvm --state-in-memoryTwo things stand out:
--state-in-memory, the LLVM backend is slower than C on three of the four titles (0.83–0.91).Colosseum's figure is the most sampled: 8 scenes spanning 12.8–37.2 fps, and the
llvmarm landed in 0.773–0.876 on every one of them.Method
BackgroundInput = False) and one emulator at a time — both mattered, see belowNoise floor was measured, not assumed
The same C module was supplied as both control and arm, so the true ratio is exactly 1.0000:
Single runs therefore carry a rare ~15% excursion. Every number above is a mean over repeated runs. Several earlier single-run figures in this work did not survive repetition and were discarded.
Caveats worth stating
SMC: chunk [0x80005300,0x80005500) hash mismatch. They boot and render normally while that chunk silently falls back to the interpreter. Anyone reproducing this should check the runtime log forhash mismatchand confirmframe_countactually advances.Why post it
--state-in-memorycurrently reads as an experimental toggle. On this hardware and these titles it is the difference between the LLVM backend being a regression against C and being a win. If that reproduces elsewhere, the default may be worth changing — and if it does not reproduce on other targets, that is worth knowing too.Happy to run additional titles or scenes if useful.