Skip to content

Measurements: --state-in-memory beats the C backend on every title and platform tested; without it the LLVM backend loses to C #29

Description

@dougchansan

Not a bug report — a measurement writeup, since --state-in-memory is opt-in and off by default and the numbers suggest that default is worth revisiting.

Everything below uses code already in main; nothing here needs a patch.

Result

Control is the C backend. Arms are LLVM variants from the same recompiler binary against the same DOL; only backend and state mode differ. Ratios are arm/control, so >1.0 is faster than C.

title --backend llvm --backend llvm --state-in-memory scenes
Luigi's Mansion (GLME01) 1.185 1.472 2
Skyward Sword (SOUE01) 0.896 1.194 2
Mario Kart: Double Dash (GM4E01) 0.906 1.145 3
Pokémon Colosseum (GC6E01) 0.827 1.040 8

Two things stand out:

  1. Without --state-in-memory, the LLVM backend is slower than C on three of the four titles (0.83–0.91).
  2. With it, LLVM beats C on all four. The swing on Colosseum is 0.827 → 1.040 across 8 independent scenes.

Colosseum's figure is the most sampled: 8 scenes spanning 12.8–37.2 fps, and the llvm arm landed in 0.773–0.876 on every one of them.

Method

  • runtime is ModernGekko with the static-recomp core; frames advanced in a fixed window, uncapped
  • each scene is a savestate; interleaved forward and reversed control/arm pairs, medians compared
  • host input pinned (no device bound, BackgroundInput = False) and one emulator at a time — both mattered, see below
  • Colosseum's MSR region-leader fix (Make the instruction after an MSR write a region leader #28) is present in all arms, so it is not what is being measured here

Noise floor was measured, not assumed

The same C module was supplied as both control and arm, so the true ratio is exactly 1.0000:

n=13 null comparisons across 4 titles / 6 scenes
12 of 13 within +/-1.72%
 1 of 13 at +14.97%  (one scene; 3 repeats of it then returned 1.0164 / 0.9972 / 1.0019)

Single runs therefore carry a rare ~15% excursion. Every number above is a mean over repeated runs. Several earlier single-run figures in this work did not survive repetition and were discarded.

Caveats worth stating

  • x86-64 Windows only. No aarch64 or macOS numbers yet, and state-in-memory could plausibly behave differently where register pressure differs.
  • Scene counts are uneven (8 for Colosseum, 2–3 elsewhere).
  • Savestate hygiene matters more than expected: of 38 Colosseum states available here, 17 carried a widescreen patch in their RAM image and reported SMC: chunk [0x80005300,0x80005500) hash mismatch. They boot and render normally while that chunk silently falls back to the interpreter. Anyone reproducing this should check the runtime log for hash mismatch and confirm frame_count actually advances.

Why post it

--state-in-memory currently reads as an experimental toggle. On this hardware and these titles it is the difference between the LLVM backend being a regression against C and being a win. If that reproduces elsewhere, the default may be worth changing — and if it does not reproduce on other targets, that is worth knowing too.

Happy to run additional titles or scenes if useful.

Activity

  1. dougchansan commented on Aug 31, 2026

    @dougchansan
    ContributorAuthor

    Correction. One claim in the original post is wrong, and the reason turns out to be more useful than the claim.

    I re-ran everything on validated scene sets — 6 scenes per title instead of 2–3 — after scanning every available savestate for boot failures, zero frame progress, and SMC hash mismatches. 38 of 105 states (36%) were rejected, including one (course-peach-beach.sav) that does not boot at all and had been used in an earlier run of mine.

    What changed

    title scenes C speed range llvm llvm --state-in-memory
    Luigi's Mansion (GLME01) 6 0.86–8.04 1.194 1.468
    Skyward Sword (SOUE01) 6 1.23–3.51 1.259 1.578
    Mario Kart DD (GM4E01) 6 0.61–2.05 0.882 1.220
    Pokémon Colosseum (GC6E01) 6 0.68–1.30 0.877 1.117

    The headline conclusion is unchanged and slightly stronger: --state-in-memory wins on all four titles, now +11.7% to +57.8% rather than +4.0% to +47.2%.

    But the original post said plain llvm loses to C on three of four titles. That is wrong — it is two of four. Skyward Sword flipped sign, 0.896 → 1.259. My earlier figure came from two scenes; with six it reverses.

    Why it flipped: the ratio depends on scene weight

    Pooling all 48 measurements and splitting by how fast the C backend runs each scene (low C speed = heavy scene):

    llvm                    heavy scenes 0.960   light scenes 1.147   r=0.73
    llvm --state-in-memory  heavy scenes 1.227   light scenes 1.464   r=0.68
    

    Per title:

    SOUE01  llvm  heavy 1.084  light 1.435
    GLME01  llvm  heavy 1.085  light 1.303
    GM4E01  llvm  heavy 0.905  light 0.859    <- flat
    GC6E01  llvm  heavy 0.881  light 0.874    <- flat
    

    LLVM gains on light scenes and gives ground on heavy ones, and the effect is concentrated in the two titles whose scene sets span a wide load range. So neither of my Skyward Sword numbers is "the" answer: the old one sampled two heavy scenes, the new one has four near-identical light scenes (C ≈ 3.4–3.5) dominating the mean.

    A per-title mean is only meaningful alongside the scene weight distribution it was drawn from. I have quoted such means twice now and they moved both times.

    What is solid

    • --state-in-memory beats C on every title and every individual scene measured — 24 of 24 scene-level comparisons.
    • Colosseum is flat across scene weight (0.881 heavy / 0.874 light), so its numbers do not depend on scene choice. It is also the tightest: --state-in-memory 1.1165, SE 0.0135, n=6.
    • Measured noise floor is unchanged at +/-1.7% (13 null comparisons, same module as both control and arm).

    What I would not claim

    • Any single per-title ratio for LLVM without --state-in-memory. It is load-dependent, and the sign can change with scene selection.
    • Anything about aarch64 or macOS. Still x86-64 Windows only.
  2. dougchansan commented on Sep 1, 2026

    @dougchansan
    ContributorAuthor

    Apple Silicon data, which the original post said it lacked.

    Same DOL (md5 c55df119fad43d7d4ee976c2242211cb, 3,779,808 bytes) and the same six validated Colosseum scenes as the x86-64 numbers, so the two platforms are directly comparable rather than approximately so.

    Host: MacBook M5, 10 cores, Apple clang 17, LLVM 20 (Homebrew llvm@20, matching the 20.1.8 used on Windows). Metal backend. 3 reps per cell, arms interleaved per scene so thermal drift hits all three equally.

    scene C llvm llvm --state-in-memory llvm/C sim/C
    clean-demo 1046 708 1188 0.677 1.135
    clean-field2 1298 992 1315 0.764 1.013
    clean-mid 1902 1596 2004 0.839 1.054
    play-Colo-Fresh-002 1661 1404 1787 0.845 1.076
    play-Colo-Fresh-009 1221 949 1233 0.777 1.010
    play-Colo-Fresh-011 969 855 1080 0.882 1.115
    llvm/C   mean 0.7976  SE 0.0301  95% CI 0.737-0.858   SLOWER than C
    sim/C    mean 1.0673  SE 0.0210  95% CI 1.025-1.109   FASTER than C
    within-cell repeatability: median 1.014x, worst 1.090x
    

    Cross-platform agreement

                      x86-64 Windows      arm64 macOS
    llvm / C              0.8775            0.7976
    llvm --state-in-memory / C
                          1.1165            1.0673
    

    Same direction and similar magnitude on both. --state-in-memory turns the LLVM backend from a regression against C into a win on AArch64 as well, which is the part I was least sure would carry over: AArch64 has 31 GPRs against x86-64's 16, so if the benefit were purely about register pressure it could plausibly have vanished here. It does not.

    Two things worth flagging for anyone reproducing on macOS

    The LLVM backend cannot build modules on macOS without #26. Every chunk aborts:

    LLVM ERROR: Global variable 'fix_pair_nan_f64' has an invalid section specifier
    '.text.unlikely.fix_pair_nan_f64': mach-o section specifier requires a segment
    and section separated by a comma.
    

    27 of these, then failure. With #26 applied: clean build, zero errors, 145.9 MB dylib. I had assumed that PR was cosmetic; it is not — it gates the whole backend on Darwin.

    Partial output from a failed run is reused. After applying #26 the errors persisted until I deleted the 665 MB of partial module output the failed attempts had left behind. Worth clearing the output directory rather than assuming a fix did not work.

    Caveat

    The macOS figures are frames-advanced over a fixed 20s window (3 reps, interleaved). The x86-64 figures came from the interleaved forward/reversed pair harness with medians. Both control for drift, but they are not the identical statistic, so I would treat the cross-platform agreement as directional rather than as a like-for-like ratio comparison.

  3. changed the title [-]Measurements: --state-in-memory beats the C backend on all four titles tested; without it the LLVM backend loses to C on three[/-] [+]Measurements: --state-in-memory beats the C backend on every title and platform tested; without it the LLVM backend loses to C[/+] on Sep 1, 2026
  4. dougchansan commented on Sep 17, 2026

    @dougchansan
    ContributorAuthor

    Ran the same comparison on Linux and got a materially different answer for MKDD, which I think sharpens the conclusion rather than undermining it: the "without --state-in-memory, LLVM loses to C" half appears to be host-dependent.

    Numbers

    MKDD race.sav, the interleaved harness ported to Linux (6 forward + 6 reversed pairs, medians decide), headless, idle box:

    host --backend llvm --backend llvm + SIM
    this issue (MKDD, 3 scenes) 0.906 1.145
    Linux, Ryzen 9 5900X 1.2002 1.5621

    Both Linux comparisons were 12/12 pairs favouring the arm, with tight deltas: +17.3%..+22.1% for llvm, +49.4%..+60.9% for SIM. Absolute: C 32.97 → LLVM 39.57 → SIM 52.03 fps.

    So on this host plain LLVM does not lose to C — it wins by 20%, and SIM adds another 30 points on top.

    Why the two disagree, most likely

    The hosts are two generations and 2× L3 apart, and the thing SIM changes is footprint:

    arm module
    C 75 MB
    LLVM 404 MB
    LLVM+SIM 166 MB

    perf stat across the three arms on the 5900X, normalised per frame (the arms run at different fps, so per-instruction normalisation flatters the slow one):

    per frame C LLVM LLVM+SIM
    cycles 175.3 M 154.2 M 122.4 M
    cache misses (L3 proxy) 4.27 M 4.01 M 2.41 M
    L1i misses 469 K 310 K 235 K
    iTLB misses 112 K 136 K 68 K
    branch misses 820 K 360 K 356 K
    IPC 1.58 1.65 1.78

    Two separable effects fall out:

    • LLVM's gain is codegen — branch misses drop 56% against C. But it costs memory-system pressure: iTLB misses rise 22%, which is what a 404 MB code module does to a TLB. LLC barely moves.
    • SIM's gain is footprint. It keeps the branch improvement (−57%, flat against plain LLVM) and additionally cuts LLC misses 40%, iTLB 50%, L1i 24% versus plain LLVM, with branch misses unchanged. That residual is a pure working-set effect and it tracks 404 MB → 166 MB.

    That predicts exactly the split observed: on a part with enough L3 to absorb plain LLVM's footprint, the codegen win is cancelled by the footprint cost and llvm lands at or below C — which is what the table in this issue shows. On a cache-starved part the footprint cost dominates first, so both arms win and SIM wins hugely.

    If that reading is right, the practical consequence cuts the way this issue argues, only harder: SIM is worth more the less cache the host has, so benchmarking it exclusively on a large-L3 part systematically understates it.

    Caveats, so these are read at the right precision

    • The counters were multiplexed at ~35% enable time (more events than PMU slots), so they are scaled estimates. The 40–50% deltas are far outside multiplexing error; I would not quote single-digit differences from that run.
    • LLC-* events are unsupported on this Zen 3 part, so cache-misses is the last-level proxy.
    • Linux ran LLVM 20.1.2 against 20.1.8 elsewhere — a patch-level codegen difference.
    • Linux was headless under Xvfb; a windowed run carries present/GX cost that a headless one does not, which shifts how CPU-bound the workload is. That is a real confound against the numbers in this issue and I would not rule it out as part of the gap.
    • Different scene selection: one savestate here versus three.

    Host: Ubuntu, Ryzen 9 5900X (12C/24T, 64 MB L3 as 2×32 MB), RTX 3090, performance governor, headless, nothing else running.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions