Skip to content

LLVM backend: Colosseum does not get through boot (x86-64 and aarch64; C backend does) #20

Description

@dougchansan

Correction (updated): I originally filed this as AArch64-specific. That was wrong. The failure reproduces on x86-64 with the same DOL and the same harness, so this is a Colosseum + LLVM-backend defect, not an ARM one. Details and measurements below.

Building Pokémon Colosseum (GC6E01) with the LLVM backend produces a module that does not get through the boot sequence. The C backend, same DOL, same runtime, same automation harness, boots normally.

Measurements

Cold boot, A-press driver, no savestate, identical harness on each platform:

platform backend frames produced fps speed
x86-64 Windows C 4262 44.77 0.749
x86-64 Windows LLVM 16 0.26 1.00
aarch64 (Pi 4) C 1417 9.74 0.166
aarch64 (Pi 4) LLVM 8 5.74 (stale) 0.634

The guest is not deadlocked — it executes (speed is nonzero, the VI keeps ticking), it just never makes forward progress. Note the LLVM arm reports a higher speed than C on the Pi while producing no frames: it is spinning cheaply instead of doing work.

Luigi's Mansion (GLME01) with the LLVM backend runs fine on the same Pi (769 frames, 6.11 fps), so the AArch64 LLVM backend is not broken in general.

Where it ends up

perf on the wedged aarch64 run puts ~58% of samples in three guest functions, which the DOL disassembly identifies as the thread scheduler:

  • 0x8009BBE0 — OSSaveContext (saves GQR1–5 via mfspr 0x393..0x397, CR, LR, MSR, CTR, XER)
  • 0x8009BC50 — OSLoadContext, which resumes a thread with rfi
  • 0x800A17E0 — run-queue enqueue / reschedule

None of these appear in the C backend's profile, which is dominated by real game work (func_800FD5E0, paired-single ops, FP helpers).

What is NOT the cause

I instrumented the shared runtime to rule these out rather than guess:

  • Thread resume is correct. ppc_rfi fires ~150M times in 70s (vs ~150k for C, ~1000x more), and at every sampled resume CPUState is correct: srr0=0x800A183C, r3=0x00000001. 0x800A183C is the instruction after bl OSSaveContext, and r3=1 is exactly the "resumed" return value the scheduler tests. The C backend resumes at the same address with the same value, so this churn is a normal idle loop — the LLVM arm is just doing ~1000x more of it.
  • Entry at that PC is valid. In the generated IR the address is a real entry case (guest_800A183C_b23), reached from native_entry, and the block loads r3 from CPUState and branches on icmp eq i32 %state2.7, 0 correctly.
  • Not an unsupported instruction. All 1643 fallbacks in the generated CSV are embedded-data; there are no unmodeled opcodes.
  • Not exception delivery. Exception counts stayed under threshold in both arms.
  • Not our open PRs. Regenerating with Register only the LLVM targets this build can emit for, and build the bench on Windows #17/Route fallback blocks through per-edge trampolines so the phi matches the CFG #18/Key the object cache on the sources that generate the objects #19 applied produces byte-identical objects (4890/4890 identical, 0 changed) and an identical module hash, so those fixes neither cause nor fix this.

Still open

I have not isolated the primitive that diverges. The dispatch-address stream diverges from the C backend within the first handful of dispatches at boot, and the LLVM arm reaches far less distinct guest code over a run (909 vs 4000+ distinct addresses in the same window).

Reproduction is easy on x86-64 now, so this no longer needs ARM hardware to chase. Happy to keep digging or hand over what I have.

Activity

  1. changed the title [-]AArch64: fixed-chunk LLVM backend hangs Colosseum during boot (C backend runs)[/-] [+]LLVM backend: Colosseum does not get through boot (x86-64 and aarch64; C backend does)[/+] on Aug 22, 2026
  2. dougchansan commented on Aug 22, 2026

    @dougchansan
    ContributorAuthor

    Further evidence, and one hypothesis eliminated.

    The DOL was not the cause. Our Windows workspace had been recompiling a modified Colosseum DOL (md5 b299467dd305) while the Pi used the retail one. I re-extracted from the retail disc image with DolphinTool (md5 c55df119fad4, matching the Pi) and rebuilt every arm from that. The defect is unchanged.

    Cold boot, x86-64, retail disc data, identical harness per arm:

    arm fps vps speed frame_count present_count
    C 59.94 59.94 1.000 5353 10755
    LLVM (upstream rewrite) 0.246 59.94 0.999 8 1575
    LLVM (AOT region backend) 0.250 59.94 0.999 8 728

    Two details that sharpen the picture:

    • The guest runs at full speed (speed ≈ 1.0) with the video interrupt ticking normally (vps 59.94). This is not a slow module — it is a module that does not render.
    • present_count keeps rising while frame_count is frozen at 8. The runtime is still swapping buffers; the game is simply re-presenting the same image because its main thread never advances.

    It also ruins a mid-game savestate scene, not just boot. Benchmarking part1-in-town.sav with the interleaved harness, all three LLVM backends are degenerate while C is healthy:

    arm samples fps median speed median
    C control 114 54.67 0.9159
    LLVM old (pre-rewrite) 11 0.34 1.3250
    LLVM new (rewrite) 17 0.58 2.1241
    LLVM aot 26 0.85 3.1201

    Worth flagging for anyone benchmarking this title: because speed counts retired guest cycles and this defect burns cycles in a spin that does no work, the metric rewards the broken arm. On speed alone the AOT arm reads as a 3.58x win over C. It is an artifact — fps and sample count expose it. I withdraw an earlier internal Colosseum figure that was measured this way.

    So the defect: affects all three LLVM backends (pre-rewrite, rewrite, and AOT), reproduces on x86-64 and aarch64, on retail data, at boot and from a savestate. The C backend is unaffected throughout.

  3. dougchansan commented on Aug 23, 2026

    @dougchansan
    ContributorAuthor

    Further investigation. No root cause yet, but three useful narrowings and one correction to something suggested along the way.

    Correction first

    A hypothesis was raised that the C backend avoids the bug because its text0 chunk fails SMC verification and the whole 0x80003100-0x80005500 range therefore runs on Dolphin's JIT rather than recompiled code. That is not what happens. Runtime counters from a C-backend run:

    [staticrecomp] shutdown: native=1494658421 fallback=0 native_exc=78253
                   hook_fb=127384220 smc_failed=0 verifications=98 reverify_events=5
    

    1.49 billion natively executed instructions, zero fallbacks, zero SMC failures. The C backend is genuinely running recompiled code for this title. Both backends report fallback=0 smc_failed=0, so the difference is not one of them quietly running on the JIT.

    New: the divergence is visible in the runtime counters

    Same title, same scene, same runtime, C versus LLVM:

    counter C LLVM
    native instructions 549,841,950 181,820,064
    native_exc 38,659 521
    hook_fb 86,909,490 108,309
    fallback 0 0
    smc_failed 0 0
    guest cycles 13,686,478,575 14,779,848,761

    The LLVM arm retires a comparable number of guest cycles while taking 74x fewer exceptions and 800x fewer hook fallbacks. It is not doing the same work. That is a much sharper signal than the frame counter and it is cheap to reproduce.

    New: the rendering ratio is 238:1

    Measured directly: 238 VI interrupts fire per rendered frame under LLVM, against 1 per frame under C. So the game's render-trigger condition is being satisfied roughly once per 238 VI callbacks rather than every one. That is consistent with the counters above — something the game does every frame is not happening.

    Ruled out this round, with evidence

    The VI exception stub at 0x800039AC was inspected in the emitted IR and is correct:

    • SPRG1/2/3 saves emit ppc_mtspr(ctx, 273/274/275, ...) with the right SPR numbers
    • SRR0 set to 0x800C0EAC, SRR1 to old_msr | 0x30
    • gpr[2] = interrupted PC, gpr[3] = 0x500, gpr[4] = srr1
    • ppc_rfi tail call with the correct CIA
    • The alignment stub at 0x80003AAC has the same structure with exception type 0x600

    Two earlier hypotheses were also refuted on inspection: 0x7C5143A6 at 0x800039AC decodes as mtspr SPRG1, r2 (correct), and the mfmsr/ori pair lowers correctly to or i32 %117, 48.

    Not yet examined

    The FPU-unavailable stubs at 0x800044AC, 0x800045AC, 0x800046AC are structurally unlike the others — they save and restore the Condition Register around the handler (mfcr / andi. / sync / mtcrf 0xFF / xoris). A CR save/restore error there would corrupt condition codes across an exception and change branch outcomes in game code, and the FPU-unavailable exception fires during startup before the FPU is enabled. That is the next place worth looking, though it is a suspicion rather than a finding.

    Happy to keep digging or hand this over. The counter comparison above is the cheapest reproduction of the divergence I have found so far.

  4. dougchansan commented on Aug 23, 2026

    @dougchansan
    ContributorAuthor

    More probing. Not resolved, but the divergence now has a number and a direction.

    The sharpest signal: the LLVM arm barely takes any exceptions

    Runtime counters, same title, same scene, same runtime, comparable guest cycles:

    counter C LLVM
    guest cycles 13,686,478,575 14,779,848,761
    native_exc 38,659 521
    hook_fb 86,909,490 108,309
    bursts 2,712,107 1,132,575
    fallback / smc_failed 0 / 0 0 / 0

    The LLVM module retires slightly more guest cycles while taking 74x fewer exceptions. External interrupts are how VI reaches the guest, so a title that renders nothing while running at full speed with the host VI ticking at 59.94 is consistent with interrupts not being delivered rather than with slow code.

    That also matches the independently measured 238 VI interrupts per rendered frame.

    What the working arm actually spends its time on

    The runtime has a built-in dispatch sampler (STATICRECOMP_DISPATCH_SAMPLES=1) that records dispatch PCs and dumps the top sites at shutdown. On the C backend the top two are:

    dispatch-site pc=8009df3c samples=51999
    dispatch-site pc=8009df64 samples=51576
    

    0x8009DF3C disassembles to:

    8009DF3C  mfmsr  r3
    8009DF40  rlwinm r4, r3, 0, 17, 15     ; clear MSR[EE]
    8009DF44  mtmsr  r4
    8009DF48  rlwinm r3, r3, 17, 31, 31    ; return previous EE
    8009DF4C  blr
    

    That is OSDisableInterrupts(), with OSRestoreInterrupts() immediately after it at 0x8009DF50. The busiest thing the working backend does is enter and leave critical sections — which is exactly the machinery that governs whether interrupts get delivered.

    Note the LLVM arm could not be sampled the same way: it never processes the automation stop command (waited 40s), so its counters never flush. That inability to shut down cleanly is itself consistent with the emulation loop being stuck.

    Ruled out this round

    mtmsr lowering is not the difference. Both backends do the same thing:

    • C: ctx->msr = ctx->gpr[rS]
    • LLVM: write_slot(DOLIR_STATE_MSR, gpr(rS)) via the IR builder

    Keeping guest state in memory does not fix it. I built the title with the --state-in-memory change from #21, which removes the entry prologue, the materialization barriers and the post-call reloads, so MSR is written to CPUState immediately rather than held in a promoted SSA value. If delayed visibility of MSR[EE] were the cause, that should have changed the outcome. It does not:

    arm fps frame_count
    C 59.94 5353
    LLVM (current) 0.246 8
    LLVM + --state-in-memory 0.249 8

    Identical failure. So the defect is not related to state hoisting.

    Gather pipe is not obviously implicated. The static recomp core routes 0xCC008000 page writes straight to GPFifo, and GPFifoManager::Write32 self-kicks via CheckGatherPipe().

    Where I would look next

    The native_exc gap is the cheapest reproduction of the divergence and does not need a benchmark run — boot either module for 70s and read the shutdown line. Anything that explains why the LLVM module takes 521 exceptions where C takes 38,659 over the same workload very likely explains the rendering failure too.

    The suggestion that this title is stuck on its THP intro video is consistent with everything above, though I have not confirmed it directly: a video that never advances would look exactly like this if the interrupts driving it are not arriving.

  5. dougchansan commented on Aug 23, 2026

    @dougchansan
    ContributorAuthor

    Ran the runtime's lockstep verifier against this title, and then ran the same thing against a title that works. Posting both because the second invalidates the first as evidence.

    Lockstep on Colosseum looked like a smoking gun

    Built the module with DOLRECOMP_LLVM_WRITE_JOURNAL=1 and ran with STATICRECOMP_LOCKSTEP=1. It diverges from Dolphin's core at the first checked region:

    DIVERGE #1 entry=0x80003158 end=0x80005498:
      r3:  N=0x6f05      I=0x6f09
      r4:  N=0x8039a77c  I=0x8039a6fc
      xer: N=0x20000000  I=0x0
      mem[0x8048e722]: N=0x20 I=0x00
    

    0x20000000 is XER[CA]. r3 is 4 lower and r4 is 0x80 higher, which is exactly 4 iterations of the 0x20-stride block-fill loop at 0x80005498. Reproducible, and unchanged when the step cap is raised from 512 to 65536, so it is not a comparison-window artefact.

    But a working title diverges identically

    Mario Kart Double Dash renders fine under the same LLVM backend (50.0 fps, 5271 frames in the same 90s window). Its lockstep output:

    DIVERGE #1 entry=0x80003158 end=0x80003500:
      mem[0x803e22e2]: N=0x20 I=0x00
      mem[0x803e22e3]: N=0x32 I=0x00
    DIVERGE #2 entry=0x80003500 end=0x80003500:
      r3:N=0x3288,I=0x32a1  r4:N=0x803642dc,I=0x80363fbc
    

    Same entry PC, same byte values (0x20, 0x32) at the analogous address, same block-fill loop signature with r3/r4 offset by a whole number of iterations. 25 divergences reported for each title.

    So lockstep divergences on this codebase are not by themselves evidence of a miscompile. They appear on a title that works perfectly. The comparison runs native code a region at a time against a stepwise interpreter, and the two are simply not sampled at equivalent points inside tight loops. Anyone using this tool here should calibrate against a known-good title first — I nearly reported the XER carry difference above as the root cause.

    What still distinguishes the two titles

    The counter gap remains the only signal I have that separates a working title from this one:

    counter C LLVM
    guest cycles 13,686,478,575 14,779,848,761
    native_exc 38,659 521
    hook_fb 86,909,490 108,309

    74x fewer exceptions over slightly more guest cycles, with the host VI ticking normally at 59.94. That is still the cheapest reproduction and still the most likely place the answer lives.

    Also confirmed this round: the module is genuinely loaded and executing recompiled code, not silently falling back. native=181820064 fallback=0 smc_failed=0, module hash matches the retail DOL.

  6. dougchansan commented on Aug 23, 2026

    @dougchansan
    ContributorAuthor

    Broke the exception counter down by vector. This is the sharpest split so far.

    Exceptions by vector, same title, same scene, 70s

    vector C backend LLVM backend
    0x800 FP unavailable 24,488 1
    0xC00 system call 59,512 19
    total 84,000 20

    Not a scaling difference — the LLVM module takes twenty exceptions where the C module takes eighty-four thousand, over comparable guest cycles at full speed.

    The first 87 exceptions are identical

    Logging the vector and SRR0 of every exception and diffing the two sequences: they match exactly for 87 exceptions, then diverge. Both arms then continue hitting the same two sc sites, but the LLVM arm's rate collapses (314 logged versus the 5000 cap for C).

    The two system-call sites are both cache/sync:

    8009B2E8  mtctr r4
    8009B2EC  dcbf  0, r3          ; DCFlushRange loop
    8009B2F0  addi  r3, r3, 0x20
    8009B2F4  bdnz  0x8009b2ec
    8009B2F8  sc                   ; SRR0 = 8009B2FC
    8009B2FC  blr
    
    80098034  sc                   ; SRR0 = 80098038, bare sync call
    80098038  blr
    

    Tested and ruled out this round

    Cache-op handling is not it. dcbf is one of the few places the backends genuinely differ: the C backend emits ppc_fallback_instruction and leaves native execution for every one, while the LLVM backend handles it inline through ppc_cache_control. That accounts for hook_fb being 86.9M on C versus 108K on LLVM, and it was a good candidate for starving the runtime of re-entry points.

    Forcing exactly those two functions through the JIT with STATICRECOMP_FALLBACK_RANGES=8009B2E0-8009B310,80098030-80098040 changes nothing: 0.2498 fps versus 0.2467 unforced, frame_count 8 either way.

    HookCacheControl is installed and correct — ICBI invalidates the iCache and the JIT, so cache operations are not silently dropped on the LLVM path.

    The number I cannot explain

    FP unavailable: 24,488 versus 1.

    The C backend emits a check before every FPU-using instruction (if (ppc_op_uses_fpu(inst->op)) → ppc_fp_available_inline). The LLVM backend checks once per region and caches the answer in fp_available_checked_, which is reset at region entry and by reloadUsedState().

    I found and fixed two genuine cache-invalidation holes while chasing this (an mtmsr write not invalidating the cache, and the native call-resume path reloading values without dropping the emitter's conclusions — both in #22). Neither moves this title: together they change one of 4890 emitted objects.

    So the caching holes are real but are not what suppresses 24,487 FP-unavailable exceptions. Something else is preventing the LLVM module from ever taking that exception, and the guest's lazy FPU context switching is built on receiving it. That seems the most promising thread for anyone with more context on this backend than I have.

    Everything eliminated so far

    SMC/JIT fallback (both run fully native, fallback=0 smc_failed=0), entry dispatch, rfi/thread resume (CPUState correct at every resume), mtmsr lowering, state hoisting (--state-in-memory is byte-for-byte the same failure), gather pipe, cache ops, DOL variant (retail disc data reproduces it), and lockstep divergences as evidence (a working title shows the same ones).

  7. dougchansan commented on Aug 23, 2026

    @dougchansan
    ContributorAuthor

    Correction to my last comment, and a much better number.

    I suggested the FP-unavailable disparity (24,488 versus 1) was the most promising thread. It is a symptom, not a cause. Probing MSR at every rfi shows why.

    MSR[FP] at thread resume

    C LLVM
    rfi resumes in 60s ~800,000 91,799,710
    MSR[FP] set 742,467 91,799,710
    MSR[FP] clear 57,533 (7.7%) 290 (0.0003%)
    ppc_fp_available helper calls 0 1

    The C arm resumes threads with FP disabled 7.7% of the time — ordinary lazy-FPU behaviour, and each of those is an FP-unavailable exception waiting to happen. The LLVM arm resumes with FP disabled 290 times out of 91.8 million.

    So the LLVM module never takes FP-unavailable because it never runs a thread that lacks the FPU. That follows directly from the real problem: it is performing 115x more thread resumes than the C backend while doing almost no work.

    What that rules out

    The FP-availability caching in the LLVM emitter is not what suppresses those exceptions. I found and fixed two genuine invalidation holes there (#22) and they change one of 4890 emitted objects on this title — consistent with this measurement, not with them being the cause.

    Also ruled out this round: forcing the DCFlushRange/sync functions through the JIT with STATICRECOMP_FALLBACK_RANGES=8009B2E0-8009B310,80098030-80098040 changes nothing (0.2498 fps versus 0.2467 unforced). So the inline-versus-fallback handling of dcbf is not it either, despite being one of the few genuine behavioural differences between the backends (hook_fb 86.9M on C versus 108K on LLVM).

    Where this leaves it

    The guest is stuck resuming a thread ~92 million times in 60 seconds, at full speed, taking essentially no exceptions and making 19 system calls where the C backend makes 59,512. Every resume carries the correct state — SRR0 0x800A183C, r3=1, which is the "resumed" return value the scheduler tests — and the emitted IR for that entry point is correct.

    So a thread is being resumed correctly and immediately blocking again, forever, without asking the OS for anything. What it is waiting on is still unknown, and that is the question I would put to anyone with more context on this runtime.

    Cheapest reproduction remains the shutdown counter line after a 70s boot: native_exc 521 versus 38,659, no benchmark needed.

  8. dougchansan commented on Aug 23, 2026

    @dougchansan
    ContributorAuthor

    Colosseum (GC6E01): what the blocked thread is waiting on

    Pinned state part1-phenac-city.sav, Windows, same runtime build for both arms.
    Baseline: C = 29.98 fps, LLVM = 0.2487 fps, VI ~59.94 in both.

    1. The thread

    Guest thread 0x803FBCB0, constant stack chain:

    803FBCB0 sp=8048E678 <- 804001F0 <- 800A19B0 <- 800D31E4 <- 801E115C <- 80005DE0 <- 800FE808
    

    0x800D31E4 is the resume point of a 3-second timed wait:

    800D31C0  stb   r0, -0x5d8f(r13)   ; clear flag
    800D31E0  bl    0x800a1990         ; block / yield
    800D31E4  lbz   r0, -0x5d8f(r13)   ; resumes here, re-reads flag
    800D31F4  lwz   r0, 0xf8(r31)      ; 0x800000F8 = bus clock
    800D31FC  srwi  r0, r0, 2          ; ticks/sec
    800D3200  divwu r0, r3, r0         ; elapsed seconds
    800D3204  cmplwi r0, 3
    800D320C  stb   r30, -0x5d8f(r13)  ; timeout -> set flag itself
    800D3218  beq   0x800d31e0         ; else keep waiting
    

    2. What sets the flag

    Only one other writer, 0x800D3F50 -- a 3-instruction leaf:

    800D3F50  li  r0, 1
    800D3F54  stb r0, -0x5d8f(r13)
    800D3F58  blr
    

    Registered at 0x800D3CA4 via 0x801C01C8, which stores the pointer into the
    VI globals at 0x80456DA8 bracketed by OSDisableInterrupts/OSRestoreInterrupts
    -- i.e. VISetPostRetraceCallback.

    So the thread waits for a VI retrace interrupt, and escapes only via its own
    3-second timeout. 0.2487 fps == one frame per ~4.0 s == the timeout path. The
    game is not slow; it is timing out once per frame.

    3. Why the retrace never arrives

    Added counters to the raise (ProcessorInterface::UpdateException) and delivery
    (PowerPCManager::CheckExternalExceptions) paths:

    arm raised delivered blocked (EE=0)
    Colosseum C 33,037 31,209 37,158
    Colosseum LLVM 484 462 1,056,463
    MKDD LLVM (control) 28,263 25,198 3,494

    raised counts only 0->1 transitions, so the low count is a consequence: the
    interrupt sits pending because MSR.EE is stuck at 0. 30% of checks are
    EE-blocked under LLVM vs 0.76% under C. MKDD under LLVM is healthy, so this is
    not a general backend fault.

    4. Where it sits with interrupts disabled

    Sampling guest PC on every 64th EE-blocked check:

    PC C LLVM
    800a128c (one-instruction blr -- null thread-switch hook) 10 11,077
    800a183c (after bl 0x8009bbd0, in __OSReschedule) 32 5,466

    Matching dispatch-site sampling: the LLVM arm spends ~97% of dispatches at
    800d31e4 + 800a128c/1930/183c. The C arm's top sites are OSDisableInterrupts
    /OSRestoreInterrupts and ordinary game code.

    The guest is parked in the OS scheduler's idle path with interrupts disabled.

    Ruled out

    • Timebase hoisting. Made mftb/DOLIR_STATE_TIMEBASE reads volatile and
      non-cacheable (branch fix/timebase-volatile, module hash changed
      625021ef -> e0e7c371). No effect: 0.2487 fps. The 3-second timeout does
      fire, so the guest clock advances correctly.
    • rfi lowering. STATICRECOMP_FALLBACK_RANGES=8009bc50-8009bcc0 forces
      OSLoadContext (ends in rfi) to the interpreter. No change (0.2479 fps).
      syncDirtyState() does flush SRR1 before ppc_rfi, and returnFromBody()
      does not re-sync, so the restored MSR is not being clobbered.
    • Earlier rounds: SMC/JIT fallback (fallback=0 in both arms), entry dispatch,
      mtmsr lowering, state hoisting, gather pipe, cache ops, DOL variant,
      FP-available cache holes, lockstep divergence signatures.

    Note

    STATICRECOMP_FALLBACK_RANGES over 8009df3c-8009df90
    (OSDisableInterrupts/OSRestoreInterrupts) or over the whole OS region
    80098000-800a4000 hangs outright (fps=0, vps=0, delivered=0). Mixed
    native/interpreter execution across MSR manipulation is itself broken, so this
    is not a usable bisect tool in that region -- worth a separate look.

    Instrumentation to revert

    • Source/Core/Core/HW/ProcessorInterface.cpp (g_dr_ext_raised)
    • Source/Core/Core/PowerPC/PowerPC_Exceptions.cpp (counters + dr_note_ee0)
    • Source/Core/Core/PowerPC/StaticRecomp/StaticRecompCore.cpp (shutdown dump)
  9. dougchansan commented on Aug 23, 2026

    @dougchansan
    ContributorAuthor

    Lockstep: an msr divergence unique to Colosseum

    STATICRECOMP_LOCKSTEP=1, both titles LLVM, both capped at 74 reports:

    title divergences of which involve msr
    MKDD (working control) 74 0
    Colosseum 74 1

    The one Colosseum msr divergence lands in OSDisableInterrupts:

    entry=0x8009E724 end=0x8009DF3C
    msr:  N=0x9932  I=0x1032
    srr1: N=0x9932  I=0x1030
    srr0: N=0x8009df3c  I=0x8009e758
    r12:  N=0x800a128c        (the null thread-switch hook)
    

    Decoding: I = ME|IR|DR|RI, EE=0. N = same plus EE(0x8000), FE0(0x800),
    FE1(0x100)
    . The native side has taken an exception (srr0/srr1 populated) that
    the interpreter has not, and the whole GPR file has diverged with it.

    FE0|FE1 being set on the native side ties this to the MSR/FP-exception-mode
    path. Not conclusive on its own -- one report against a 74-report cap -- but it
    is the only msr divergence across both titles and it is in exactly the function
    that gates interrupt delivery.

  10. dougchansan commented on Aug 23, 2026

    @dougchansan
    ContributorAuthor

    Lockstep, uncapped and filtered to msr

    The report cap was never the limit -- STATICRECOMP_LOCKSTEP_MAXREPORT already
    defaults to unlimited. The real bound is m_ls_checked, which dedupes by entry
    PC, so "74 divergences" meant 74 distinct diverging entry PCs. Added two
    options to get past it:

    • STATICRECOMP_LOCKSTEP_FILTER=<substr> -- only report diffs containing the
      substring (suppressed reports counted, not printed)
    • STATICRECOMP_LOCKSTEP_NODEDUP=1 -- re-check entry PCs instead of once each

    Also raised STATICRECOMP_LOCKSTEP_STEPCAP to 100000; most of the original 74
    were CTRLFLOW:...steps=512 artifacts of the interpreter running out of steps.

    Result: 44 msr divergences, 42 of them one bit

    180 s run, Colosseum LLVM, part1-phenac-city.sav:

    count native interpreter delta
    42 0x9932 0x1932 0x8000 = EE
    1 0x9932 0x1032 EE + FP
    1 0x3932 0x1030 --

    Entry PCs:

    count entry end
    21 0x800FEA98 0x800FEA98
    21 0x800FEA98 0x800FEA90
    1 0x800B8E28 0x800A128C
    1 0x8009E724 0x8009DF3C

    42 of 44 come from one loop. 0x800FEA90-0x800FEAE4 walks a linked list at
    r13-0x5B88 and bctrls each node's handler at offset 20, then calls
    OSDisableInterrupts -- a callback dispatcher running in interrupt context,
    where EE should be 0. The interpreter agrees; native has it set.

    Why that is plausible as a bug

    The OS interrupt primitives are all mfmsr/mtmsr pairs:

    8009DF3C  mfmsr r3; rlwinm r4,r3,0,17,15; mtmsr r4    ; OSDisableInterrupts
    8009DF50  mfmsr r3; ori r4,r3,0x8000;     mtmsr r4    ; OSEnableInterrupts
    8009DF64  cmpwi r3,0; mfmsr r4; ...;      mtmsr r5    ; OSRestoreInterrupts
    

    cpu->msr is written from outside generated code (the runtime vectoring an
    exception, or ppc_rfi). If mfmsr reads an SSA-promoted copy hoisted at region
    entry, the following mtmsr writes the stale value straight back and clobbers
    MSR[EE]. Same bug class as the FP-available cache and the timebase.

    Fix attempted, and its result

    Branch fix/msr-in-memory: slotInMemory() returns true for DOLIR_STATE_MSR
    unconditionally, so MSR always lives in CPUState and is never SSA-promoted.
    Module regenerated (hash 94f97ee7).

    It does not fix Colosseum. Three runs: 0.4126, 0.2487, 0.2487 fps. The first
    reading was an outlier; the reproducible value is 0.2487, identical to before
    the change. blocked_ee essentially unchanged (1.013M vs 1.058M).

    The divergence signature did change -- the EE-only pattern is no longer dominant;
    a N=0x3932,I=0x1030 pattern (native has FP/FE0/FE1/RI set where the interpreter
    is in exception-entry state) now accounts for 42011 of 42360. Whether that is a
    genuine second bug or an artifact of the shadow interpreter is not established.

    Caveat on lockstep as evidence

    A Mario Kart lockstep run crashed with Invalid read from 0x00000054 at
    PC=0x801a1bc8 and threw modal dialogs. The shadow interpreter re-executes
    blocks and can double-issue side effects, so lockstep output in this region is
    suggestive, not authoritative. (Harness scripts now write
    UsePanicHandlers = False so runs cannot raise blocking dialogs.)

  11. dougchansan commented on Aug 23, 2026

    @dougchansan
    ContributorAuthor

    Exception histogram by vector (new, and the sharpest result so far)

    Added a per-vector histogram in ppc_take_exception (GXRuntime/src/core/cpu_exception.c)
    and regenerated both Colosseum modules. Same scene, same savestate:

    arm total vec 0x800 (FP unavailable) vec 0xC00 (system call)
    C 19,858 9,297 10,561
    LLVM < 32 -- --

    The C arm raises ~20k guest exceptions of exactly two kinds. The LLVM arm raises
    fewer than 32 in total -- the histogram never even reached a dump threshold of 32.
    Both sc and FP-unavailable are correctly lowered in the LLVM backend
    (src/backend/llvm/control_flow.cpp:202 calls ppc_system_call_exception;
    emitFPAvailable in special_instructions.cpp), so this is not missing codegen --
    the guest simply never reaches that code.

    Consistent with this, the one non-EE lockstep msr pattern is N=0x3932, I=0x1030:
    native has MSR[FP] (0x2000) set where the interpreter has it clear. A guest with
    FP permanently enabled never traps, so the OS lazy-FPU context switch never engages.
    Note the FP-available cache fix did not change this, so the suspicion is MSR[FP]
    itself being wrong in native state rather than the cached answer.

    Thread scheduling is identical

    Tracing OSCurrentThread (0x800000E4) transitions from the run loop, both arms
    alternate between exactly the same two threads -- 803FBCB0 (the blocked one) and
    803FB998 -- in the same pattern. The LLVM arm just switches ~8x less often. No
    thread is starved out of the run queue; the structure is the same.

    The framerate is bimodal

  12. dougchansan commented on Aug 23, 2026

    @dougchansan
    ContributorAuthor

    Root cause found and fixed: interrupt-delivery starvation by burst phase-locking

    The exception histogram was the right thread to pull, but it was a symptom. The chain, fully traced:

    1. Only 3 sc instructions exist in the whole DOL -- 0x8009B2F8 is DCFlushRange's flush-via-syscall. The ~10k missing syscalls (and the missing FP-unavailable traps) were per-frame game work that never ran. Symptom, not cause.
    2. 0x800A1990 -- the "block/yield" in the retrace wait loop -- disassembles to OSYieldThread: OSDisableInterrupts; SelectThread(1); OSRestoreInterrupts. Colosseum's retrace wait is a yield-spin, not a queue sleep. Dispatch sampling shows it iterating at ~190k/s under LLVM. The other tested titles block on wait queues -- this is why only Colosseum failed.
    3. Under the LLVM backend the entire critical section runs inside native calls within one dispatch. The only burst-ending points are the budget guards, which sit at call sites -- in OS code, almost entirely inside the interrupt-disabled window. Burst boundaries phase-lock to EE=0.
    4. External interrupts are only deliverable when the host regains control with EE=1: the VI interrupt sat pending-but-blocked at 30% of checks, only 1.2% of raises were ever delivered, the post-retrace callback never ran, and every frame escaped via the game's 3-second timeout. 1 frame / ~4 s = 0.2487 fps exactly. This also explains the bimodal 0.41/0.47/0.62 runs (occasional lucky boundary in the EE=1 zone) and the THP cold-boot hang (same wait pattern).

    mtmsr was lowered as a plain state write; Dolphin's own JITs end the block at mtmsr and check for pending interrupts. Two-part fix:

    Results

    before after
    Colosseum LLVM 0.2487 fps 27.4-28.8 fps
    Colosseum C (control) 29.98 29.99
    MKDD LLVM (control) 48.65 51.0
    Colosseum LLVM cold boot hung at THP videos ~55 fps, boots

    Interrupt counters: raised 484 -> 31,777, delivered 462 -> 29,856, blocked 30% -> 0.5% -- matching the C profile. Renders the Phenac City scene correctly.

    Remaining loose end from this investigation: forcing STATICRECOMP_FALLBACK_RANGES across OSDisableInterrupts/OSRestoreInterrupts still hangs outright -- mixed native/interpreter execution across MSR manipulation deserves its own issue.

  13. dougchansan commented on Aug 24, 2026

    @dougchansan
    ContributorAuthor

    Uncapped runs (EmulationSpeed = 0): AOT is the fastest arm, and a second scene-triggered stall surfaced

    Same pinned scene, 45 s uncapped windows, order-balanced, two runs per arm.

    arm run 1 fps (speed) run 2 fps (speed) end frame
    C 42.08 (1.41x) 42.17 (1.40x) 95058 / 95008
    LLVM AOT + fix 46.95 (1.56x) 43.39 (1.47x) 95219 / 95224
    LLVM new + SIM + fix 34.65 (4.37x) 36.10 (4.32x) 93395 / 94212
    LLVM new + fix 28.51 (0.36x) 27.60 (2.69x) 93763 / 94212
    LLVM old + fix 22.41 (0.74x) 25.80 (0.85x) 94316 / 94308

    Clean readings: C 42.1 fps / 1.41x, stable to 0.2% across runs. AOT 43-47 fps / ~1.5x -- the only arm faster than C on this title, up to 1.12x.

    The other three rows are not valid speed measurements. Uncapped windows cover 2-4x more game time than the capped runs did, and they reach scene positions the capped windows never hit. At specific frame counts (93395, 93763, 94212 -- game-state positions, reproducible across arms and builds) the new and state-in-memory arms hard-stall rendering: no frames present (screenshot requests return nothing), while the guest CPU spins at up to 4.4x realtime -- which is why SIM "reads" speed=4.3 with fewer fps than C. C, AOT and old pass the same positions cleanly.

    Two corrections this forces:

    1. The earlier claim that the two parked codegen optimizations "wedge Colosseum deterministically" attributed the stall to those changes. The stall frame numbers recur on the unmodified PR variant -- they are scene positions, and the faster variants simply reached them within the measurement window. Those optimizations may be fine; they need re-testing once the stall is fixed.
    2. The mtmsr delivery fix cures the steady-state retrace wait (the 0.25 fps symptom), but a second, scene-triggered render stall remains on the new-base LLVM arms -- same family as the original THP-video hang, and now the top open Colosseum issue. An edge-gated delivery variant was tried and does not prevent it.

    Capped behavior is unaffected: the shipped variant still runs the scene at 27.9 fps with screenshots after these experiments.

  14. dougchansan commented on Aug 24, 2026

    @dougchansan
    ContributorAuthor

    Scene clarification: no C regression, and AOT confirmed fastest on the historical benchmark scene

    A "C used to hit 60 on Colosseum" concern turned out to be scene identity. All the Aug 20-21 Colosseum campaigns in games/pkmn-colo-recomp/build/benchmarks/ used part1-in-town.sav (uncapped C median 53-55 fps, speed 0.89-0.92); the exact-59.94 readings were the capped cold-boot intro, which is 60 Hz content. The matrix earlier in this thread used part1-phenac-city.sav, a heavier cutscene (C: 42). Same backend, different scenes -- willie-battle.sav reads 35 for scale.

    Re-ran today's builds uncapped on the historical scene:

    arm today, part1-in-town.sav Aug 21 campaign
    C 53.5 fps (speed 0.898) 53-55 (0.89-0.92) -- unchanged
    LLVM AOT + fix 56.5 (0.944) 0.85 (pre-fix, broken)
    LLVM new + SIM + fix 45.9 0.85
    LLVM new + fix 37.6 0.85

    So on both scenes measured, AOT is the fastest arm (1.06x C in town, up to 1.12x on the cutscene). And the scene-stall pathology reproduces here too: the SIM and plain-new runs both ended frozen at frame 176641 -- a second reproducible stall position for the new-base non-AOT arms, strengthening that as the top open Colosseum issue.

  15. dougchansan commented on Aug 24, 2026

    @dougchansan
    ContributorAuthor

    Stall investigation: runtime half fixed and pushed; the new-emitter half is characterized but not yet root-caused

    Fixed and pushed (RecompCore PR #18, new commit): external-interrupt delivery is now restricted to post-mtmsr boundaries, matching the block-ending JITs. Delivering at every EE=1 boundary could preempt the AX audio callback mid-work at scene transitions and consume any backend probabilistically -- the C arm crawled in one of two uncapped runs before, and passes 3/3 at 53.6-56.8 fps after. AOT stays clean, and the capped Colosseum fix is unaffected (25-29 fps).

    Still open: the deterministic scene-transition stall on the new-base LLVM arms (plain and --state-in-memory; the SIM arm is the upstream base + state-in-memory, so this is one bug, not two). AOT, old and C all pass. State of knowledge:

    • Stall positions are game-state-deterministic (frame 176641 from part1-in-town.sav, 94212/93763/93395 on the Phenac scene).
    • The stalled state is persistent: a savestate taken in it stalls even the C backend on load (states/pre-stall.sav), so a guest-visible structure goes wrong on the approach and nothing recovers.
    • In the stalled state interrupts go quiet, not stormy: <4k raises/20 s vs 20k healthy (SI/DSP/VI/PE all ~stopped), zero deliveries, GPU never finishes a frame, while the guest spins through the interrupt-callback walker, the AX voice list and the yield loop at full speed.
    • Ruled out by direct experiment: the inline paired-single FMA path (disabled -- still stalls), all vectorized paired-single arithmetic (disabled -- still stalls), the interrupt-callback list itself (dumped in the stalled state: 4 nodes, well-formed, identical to healthy), delivery eagerness (scoped -- C/AOT cured, new/SIM not), edge-gated delivery (no effect).
    • Lockstep float divergences at the sound-engine entries (0x8016F474 cluster) persist even with all vector paired paths disabled, so their origin is elsewhere (psq quantized load/store or scalar float lowering are the remaining candidates); they may still be symptoms rather than cause.

    Repro assets: states/pre-stall.sav (stalls any backend instantly), part1-in-town.sav (new/SIM stall ~15 frames after load, capped or uncapped). Next probe queued: run the C arm from part1-in-town.sav saving a state every second, find the last state from which the new arm still passes, then RAM-diff consecutive states across the new arm's execution to catch the first corrupted structure and its writer.

  16. dougchansan commented on Aug 24, 2026

    @dougchansan
    ContributorAuthor

    RAM-level investigation: the stall is a phantom thread, and the goal metric is met by the AOT arm

    LLVM is now faster than C on Colosseum wherever it runs correctly: the AOT arm measures 56.5 vs C's 53.5 uncapped on part1-in-town.sav and 43-47 vs 42 on the Phenac scene, clean in every run. The remaining defect is specific to the new-base emitter (plain and state-in-memory), and this session pinned down what actually breaks.

    What the corruption is

    Thread-state dumps in the stalled state:

    healthy stalled
    OSCurrentThread (0x800000E4) 803FBCB0 (main, state RUNNING) 8048E010 -- an address inside the main thread's stack
    real threads 803FBCB0/803FB998 RUNNING / READY both READY, never scheduled again

    A change tracker on 0x800000E4 caught the write: 803FBCB0 -> 8048E010, dispatch site = 0x800A1930 -- the scheduler's own thread-switch store (stw r30, OSCurrentThread). The scheduler selected the phantom from a run queue, meaning something upstream enqueued a stack address as a thread -- the signature of OSWakeupThread running against a wait object that lived on a stack frame which had already been unwound.

    The phantom's garbage context then executes: depending on module timing the visible result is the hard wedge (frame 176641), or a 0.35-0.5 fps crawl where interrupts sit pending and unmasked (pi_cause=0x548 under pi_mask=0xFFC, EE set) while the guest runs garbage state -- the same crawl signature as the original retrace-timeout bug, which is why the presentations looked related.

    Ruled out this session, each by direct experiment

    • OS interrupt masks (0x800000C4/C8) and PI mask: byte-identical healthy vs stalled.
    • Run-queue heads and both threads' link fields: polled every dispatch from 2.5M on -- never hold a stack-range value at a dispatch boundary; the enqueue-and-consume happens inside a single native burst.
    • Timebase staleness (volatile, uncache mftb reads): no effect.
    • Inline paired-single FMA, and all vectorized paired-single arithmetic: disabled wholesale, still stalls.
    • Delivery policy (eager / edge-gated / post-mtmsr-scoped): changes which failure mode appears, never prevents the corruption.

    Tooling landed along the way

    • ModernGekko PR LLVM-backend module stalls in the audio interrupt path; C backend unaffected #30: DOLRECOMP_LLVM_WRITE_JOURNAL was missing from the module reuse identity, so toggling it silently reused a non-journaled module.
    • The module RAM write journal works for watchpoints, but the journaled module's timing shifts the race into the crawl mode instead of the phantom, so the write of the phantom pointer has not been caught red-handed yet.

    The remaining hunt, precisely scoped

    Find the guest code that, under new-base codegen only, wakes a stale on-stack wait object at the scene transition. Best next steps: (1) a burst-boundary ring of the last N dispatch sites, dumped when the tracker fires, to identify the burst that performed the enqueue; (2) diff the new emitter against the AOT emitter for the specific functions in that burst. Repro assets: states/pre-stall.sav (poisoned state, stalls every backend), part1-in-town.sav (+~15 frames to failure), tracker/probes in this thread's history.

  17. dougchansan commented on Aug 24, 2026

    @dougchansan
    ContributorAuthor

    The historical 60 fps work, and what the ring probe changed about the stall

    The dispatch-ring result: the "phantom" was never a phantom

    The ring capture at the corruption moment, walked backward, names the whole flow:

    1. An init sequence (801BDCC4 bl-table calling 800B85B4 repeatedly) -- scene loading.
    2. FP-unavailable exceptions (vector 0x800) around the music/event engine (801A6F78, 801B4B40).
    3. Twelve dispatches inside 8000550C -- which disassembles to memcpy's inner loop, called
      from 0x80006410 copying a 712-byte (OSContext-sized) block to 0x803A1700.
    4. 0x80006450+ then does: OSCreateThread(&thread_on_stack /* r1+120 */, ..., prio 1),
      OSResumeThread, OSJoinThread
      -- and installs a custom handler for exception Emitter fix series + Dolphin-differential test suite from a downstream runtime project #8
      (FP-unavailable) at 0x800064A0, which is just OSLoadContext(saved) -- a crash-recovery
      longjmp. 0x800064C4 is the game's crash-report printer.

    So OSCurrentThread = 0x8048E010 is the game's own scene-loader worker thread, whose OSThread
    legitimately lives on the caller's stack
    -- not corruption. The stall is Colosseum's crash
    guard catching a fault in that worker
    under new-base codegen: the ring's final entry shows the
    switch to the worker ending inside the crash-report path. C, AOT and old-LLVM run the worker
    without faulting. The open question is now precisely "what does the worker execute differently
    under the new emitter that trips the crash guard" -- with the crash window fully mapped
    (0x80006378-0x80006520 plus the worker entry).

    The 60 fps / 6-scenes work, located

    Three layers, in different states:

    1. The matching decompilation (C:\Users\douglaswhittingham\pkmn-colosseum): 88.05% fuzzy,
      6,786/8,603 functions matched, with hundreds of per-function promotion worktrees
      (phase3/4/5-close-* covering HSD graphics, MusyX audio, SDK, fight, field...). This is the
      scene-by-scene campaign the "targeted 60 fps on 6 scenes" memory refers to.
    2. The profiling distillate: games/pkmn-colo-recomp/patches/GC6E01-native-substitutions.csv
      -- the seven hottest guest functions from scene traces: PSMTXConcat (6.1M calls in one traced
      scene
      ), PSMTXCopy/Identity/MultVec, HSD_MtxScaledAdd, memset, and memcpy (3.1M calls, most
      through the L1 locked cache)
      . Verified today: no consumer of this CSV exists anywhere in the
      current toolchain
      -- no dolrecomp_native_* symbol appears in any built module. It is a
      worklist, not a shipped feature. (The original C:...\pkmn-colo-recomp project dir was emptied
      on Aug 20 during the move to G:.)
    3. The widescreen code patches (GC6E01-native-widescreen.csv): these ARE live -- verified
      baked into patched-main.dol inside every module we have benchmarked (word at 0x80005300 is
      the patched value). --no-mods never excluded them.

    Why the substitution list matters now

    It attacks exactly the costs the backend benchmarks measure and cannot remove:

    • The C arm's 7.7M hook-fallback instructions per run are dominated by the locked-cache memcpy
      path the CSV names.
    • PSMTX*/HSD are paired-single-heavy guest code -- the biggest per-scene CPU item.
    • Uncapped Colosseum today: C 0.90x realtime, AOT 0.94x. Native substitution of these seven
      functions stacks with any backend and is the credible route past 1.0x -- and the decomp already
      holds matched C sources for most of them, so the substitutes largely exist.
    • Bonus: substituting the FP-heavy PSMTX*/HSD functions removes guest FP from exactly the code the
      scene-loader worker runs, which may sidestep the new-emitter crash window entirely.

    Attempting to run with the runner's mod layer enabled (dropping --no-mods) crawls both backends
    -- the runtime mod path is not how these were meant to be applied; the substitution mechanism
    needs to be built (link native implementations over the seven guest addresses at module build).

  18. dougchansan commented on Aug 24, 2026

    @dougchansan
    ContributorAuthor

    Correction: Colosseum freezes on every LLVM arm; Skyward Sword now works

    Correction to the earlier "Colosseum fixed at 28 fps" claim

    The mtmsr/EE-delivery fix did cure the retrace-interrupt starvation (0.25 fps -> the
    game renders and cold-boots past the THP intros). But from a savestate, every LLVM
    arm still freezes after a few seconds
    , and the fps averages reported earlier were
    measuring "how long until it froze", not sustained throughput.

    Evidence: uncapped 25 s runs, part1-phenac-nw.sav, repeated back to back.

    arm run 1 end frame run 2 end frame speed
    C backend 278262 278285 1.22
    LLVM, upstream main 277678 277678 0.06
    LLVM, pre-upstream fix build 277678 277678 0.06

    C reaches a different frame each run because it keeps running. Every LLVM arm halts
    at a byte-identical frame with the guest CPU at 6% of realtime. Reproduced on three
    states so far: part1-phenac-nw (277678), willie-battle (67679), part1-in-town
    (176641). Not upstream-specific -- the pre-upstream build freezes at the same frame --
    and not caused by native substitutions, which were not present in the control arms.

    This is the same phantom-thread / crash-guard defect tracked above, but its scope is
    wider than previously recorded: it is not limited to scene transitions, it happens
    from ordinary gameplay states. Colosseum LLVM performance numbers should not be
    quoted until this is fixed, because none of them represent sustained execution.

    Cold boot is the exception: a 150 s cold-boot run rendered throughout at ~55 fps, so
    the trigger is tied to savestate-restored state rather than to any scene reached
    by playing forward.

    Skyward Sword (SOUE01, Wii/Broadway) now runs on LLVM

    Previously recorded as producing no usable LLVM number at all. On current upstream
    main it renders correctly (verified by screenshot: full HUD, hearts, control
    prompts, textures) and runs freely -- frame counts advance differently every run,
    no freeze.

    arm run 1 run 2 speed
    LLVM 33.82 fps 35.08 fps 1.18
    C 43.87 fps 43.07 fps 1.50

    LLVM is at 0.79x of C here, so there is headroom to chase, but the title is no
    longer broken on the LLVM backend.

    Native substitutions: implemented, not yet measurable

    Seven profile-identified hot guest functions (PSMTXConcat/Copy/Identity/MultVec,
    HSD_MtxScaledAdd, memcpy, memset) can now be replaced with native implementations
    via DOLRECOMP_SUBSTITUTIONS=addr=symbol,.... The mechanism works end to end and is
    semantically correct -- the substituted Colosseum module renders the reference scene
    pixel-identically to the unsubstituted one. Its performance effect cannot be measured
    until the freeze above is resolved.

    Two bugs found while building it, both worth noting for anyone extending the IR:

    • src/ir/dolir_builder.c had no <stdlib.h>, so a newly added getenv was
      implicitly declared as returning int; the truncated pointer segfaulted the
      recompiler. MSVC reports this only as warning C4013.
    • control_flow.cpp dereferences values_[term.condition] unconditionally for
      DOLIR_TERM_INDIRECT, so an unconditional indirect edge still needs a
      materialized constant-true condition. DOLIR_NO_VALUE crashes the emitter.
  19. dougchansan commented on Aug 24, 2026

    @dougchansan
    ContributorAuthor

    Decrementer starvation (real, unresolved) + Apple Silicon results

    The decrementer is never delivered from native code

    Colosseum's frozen guest spins at 800feb9c + 8009df80 -- the same pair that dominated
    the earlier lockstep run (728k divergences at entry=0x800FEB9C end=0x8009DF80, native
    MSR 0x9932 vs interpreter 0x1932). Disassembling 0x800FEA74-0x800FEB9C shows an
    infinite OS dispatcher loop: walk a callback list, call each ready callback, then
    under OSDisableInterrupts merge a pending list into it by priority, OSRestoreInterrupts,
    and branch unconditionally back to the top. It never returns; it only stops hogging the
    CPU when an interrupt preempts it.

    The runtime's asynchronous exception set is

    SYNC_EXCEPTION_MASK = ~(EXCEPTION_EXTERNAL_INT | EXCEPTION_DECREMENTER |
                            EXCEPTION_PERFORMANCE_MONITOR)
    

    but the delivery gate added for the retrace fix tests EXCEPTION_EXTERNAL_INT only. So
    the decrementer -- which drives GameCube OS alarms and thread timeslicing -- is never
    delivered while the guest runs native code.
    Titles driven by VI retrace (Mario Kart,
    Luigi's Mansion, Skyward Sword) are unaffected; Colosseum's dispatcher waits on
    alarm-driven preemption that never arrives. That is a concrete answer to "why only
    Colosseum".

    Widening the gate to (ppc.Exceptions & ~SYNC_EXCEPTION_MASK) measurably changes the
    failure: the previously deterministic halt (three states, always the identical end frame)
    became intermittent, and one run reached speed 1.05 -- full realtime with a fresh end
    frame.

    variant end frames across runs verdict
    external-int only (shipped) 277678, 277678, 277678 always frozen
    + decrementer, post-mtmsr gate 277865, 277454, 277678 intermittent
    + decrementer, any EE=1 boundary 277505 (speed 1.05), 277678, 277678 intermittent

    Not a fix on its own, and currently reverted in the tree, but it is a genuine gap and the
    most specific lead so far.

    Correction: the revert was based on a contaminated signal

    The widened gate was reverted because Skyward Sword appeared to drop to 0 fps. That
    reading was wrong. A background agent was regenerating SOUE01 modules for a separate arm
    matrix at the same time -- active-module.txt had been repointed from 8a49f3e1... to
    759fcc8d... underneath the benchmark, and a second emulator was competing for the GPU.
    Skyward Sword was fine throughout. The decrementer change must be re-tested on a quiet
    machine before any conclusion about it is drawn.

    Method rule this produces: never benchmark a title while another agent or session is
    generating or running that same title, and check the module hash in active-module.txt
    before and after a benchmark series.

    Regression check (concurrent load present, numbers not quotable)

    Both titles run and advance frames -- no freeze, no regression to the broken state:

    title LLVM C
    Mario Kart DD 41.3 / 44.3 fps 52.1 / 19.9 fps
    Luigi's Mansion 33.4 / 37.3 fps 52.8 / 48.6 fps

    The C 19.9 fps outlier shows the load contamination; these are health checks, not scores.

    Apple Silicon, upstream main (M5, 10 cores)

    Luigi's Mansion foyer.sav, Null graphics, headless, uncapped, interleaved blocks with a
    discarded warm run per arm.

    • LLVM median 81.18 fps (79.95, 81.15, 81.21, 81.45)
    • C median 79.10 fps (74.48, 77.68, 79.10, 79.27, 79.56)
    • Ratio 1.026x (1.016x restricting to the quietest campaign)

    Versus the previous Apple Silicon campaign's baseline of ~71.5 fps on this state, that is
    about +13.5%, and it replaces the earlier 0.93x regression with parity-to-slightly-ahead.
    No freeze: end frames varied every run (6678...7018), ~2400 frames per 30 s window.

    Caveats: the C control is a pre-existing module from an older dolrecomp revision, not a
    same-commit rebuild, because the C backend at 1bec355 emits calls to
    ppc_fp_available_inline which does not exist in that checkout's GXRuntime. The Mac is a
    shared, noisy host -- one campaign was discarded entirely at load average 15.5 -- so
    anything under ~5% there needs many more pairs.

    macOS build blocker found

    The LLVM backend does not build a module on macOS at 1bec355. Generation dies at
    chunk 57/4164:

    LLVM ERROR: Global variable 'fix_pair_nan_f64' has an invalid section specifier
    '.text.unlikely.fix_pair_nan_f64': mach-o section specifier requires a segment and
    section separated by a comma.
    

    markColdFixup() in src/backend/llvm/fp_fixups.cpp:21 sets an ELF-style section name
    unconditionally. Skipping the cold-section hint on Mach-O targets lets all 4164 chunks
    generate. Worth fixing upstream.

  20. dougchansan commented on Aug 24, 2026

    @dougchansan
    ContributorAuthor

    Colosseum: the freeze is a guest crash dump, not a scheduling bug

    What the guest is actually doing

    Instrumenting the stuck state and dumping the buffer the guest is writing settles it.
    The LLVM arm parks in the C runtime's line-buffered stdio write loop
    (0x800C76F0-0x800C77E0, which calls memchr(buf, '\n', n) at 0x800C811C), and the
    bytes it is emitting are:

    r26=80313900 r29=0000001A r30=80310A14 r31=00000001 msr=00001030
    text=Non-recoverable Exception %d....Unhandled Exception %d...
         DSISR = 0x%08x   DAR = 0x%08x.....TB = 0x%016llx...
         Instruction at 0x%x (read from SRR0) attempted to acce[ss]...
    

    r29 = 0x1A = 26 bytes, exactly the length of "Non-recoverable Exception ". That is the
    GameCube SDK's unhandled-exception dump, and msr = 0x1030 (EE clear, FP clear, ME/IR/DR
    set) is exception context.

    So the LLVM arm takes a guest exception the OS treats as unhandled, and then hangs
    inside its own crash dump
    , whose output never drains, looping forever with interrupts
    disabled. Everything previously described as "the freeze" is that hang. The apparent
    scheduling symptoms were downstream of it.

    What this retires

    • The decrementer theory is dead. Counters over a stuck run: dec_seen=30,
      dec_del=22 against ext_seen=393828. The decrementer is essentially never pending, so
      it cannot be what starves the guest. The static check that motivated it was also wrong:
      Colosseum, Mario Kart and Luigi's Mansion all have identical decrementer usage
      (2 mtspr DEC, 1 mfspr DEC -- the same SDK leaves), so "only Colosseum uses alarms" was
      never true. Branch fix/async-exception-delivery holds the widening; it is a dead end
      for this bug and should not be merged on these grounds.
    • The phantom-thread reading is superseded. OSCurrentThread pointing into a stack was
      the game's own scene-loader worker; the crash guard around it is the SDK error path, not
      corruption.

    The live signal

    Blocked-interrupt PC sampling, same stuck run, is unambiguous:

    arm ext_seen ext_del blocked (EE=0) top EE=0 site
    LLVM 393,828 8,933 384,903 (98%) 800c77e0 -- 1476 of 1503 samples
    C 79,255 30,353 48,970 (62%) 80164f1c -- 88 samples; 800c77e0 absent

    C never blocks at 800c77e0 at all. The site is the crash-dump loop, so this is simply
    another view of the same hang.

    What is still unknown

    Which exception, and at what SRR0. Three probes failed to catch it and the reason is
    instructive: the exception vectors live in low RAM, copied there at boot, so they are not
    part of the recompiled DOL, are not natively dispatchable, and never appear either as a
    native dispatch target or as ppc.pc at the sampling points tried. A probe that catches
    this has to hook the fallback/interpreter path, or the runtime's ppc_take_exception
    inside the module (module-side, so it needs a regenerated module).

    That is the next step, and it is now a narrow one: name the exception and its faulting
    instruction, and the defect in the generated code follows directly.

    Method notes worth keeping

    • The exception vectors are not in the DOL. Any probe that samples only native dispatch
      will miss all guest exception entry.
    • frame_count equality across runs is good evidence of a hang but not proof: on Skyward
      Sword two healthy runs matched by coincidence. Confirm with present_count and state=.
    • A module's directory hash is derived from dolrecomp_binary_sha256, not from emitter
      options, so "the hash changed" does not prove a codegen flag took effect. Diff an actual
      emitted chunk instead.
  21. dougchansan commented on Aug 24, 2026

    @dougchansan
    ContributorAuthor

    Module-side exception probe: every exception taken is benign

    Instrumenting ppc_take_exception inside the module (the only place that sees guest
    exception entry, since the vectors are not natively dispatchable) and regenerating the
    Colosseum LLVM module gives the first 10 exceptions of a stuck run:

    vector=0xC00 exception=0x08 srr0=0x80098038   (system call -- sc at 0x80098034)
    vector=0x800 exception=0x20 srr0=0x80165330   (FP unavailable, lazy FPU)
    vector=0x800 exception=0x20 srr0=0x8016CBB0   (FP unavailable)
    vector=0xC00 exception=0x08 srr0=0x8009B2FC   (system call -- DCFlushRange)
    ... same three kinds repeating
    

    No DSI (0x300), no ISI (0x400), no alignment (0x600), no program (0x700). Only system
    calls and lazy-FPU traps, both entirely normal. So within the observation window the guest
    takes no faulting exception at all, yet it is simultaneously emitting
    "Non-recoverable Exception " (26 bytes, matching r29=0x1A exactly).

    Two possibilities remain, and they are cheap to separate:

    1. The crash-dump write begins before the first dispatch of the window -- i.e. immediately
      on savestate restore -- so the faulting exception is never observed. Against this: the C
      backend restores the same state and runs normally, so the state itself is healthy.
    2. Both arms enter this print path, but the C arm's write drains and the LLVM arm's does
      not. The loop retries via bl 0x800C7454 at 0x800C77B0 and only exits when that
      returns non-zero, so a flush that never reports progress would hang exactly here while
      the same code on C completes. This would make the hang a stdio/flush-path defect rather
      than a crash at all, and would explain the absence of any faulting exception.

    Distinguishing them needs the buffer probe run against the C module as well: if C also
    writes "Non-recoverable Exception " and proceeds, possibility 2 is confirmed and the
    investigation moves to 0x800C7454 and what it depends on.

  22. dougchansan commented on Aug 24, 2026

    @dougchansan
    ContributorAuthor

    Colosseum: the divergence is pinned to one point

    C never enters the crash path at all

    A host-side probe counting entries into the stdio write loop (0x800C76F0-0x800C77E4),
    run on both arms from part1-phenac-nw.sav:

    arm write-loop entries text being written
    LLVM many; parks there "Non-recoverable Exception " (26 bytes, r29=0x1A)
    C zero --

    So the earlier "maybe both arms print it and only C drains" theory is dead. Only the LLVM
    arm reaches the unhandled-exception path.

    The path into it

    Dispatch ring captured at the first write-loop entry, oldest first:

    80164F1C 800C4824 801653AC 80164478 8015D714 80162254 8009E758 800AE9FC
    80163F98 800AEBD8 8009E724 8009DF5C 801643C8 8009A190 8009A1E0
    80006378 80005530 8000642C 8009C6F0 8009A190 8009A1E0 800C8864
    

    MusyX audio code (0x8015xxxx-0x8016xxxx) -> context save (8009A190/8009A1E0) ->
    the game's error handler (80006378 -> 80005530 memcpy -> 8000642C) -> crash printer.

    The exceptions are legitimate

    Module-side probe on ppc_take_exception (the only place that sees guest exception entry,
    since the vectors live in low RAM outside the recompiled DOL). Both FP-unavailable SRR0
    values point at real FPU instructions, so the traps and their reported addresses are
    correct, not corrupted:

    0x80165330  lfs   r0,-25008(r2)
    0x8016CBB0  fsubs f2,f23,f26
    

    The divergence point

    First twelve exceptions, same savestate, same runtime binary:

    # C arm LLVM arm
    1 0xC00 srr0=80098038 0xC00 srr0=80098038
    2 0x800 srr0=80165330 0x800 srr0=80165330
    3 0x800 srr0=800A3250 0x800 srr0=8016CBB0
    4 0xC00 srr0=80098038 0xC00 srr0=80098038
    5 0xC00 srr0=80098038 0xC00 srr0=80098038
    6 0xC00 srr0=8009B2FC 0xC00 srr0=8009B2FC
    7-12 identical identical

    The two arms are in lockstep through exception #2 and diverge at #3: C's next FP trap is in
    the SDK matrix library (0x800A3250, inside the PSMTX* range), the LLVM arm's is in MusyX
    (0x8016CBB0). Everything before that point matches instruction-for-instruction as far as
    this probe can see.

    That is a narrow, well-defined window: between the FP-unavailable trap at 0x80165330
    and the following trap. Whatever the LLVM backend does differently happens inside it. The
    next step is lockstep restricted to that window -- STATICRECOMP_LOCKSTEP_START set just
    before exception #2 with _NODEDUP=1 -- rather than the whole-run lockstep that previously
    drowned in step-cap artifacts.

    Housekeeping

    GXRuntime/src/core/cpu_exception.c carries a temporary probe; both GC6E01 modules were
    regenerated with it and print to stderr on every guest exception. Revert the file and
    regenerate both modules before quoting any benchmark from them.

  23. dougchansan commented on Aug 24, 2026

    @dougchansan
    ContributorAuthor

    Windowed lockstep, and a direct test of the block it names

    Divergence pinned to a 244-dispatch window (LLVM exc#2 at dispatch 18992, exc#3 at 19236;
    the C arm takes 2074 dispatches over the same stretch). Lockstep aimed at it
    (_START=18900 _LIMIT=19400 _NODEDUP=1 _STEPCAP=200000) gave 258 divergences dominated by
    one repeating pair:

    entry=0x8016CB7C end=0x8016C0B0:
      r5:N=0xcc010000,I=0x4   r6:N=0x0,I=0x803fc860
      r7:N=0x0,I=0xff0000     r8:N=0x1,I=0x9000580
      mmio#:N=27,I=18  writes at 0xCC008000
    

    0x8016CB7C is write-gather-pipe code, and 0x8016CBB0 -- where LLVM's exception #3 fires
    -- is inside it:

    8016CBB0  fsubs f2,f23,f26
    8016CBB4  lis   r3,0xcc01
    8016CBC0  stfs  f2,-32768(r3)     ; 0xCC008000, the GX write-gather pipe
    

    The MMIO count mismatch is probably not real: emitGuestStore issues exactly one write per
    path (MEM1 / MEM2 / external) before joining, and the lockstep shadow interpreter is known
    to re-run blocks and double-issue side effects.

    Direct test, independent of lockstep -- forcing that range to the interpreter:

    configuration speed end frame
    LLVM, unmodified 0.06 277678 every run
    LLVM, 0x8016c000-0x8016d000 interpreted 0.84 - 0.88 277433 every run

    CPU throughput goes from 6% of realtime to ~87%, so the block is implicated, but frames
    still do not advance -- the failure moves rather than clearing. Caveat: forcing ranges to
    the interpreter has hung this title on its own before, so mixed-mode is not a clean
    instrument.

  24. dougchansan commented on Aug 24, 2026

    @dougchansan
    ContributorAuthor

    Resolved: Colosseum runs at full speed on the LLVM backend. The failure was a modified DOL.

    The extracted game this investigation was run against, games/pkmn-colo-recomp/extracted/Pokemon-Colosseum-USA, has a widescreen mod baked into main.dol:

    address clean what we were benchmarking
    0x800A3930 FFA01090 (fmr f29,f2) 4BF619D0 (b 0x80005300)
    0x80005300 00000000 C3A2B084 (lfs f29,-20348(r2))
    0x80005304 00000000 EFBD00B2 (fp single op)
    0x80005308 00000000 4809E62C (b 0x800A3934)

    A 16:9 hook that diverts the projection-matrix path into hand-injected FP code placed in former zero padding. It renders visibly wrong -- the boot logo is squashed and menus collapse into a vertical strip -- and it sits in exactly the FP/matrix region where this issue kept finding "divergences". Two other extractions in the same tree (Colo-Fresh, Colo-PiDOL) are clean.

    Measured against a clean DOL

    Colo-Fresh, with fresh savestates captured on that DOL (the previous states were captured on the modded binary, so their RAM carried the patched code and were never valid against a clean module). Interleaved, warm, uncapped:

    savestate plain LLVM LLVM + state-in-memory C backend
    title 49.26 61.22 86.45
    early 46.93 61.01 92.16
    mid 49.76 64.48 86.64

    state-in-memory is worth 1.24x - 1.30x on this title, and its uncapped speed is above realtime in every state (1.006 / 1.024 / 1.078). Capped, it holds the frame rate exactly:

    savestate fps vps speed
    title 60.07 60.05 1.001
    early 60.12 60.11 1.002
    mid 59.94 59.94 1.000

    Rendering verified by screenshot: correct aspect, correct text, no distortion. No freeze in any state, on any arm.

    C is still ahead uncapped (LLVM/C ~0.66 - 0.74 with state-in-memory), so there is optimization headroom, but it no longer affects playability -- the title is 60 Hz content and the LLVM backend sustains it.

    What this retracts

    Everything in this issue measured against Pokemon-Colosseum-USA is unreliable. Specifically retracted: the deterministic freeze and its "identical end frame" evidence; the 0.25 fps retrace-timeout analysis; the interrupt starvation counters (blocked_ee ratios, raised/delivered); the __OSUnhandledException crash-dump finding; the phantom-thread reading; the decrementer theory; and the 0x8016CB7C gather-pipe divergence. Some of those may describe real effects of the injected code, but none of them describe the shipping game.

    ExpansionPak/RecompCore#18 has been withdrawn for the same reason -- a clean DOL runs at full speed with no delivery fix present at all.

    What remains worth doing

    • Low priority, real: on the modded DOL the C backend copes with the injected code at 0x80005300 and the LLVM backend does not. That is a genuine backend difference in handling code placed in former padding, and deserves its own issue if anyone wants the widescreen mod to work.
    • Optimization: closing the remaining LLVM/C gap on this title.

    Method changes adopted

    • Benchmark only from cleanly extracted, unmodified DOLs; verify main.dol against the untouched extraction first.
    • Savestates are only valid against a module built from the same DOL.
    • moderngekko-port applies Data/Sys/GameSettings/<id>.ini [OnFrame_Enabled] patches with no opt-out flag; for GC6E01 that is the memory-card patch. The widescreen hook came from the extraction itself, not from that mechanism.
  25. siahisaforker commented on Sep 5, 2026

    @siahisaforker
    Contributor

    sorry doug but i aint reading allat, good luck on your recomp lmao

  26. dougchansan commented on Sep 5, 2026

    @dougchansan
    ContributorAuthor

    You're good. This is just tracking. The recomp is getting closer.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions