Repository navigation
LLVM backend: Colosseum does not get through boot (x86-64 and aarch64; C backend does) #20
Description
Activity
- changed the title
[-]AArch64: fixed-chunk LLVM backend hangs Colosseum during boot (C backend runs)[/-][+]LLVM backend: Colosseum does not get through boot (x86-64 and aarch64; C backend does)[/+]on Aug 22, 2026 Further evidence, and one hypothesis eliminated.
The DOL was not the cause. Our Windows workspace had been recompiling a modified Colosseum DOL (
md5 b299467dd305) while the Pi used the retail one. I re-extracted from the retail disc image with DolphinTool (md5 c55df119fad4, matching the Pi) and rebuilt every arm from that. The defect is unchanged.Cold boot, x86-64, retail disc data, identical harness per arm:
arm fps vps speed frame_count present_count C 59.94 59.94 1.000 5353 10755 LLVM (upstream rewrite) 0.246 59.94 0.999 8 1575 LLVM (AOT region backend) 0.250 59.94 0.999 8 728 Two details that sharpen the picture:
- The guest runs at full speed (
speed≈ 1.0) with the video interrupt ticking normally (vps59.94). This is not a slow module — it is a module that does not render. present_countkeeps rising whileframe_countis frozen at 8. The runtime is still swapping buffers; the game is simply re-presenting the same image because its main thread never advances.
It also ruins a mid-game savestate scene, not just boot. Benchmarking
part1-in-town.savwith the interleaved harness, all three LLVM backends are degenerate while C is healthy:arm samples fps median speed median C control 114 54.67 0.9159 LLVM old (pre-rewrite) 11 0.34 1.3250 LLVM new (rewrite) 17 0.58 2.1241 LLVM aot 26 0.85 3.1201 Worth flagging for anyone benchmarking this title: because
speedcounts retired guest cycles and this defect burns cycles in a spin that does no work, the metric rewards the broken arm. On speed alone the AOT arm reads as a 3.58x win over C. It is an artifact — fps and sample count expose it. I withdraw an earlier internal Colosseum figure that was measured this way.So the defect: affects all three LLVM backends (pre-rewrite, rewrite, and AOT), reproduces on x86-64 and aarch64, on retail data, at boot and from a savestate. The C backend is unaffected throughout.
- The guest runs at full speed (
Further investigation. No root cause yet, but three useful narrowings and one correction to something suggested along the way.
Correction first
A hypothesis was raised that the C backend avoids the bug because its text0 chunk fails SMC verification and the whole 0x80003100-0x80005500 range therefore runs on Dolphin's JIT rather than recompiled code. That is not what happens. Runtime counters from a C-backend run:
[staticrecomp] shutdown: native=1494658421 fallback=0 native_exc=78253 hook_fb=127384220 smc_failed=0 verifications=98 reverify_events=51.49 billion natively executed instructions, zero fallbacks, zero SMC failures. The C backend is genuinely running recompiled code for this title. Both backends report
fallback=0 smc_failed=0, so the difference is not one of them quietly running on the JIT.New: the divergence is visible in the runtime counters
Same title, same scene, same runtime, C versus LLVM:
counter C LLVM native instructions 549,841,950 181,820,064 native_exc38,659 521 hook_fb86,909,490 108,309 fallback0 0 smc_failed0 0 guest cycles 13,686,478,575 14,779,848,761 The LLVM arm retires a comparable number of guest cycles while taking 74x fewer exceptions and 800x fewer hook fallbacks. It is not doing the same work. That is a much sharper signal than the frame counter and it is cheap to reproduce.
New: the rendering ratio is 238:1
Measured directly: 238 VI interrupts fire per rendered frame under LLVM, against 1 per frame under C. So the game's render-trigger condition is being satisfied roughly once per 238 VI callbacks rather than every one. That is consistent with the counters above — something the game does every frame is not happening.
Ruled out this round, with evidence
The VI exception stub at
0x800039ACwas inspected in the emitted IR and is correct:- SPRG1/2/3 saves emit
ppc_mtspr(ctx, 273/274/275, ...)with the right SPR numbers - SRR0 set to
0x800C0EAC, SRR1 toold_msr | 0x30 gpr[2]= interrupted PC,gpr[3]= 0x500,gpr[4]= srr1ppc_rfitail call with the correct CIA- The alignment stub at
0x80003AAChas the same structure with exception type 0x600
Two earlier hypotheses were also refuted on inspection:
0x7C5143A6at0x800039ACdecodes asmtspr SPRG1, r2(correct), and themfmsr/oripair lowers correctly toor i32 %117, 48.Not yet examined
The FPU-unavailable stubs at
0x800044AC,0x800045AC,0x800046ACare structurally unlike the others — they save and restore the Condition Register around the handler (mfcr/andi./sync/mtcrf 0xFF/xoris). A CR save/restore error there would corrupt condition codes across an exception and change branch outcomes in game code, and the FPU-unavailable exception fires during startup before the FPU is enabled. That is the next place worth looking, though it is a suspicion rather than a finding.Happy to keep digging or hand this over. The counter comparison above is the cheapest reproduction of the divergence I have found so far.
- SPRG1/2/3 saves emit
More probing. Not resolved, but the divergence now has a number and a direction.
The sharpest signal: the LLVM arm barely takes any exceptions
Runtime counters, same title, same scene, same runtime, comparable guest cycles:
counter C LLVM guest cycles 13,686,478,575 14,779,848,761 native_exc38,659 521 hook_fb86,909,490 108,309 bursts2,712,107 1,132,575 fallback/smc_failed0 / 0 0 / 0 The LLVM module retires slightly more guest cycles while taking 74x fewer exceptions. External interrupts are how VI reaches the guest, so a title that renders nothing while running at full speed with the host VI ticking at 59.94 is consistent with interrupts not being delivered rather than with slow code.
That also matches the independently measured 238 VI interrupts per rendered frame.
What the working arm actually spends its time on
The runtime has a built-in dispatch sampler (
STATICRECOMP_DISPATCH_SAMPLES=1) that records dispatch PCs and dumps the top sites at shutdown. On the C backend the top two are:dispatch-site pc=8009df3c samples=51999 dispatch-site pc=8009df64 samples=515760x8009DF3Cdisassembles to:8009DF3C mfmsr r3 8009DF40 rlwinm r4, r3, 0, 17, 15 ; clear MSR[EE] 8009DF44 mtmsr r4 8009DF48 rlwinm r3, r3, 17, 31, 31 ; return previous EE 8009DF4C blrThat is
OSDisableInterrupts(), withOSRestoreInterrupts()immediately after it at0x8009DF50. The busiest thing the working backend does is enter and leave critical sections — which is exactly the machinery that governs whether interrupts get delivered.Note the LLVM arm could not be sampled the same way: it never processes the automation stop command (waited 40s), so its counters never flush. That inability to shut down cleanly is itself consistent with the emulation loop being stuck.
Ruled out this round
mtmsrlowering is not the difference. Both backends do the same thing:- C:
ctx->msr = ctx->gpr[rS] - LLVM:
write_slot(DOLIR_STATE_MSR, gpr(rS))via the IR builder
Keeping guest state in memory does not fix it. I built the title with the
--state-in-memorychange from #21, which removes the entry prologue, the materialization barriers and the post-call reloads, so MSR is written toCPUStateimmediately rather than held in a promoted SSA value. If delayed visibility of MSR[EE] were the cause, that should have changed the outcome. It does not:arm fps frame_count C 59.94 5353 LLVM (current) 0.246 8 LLVM + --state-in-memory0.249 8 Identical failure. So the defect is not related to state hoisting.
Gather pipe is not obviously implicated. The static recomp core routes
0xCC008000page writes straight toGPFifo, andGPFifoManager::Write32self-kicks viaCheckGatherPipe().Where I would look next
The
native_excgap is the cheapest reproduction of the divergence and does not need a benchmark run — boot either module for 70s and read the shutdown line. Anything that explains why the LLVM module takes 521 exceptions where C takes 38,659 over the same workload very likely explains the rendering failure too.The suggestion that this title is stuck on its THP intro video is consistent with everything above, though I have not confirmed it directly: a video that never advances would look exactly like this if the interrupts driving it are not arriving.
- C:
Ran the runtime's lockstep verifier against this title, and then ran the same thing against a title that works. Posting both because the second invalidates the first as evidence.
Lockstep on Colosseum looked like a smoking gun
Built the module with
DOLRECOMP_LLVM_WRITE_JOURNAL=1and ran withSTATICRECOMP_LOCKSTEP=1. It diverges from Dolphin's core at the first checked region:DIVERGE #1 entry=0x80003158 end=0x80005498: r3: N=0x6f05 I=0x6f09 r4: N=0x8039a77c I=0x8039a6fc xer: N=0x20000000 I=0x0 mem[0x8048e722]: N=0x20 I=0x000x20000000is XER[CA]. r3 is 4 lower and r4 is 0x80 higher, which is exactly 4 iterations of the0x20-stride block-fill loop at0x80005498. Reproducible, and unchanged when the step cap is raised from 512 to 65536, so it is not a comparison-window artefact.But a working title diverges identically
Mario Kart Double Dash renders fine under the same LLVM backend (50.0 fps, 5271 frames in the same 90s window). Its lockstep output:
DIVERGE #1 entry=0x80003158 end=0x80003500: mem[0x803e22e2]: N=0x20 I=0x00 mem[0x803e22e3]: N=0x32 I=0x00 DIVERGE #2 entry=0x80003500 end=0x80003500: r3:N=0x3288,I=0x32a1 r4:N=0x803642dc,I=0x80363fbcSame entry PC, same byte values (
0x20,0x32) at the analogous address, same block-fill loop signature with r3/r4 offset by a whole number of iterations. 25 divergences reported for each title.So lockstep divergences on this codebase are not by themselves evidence of a miscompile. They appear on a title that works perfectly. The comparison runs native code a region at a time against a stepwise interpreter, and the two are simply not sampled at equivalent points inside tight loops. Anyone using this tool here should calibrate against a known-good title first — I nearly reported the XER carry difference above as the root cause.
What still distinguishes the two titles
The counter gap remains the only signal I have that separates a working title from this one:
counter C LLVM guest cycles 13,686,478,575 14,779,848,761 native_exc38,659 521 hook_fb86,909,490 108,309 74x fewer exceptions over slightly more guest cycles, with the host VI ticking normally at 59.94. That is still the cheapest reproduction and still the most likely place the answer lives.
Also confirmed this round: the module is genuinely loaded and executing recompiled code, not silently falling back.
native=181820064 fallback=0 smc_failed=0, module hash matches the retail DOL.Broke the exception counter down by vector. This is the sharpest split so far.
Exceptions by vector, same title, same scene, 70s
vector C backend LLVM backend 0x800FP unavailable24,488 1 0xC00system call59,512 19 total 84,000 20 Not a scaling difference — the LLVM module takes twenty exceptions where the C module takes eighty-four thousand, over comparable guest cycles at full speed.
The first 87 exceptions are identical
Logging the vector and SRR0 of every exception and diffing the two sequences: they match exactly for 87 exceptions, then diverge. Both arms then continue hitting the same two
scsites, but the LLVM arm's rate collapses (314 logged versus the 5000 cap for C).The two system-call sites are both cache/sync:
8009B2E8 mtctr r4 8009B2EC dcbf 0, r3 ; DCFlushRange loop 8009B2F0 addi r3, r3, 0x20 8009B2F4 bdnz 0x8009b2ec 8009B2F8 sc ; SRR0 = 8009B2FC 8009B2FC blr 80098034 sc ; SRR0 = 80098038, bare sync call 80098038 blrTested and ruled out this round
Cache-op handling is not it.
dcbfis one of the few places the backends genuinely differ: the C backend emitsppc_fallback_instructionand leaves native execution for every one, while the LLVM backend handles it inline throughppc_cache_control. That accounts forhook_fbbeing 86.9M on C versus 108K on LLVM, and it was a good candidate for starving the runtime of re-entry points.Forcing exactly those two functions through the JIT with
STATICRECOMP_FALLBACK_RANGES=8009B2E0-8009B310,80098030-80098040changes nothing: 0.2498 fps versus 0.2467 unforced,frame_count8 either way.HookCacheControlis installed and correct — ICBI invalidates the iCache and the JIT, so cache operations are not silently dropped on the LLVM path.The number I cannot explain
FP unavailable: 24,488 versus 1.
The C backend emits a check before every FPU-using instruction (
if (ppc_op_uses_fpu(inst->op))→ppc_fp_available_inline). The LLVM backend checks once per region and caches the answer infp_available_checked_, which is reset at region entry and byreloadUsedState().I found and fixed two genuine cache-invalidation holes while chasing this (an
mtmsrwrite not invalidating the cache, and the native call-resume path reloading values without dropping the emitter's conclusions — both in #22). Neither moves this title: together they change one of 4890 emitted objects.So the caching holes are real but are not what suppresses 24,487 FP-unavailable exceptions. Something else is preventing the LLVM module from ever taking that exception, and the guest's lazy FPU context switching is built on receiving it. That seems the most promising thread for anyone with more context on this backend than I have.
Everything eliminated so far
SMC/JIT fallback (both run fully native,
fallback=0 smc_failed=0), entry dispatch,rfi/thread resume (CPUState correct at every resume),mtmsrlowering, state hoisting (--state-in-memoryis byte-for-byte the same failure), gather pipe, cache ops, DOL variant (retail disc data reproduces it), and lockstep divergences as evidence (a working title shows the same ones).Correction to my last comment, and a much better number.
I suggested the FP-unavailable disparity (24,488 versus 1) was the most promising thread. It is a symptom, not a cause. Probing MSR at every
rfishows why.MSR[FP] at thread resume
C LLVM rfiresumes in 60s~800,000 91,799,710 MSR[FP] set 742,467 91,799,710 MSR[FP] clear 57,533 (7.7%) 290 (0.0003%) ppc_fp_availablehelper calls0 1 The C arm resumes threads with FP disabled 7.7% of the time — ordinary lazy-FPU behaviour, and each of those is an FP-unavailable exception waiting to happen. The LLVM arm resumes with FP disabled 290 times out of 91.8 million.
So the LLVM module never takes FP-unavailable because it never runs a thread that lacks the FPU. That follows directly from the real problem: it is performing 115x more thread resumes than the C backend while doing almost no work.
What that rules out
The FP-availability caching in the LLVM emitter is not what suppresses those exceptions. I found and fixed two genuine invalidation holes there (#22) and they change one of 4890 emitted objects on this title — consistent with this measurement, not with them being the cause.
Also ruled out this round: forcing the
DCFlushRange/sync functions through the JIT withSTATICRECOMP_FALLBACK_RANGES=8009B2E0-8009B310,80098030-80098040changes nothing (0.2498 fps versus 0.2467 unforced). So the inline-versus-fallback handling ofdcbfis not it either, despite being one of the few genuine behavioural differences between the backends (hook_fb86.9M on C versus 108K on LLVM).Where this leaves it
The guest is stuck resuming a thread ~92 million times in 60 seconds, at full speed, taking essentially no exceptions and making 19 system calls where the C backend makes 59,512. Every resume carries the correct state — SRR0
0x800A183C,r3=1, which is the "resumed" return value the scheduler tests — and the emitted IR for that entry point is correct.So a thread is being resumed correctly and immediately blocking again, forever, without asking the OS for anything. What it is waiting on is still unknown, and that is the question I would put to anyone with more context on this runtime.
Cheapest reproduction remains the shutdown counter line after a 70s boot:
native_exc521 versus 38,659, no benchmark needed.Colosseum (GC6E01): what the blocked thread is waiting on
Pinned state
part1-phenac-city.sav, Windows, same runtime build for both arms.
Baseline: C = 29.98 fps, LLVM = 0.2487 fps, VI ~59.94 in both.1. The thread
Guest thread
0x803FBCB0, constant stack chain:803FBCB0 sp=8048E678 <- 804001F0 <- 800A19B0 <- 800D31E4 <- 801E115C <- 80005DE0 <- 800FE8080x800D31E4is the resume point of a 3-second timed wait:800D31C0 stb r0, -0x5d8f(r13) ; clear flag 800D31E0 bl 0x800a1990 ; block / yield 800D31E4 lbz r0, -0x5d8f(r13) ; resumes here, re-reads flag 800D31F4 lwz r0, 0xf8(r31) ; 0x800000F8 = bus clock 800D31FC srwi r0, r0, 2 ; ticks/sec 800D3200 divwu r0, r3, r0 ; elapsed seconds 800D3204 cmplwi r0, 3 800D320C stb r30, -0x5d8f(r13) ; timeout -> set flag itself 800D3218 beq 0x800d31e0 ; else keep waiting2. What sets the flag
Only one other writer,
0x800D3F50-- a 3-instruction leaf:800D3F50 li r0, 1 800D3F54 stb r0, -0x5d8f(r13) 800D3F58 blrRegistered at
0x800D3CA4via0x801C01C8, which stores the pointer into the
VI globals at0x80456DA8bracketed by OSDisableInterrupts/OSRestoreInterrupts
-- i.e. VISetPostRetraceCallback.So the thread waits for a VI retrace interrupt, and escapes only via its own
3-second timeout. 0.2487 fps == one frame per ~4.0 s == the timeout path. The
game is not slow; it is timing out once per frame.3. Why the retrace never arrives
Added counters to the raise (
ProcessorInterface::UpdateException) and delivery
(PowerPCManager::CheckExternalExceptions) paths:arm raised delivered blocked (EE=0) Colosseum C 33,037 31,209 37,158 Colosseum LLVM 484 462 1,056,463 MKDD LLVM (control) 28,263 25,198 3,494 raisedcounts only 0->1 transitions, so the low count is a consequence: the
interrupt sits pending because MSR.EE is stuck at 0. 30% of checks are
EE-blocked under LLVM vs 0.76% under C. MKDD under LLVM is healthy, so this is
not a general backend fault.4. Where it sits with interrupts disabled
Sampling guest PC on every 64th EE-blocked check:
PC C LLVM 800a128c(one-instructionblr-- null thread-switch hook)10 11,077 800a183c(afterbl 0x8009bbd0, in__OSReschedule)32 5,466 Matching dispatch-site sampling: the LLVM arm spends ~97% of dispatches at
800d31e4+800a128c/1930/183c. The C arm's top sites are OSDisableInterrupts
/OSRestoreInterrupts and ordinary game code.The guest is parked in the OS scheduler's idle path with interrupts disabled.
Ruled out
- Timebase hoisting. Made
mftb/DOLIR_STATE_TIMEBASEreads volatile and
non-cacheable (branchfix/timebase-volatile, module hash changed
625021ef -> e0e7c371). No effect: 0.2487 fps. The 3-second timeout does
fire, so the guest clock advances correctly. rfilowering.STATICRECOMP_FALLBACK_RANGES=8009bc50-8009bcc0forces
OSLoadContext (ends inrfi) to the interpreter. No change (0.2479 fps).
syncDirtyState()does flush SRR1 beforeppc_rfi, andreturnFromBody()
does not re-sync, so the restored MSR is not being clobbered.- Earlier rounds: SMC/JIT fallback (fallback=0 in both arms), entry dispatch,
mtmsrlowering, state hoisting, gather pipe, cache ops, DOL variant,
FP-available cache holes, lockstep divergence signatures.
Note
STATICRECOMP_FALLBACK_RANGESover8009df3c-8009df90
(OSDisableInterrupts/OSRestoreInterrupts) or over the whole OS region
80098000-800a4000hangs outright (fps=0, vps=0, delivered=0). Mixed
native/interpreter execution across MSR manipulation is itself broken, so this
is not a usable bisect tool in that region -- worth a separate look.Instrumentation to revert
Source/Core/Core/HW/ProcessorInterface.cpp(g_dr_ext_raised)Source/Core/Core/PowerPC/PowerPC_Exceptions.cpp(counters + dr_note_ee0)Source/Core/Core/PowerPC/StaticRecomp/StaticRecompCore.cpp(shutdown dump)
- Timebase hoisting. Made
Lockstep: an msr divergence unique to Colosseum
STATICRECOMP_LOCKSTEP=1, both titles LLVM, both capped at 74 reports:title divergences of which involve msrMKDD (working control) 74 0 Colosseum 74 1 The one Colosseum msr divergence lands in
OSDisableInterrupts:entry=0x8009E724 end=0x8009DF3C msr: N=0x9932 I=0x1032 srr1: N=0x9932 I=0x1030 srr0: N=0x8009df3c I=0x8009e758 r12: N=0x800a128c (the null thread-switch hook)Decoding: I = ME|IR|DR|RI, EE=0. N = same plus EE(0x8000), FE0(0x800),
FE1(0x100). The native side has taken an exception (srr0/srr1 populated) that
the interpreter has not, and the whole GPR file has diverged with it.FE0|FE1 being set on the native side ties this to the MSR/FP-exception-mode
path. Not conclusive on its own -- one report against a 74-report cap -- but it
is the only msr divergence across both titles and it is in exactly the function
that gates interrupt delivery.Lockstep, uncapped and filtered to
msrThe report cap was never the limit --
STATICRECOMP_LOCKSTEP_MAXREPORTalready
defaults to unlimited. The real bound ism_ls_checked, which dedupes by entry
PC, so "74 divergences" meant 74 distinct diverging entry PCs. Added two
options to get past it:STATICRECOMP_LOCKSTEP_FILTER=<substr>-- only report diffs containing the
substring (suppressed reports counted, not printed)STATICRECOMP_LOCKSTEP_NODEDUP=1-- re-check entry PCs instead of once each
Also raised
STATICRECOMP_LOCKSTEP_STEPCAPto 100000; most of the original 74
wereCTRLFLOW:...steps=512artifacts of the interpreter running out of steps.Result: 44 msr divergences, 42 of them one bit
180 s run, Colosseum LLVM,
part1-phenac-city.sav:count native interpreter delta 42 0x99320x19320x8000 = EE 1 0x99320x1032EE + FP 1 0x39320x1030-- Entry PCs:
count entry end 21 0x800FEA980x800FEA9821 0x800FEA980x800FEA901 0x800B8E280x800A128C1 0x8009E7240x8009DF3C42 of 44 come from one loop.
0x800FEA90-0x800FEAE4walks a linked list at
r13-0x5B88andbctrls each node's handler at offset 20, then calls
OSDisableInterrupts-- a callback dispatcher running in interrupt context,
where EE should be 0. The interpreter agrees; native has it set.Why that is plausible as a bug
The OS interrupt primitives are all mfmsr/mtmsr pairs:
8009DF3C mfmsr r3; rlwinm r4,r3,0,17,15; mtmsr r4 ; OSDisableInterrupts 8009DF50 mfmsr r3; ori r4,r3,0x8000; mtmsr r4 ; OSEnableInterrupts 8009DF64 cmpwi r3,0; mfmsr r4; ...; mtmsr r5 ; OSRestoreInterruptscpu->msris written from outside generated code (the runtime vectoring an
exception, orppc_rfi). Ifmfmsrreads an SSA-promoted copy hoisted at region
entry, the followingmtmsrwrites the stale value straight back and clobbers
MSR[EE]. Same bug class as the FP-available cache and the timebase.Fix attempted, and its result
Branch
fix/msr-in-memory:slotInMemory()returns true forDOLIR_STATE_MSR
unconditionally, so MSR always lives inCPUStateand is never SSA-promoted.
Module regenerated (hash94f97ee7).It does not fix Colosseum. Three runs: 0.4126, 0.2487, 0.2487 fps. The first
reading was an outlier; the reproducible value is 0.2487, identical to before
the change.blocked_eeessentially unchanged (1.013M vs 1.058M).The divergence signature did change -- the EE-only pattern is no longer dominant;
aN=0x3932,I=0x1030pattern (native has FP/FE0/FE1/RI set where the interpreter
is in exception-entry state) now accounts for 42011 of 42360. Whether that is a
genuine second bug or an artifact of the shadow interpreter is not established.Caveat on lockstep as evidence
A Mario Kart lockstep run crashed with
Invalid read from 0x00000054at
PC=0x801a1bc8and threw modal dialogs. The shadow interpreter re-executes
blocks and can double-issue side effects, so lockstep output in this region is
suggestive, not authoritative. (Harness scripts now write
UsePanicHandlers = Falseso runs cannot raise blocking dialogs.)Exception histogram by vector (new, and the sharpest result so far)
Added a per-vector histogram in
ppc_take_exception(GXRuntime/src/core/cpu_exception.c)
and regenerated both Colosseum modules. Same scene, same savestate:arm total vec 0x800 (FP unavailable) vec 0xC00 (system call) C 19,858 9,297 10,561 LLVM < 32 -- -- The C arm raises ~20k guest exceptions of exactly two kinds. The LLVM arm raises
fewer than 32 in total -- the histogram never even reached a dump threshold of 32.
Bothscand FP-unavailable are correctly lowered in the LLVM backend
(src/backend/llvm/control_flow.cpp:202callsppc_system_call_exception;
emitFPAvailableinspecial_instructions.cpp), so this is not missing codegen --
the guest simply never reaches that code.Consistent with this, the one non-EE lockstep msr pattern is
N=0x3932, I=0x1030:
native has MSR[FP] (0x2000) set where the interpreter has it clear. A guest with
FP permanently enabled never traps, so the OS lazy-FPU context switch never engages.
Note the FP-available cache fix did not change this, so the suspicion is MSR[FP]
itself being wrong in native state rather than the cached answer.Thread scheduling is identical
Tracing
OSCurrentThread(0x800000E4) transitions from the run loop, both arms
alternate between exactly the same two threads --803FBCB0(the blocked one) and
803FB998-- in the same pattern. The LLVM arm just switches ~8x less often. No
thread is starved out of the run queue; the structure is the same.The framerate is bimodal
Root cause found and fixed: interrupt-delivery starvation by burst phase-locking
The exception histogram was the right thread to pull, but it was a symptom. The chain, fully traced:
- Only 3
scinstructions exist in the whole DOL --0x8009B2F8is DCFlushRange's flush-via-syscall. The ~10k missing syscalls (and the missing FP-unavailable traps) were per-frame game work that never ran. Symptom, not cause. 0x800A1990-- the "block/yield" in the retrace wait loop -- disassembles to OSYieldThread:OSDisableInterrupts; SelectThread(1); OSRestoreInterrupts. Colosseum's retrace wait is a yield-spin, not a queue sleep. Dispatch sampling shows it iterating at ~190k/s under LLVM. The other tested titles block on wait queues -- this is why only Colosseum failed.- Under the LLVM backend the entire critical section runs inside native calls within one dispatch. The only burst-ending points are the budget guards, which sit at call sites -- in OS code, almost entirely inside the interrupt-disabled window. Burst boundaries phase-lock to EE=0.
- External interrupts are only deliverable when the host regains control with EE=1: the VI interrupt sat pending-but-blocked at 30% of checks, only 1.2% of raises were ever delivered, the post-retrace callback never ran, and every frame escaped via the game's 3-second timeout. 1 frame / ~4 s = 0.2487 fps exactly. This also explains the bimodal 0.41/0.47/0.62 runs (occasional lucky boundary in the EE=1 zone) and the THP cold-boot hang (same wait pattern).
mtmsrwas lowered as a plain state write; Dolphin's own JITs end the block atmtmsrand check for pending interrupts. Two-part fix:- DolRecomp LLVM: exit to dispatcher when mtmsr enables MSR[EE] #23 -- EE-enabling
mtmsrtakes a side exit to the dispatcher - RecompCore Route fallback blocks through per-edge trampolines so the phi matches the CFG #18 -- the burst loop delivers a pending external interrupt when a boundary shows EE=1 (plus the lockstep FILTER/NODEDUP options used in this hunt)
Results
before after Colosseum LLVM 0.2487 fps 27.4-28.8 fps Colosseum C (control) 29.98 29.99 MKDD LLVM (control) 48.65 51.0 Colosseum LLVM cold boot hung at THP videos ~55 fps, boots Interrupt counters: raised 484 -> 31,777, delivered 462 -> 29,856, blocked 30% -> 0.5% -- matching the C profile. Renders the Phenac City scene correctly.
Remaining loose end from this investigation: forcing
STATICRECOMP_FALLBACK_RANGESacross OSDisableInterrupts/OSRestoreInterrupts still hangs outright -- mixed native/interpreter execution across MSR manipulation deserves its own issue.- Only 3
Uncapped runs (
EmulationSpeed = 0): AOT is the fastest arm, and a second scene-triggered stall surfacedSame pinned scene, 45 s uncapped windows, order-balanced, two runs per arm.
arm run 1 fps (speed) run 2 fps (speed) end frame C 42.08 (1.41x) 42.17 (1.40x) 95058 / 95008 LLVM AOT + fix 46.95 (1.56x) 43.39 (1.47x) 95219 / 95224 LLVM new + SIM + fix 34.65 (4.37x) 36.10 (4.32x) 93395 / 94212 LLVM new + fix 28.51 (0.36x) 27.60 (2.69x) 93763 / 94212 LLVM old + fix 22.41 (0.74x) 25.80 (0.85x) 94316 / 94308 Clean readings: C 42.1 fps / 1.41x, stable to 0.2% across runs. AOT 43-47 fps / ~1.5x -- the only arm faster than C on this title, up to 1.12x.
The other three rows are not valid speed measurements. Uncapped windows cover 2-4x more game time than the capped runs did, and they reach scene positions the capped windows never hit. At specific frame counts (93395, 93763, 94212 -- game-state positions, reproducible across arms and builds) the new and state-in-memory arms hard-stall rendering: no frames present (screenshot requests return nothing), while the guest CPU spins at up to 4.4x realtime -- which is why SIM "reads" speed=4.3 with fewer fps than C. C, AOT and old pass the same positions cleanly.
Two corrections this forces:
- The earlier claim that the two parked codegen optimizations "wedge Colosseum deterministically" attributed the stall to those changes. The stall frame numbers recur on the unmodified PR variant -- they are scene positions, and the faster variants simply reached them within the measurement window. Those optimizations may be fine; they need re-testing once the stall is fixed.
- The mtmsr delivery fix cures the steady-state retrace wait (the 0.25 fps symptom), but a second, scene-triggered render stall remains on the new-base LLVM arms -- same family as the original THP-video hang, and now the top open Colosseum issue. An edge-gated delivery variant was tried and does not prevent it.
Capped behavior is unaffected: the shipped variant still runs the scene at 27.9 fps with screenshots after these experiments.
Scene clarification: no C regression, and AOT confirmed fastest on the historical benchmark scene
A "C used to hit 60 on Colosseum" concern turned out to be scene identity. All the Aug 20-21 Colosseum campaigns in
games/pkmn-colo-recomp/build/benchmarks/usedpart1-in-town.sav(uncapped C median 53-55 fps, speed 0.89-0.92); the exact-59.94 readings were the capped cold-boot intro, which is 60 Hz content. The matrix earlier in this thread usedpart1-phenac-city.sav, a heavier cutscene (C: 42). Same backend, different scenes --willie-battle.savreads 35 for scale.Re-ran today's builds uncapped on the historical scene:
arm today, part1-in-town.savAug 21 campaign C 53.5 fps (speed 0.898) 53-55 (0.89-0.92) -- unchanged LLVM AOT + fix 56.5 (0.944) 0.85 (pre-fix, broken) LLVM new + SIM + fix 45.9 0.85 LLVM new + fix 37.6 0.85 So on both scenes measured, AOT is the fastest arm (1.06x C in town, up to 1.12x on the cutscene). And the scene-stall pathology reproduces here too: the SIM and plain-new runs both ended frozen at frame 176641 -- a second reproducible stall position for the new-base non-AOT arms, strengthening that as the top open Colosseum issue.
Stall investigation: runtime half fixed and pushed; the new-emitter half is characterized but not yet root-caused
Fixed and pushed (RecompCore PR #18, new commit): external-interrupt delivery is now restricted to post-
mtmsrboundaries, matching the block-ending JITs. Delivering at every EE=1 boundary could preempt the AX audio callback mid-work at scene transitions and consume any backend probabilistically -- the C arm crawled in one of two uncapped runs before, and passes 3/3 at 53.6-56.8 fps after. AOT stays clean, and the capped Colosseum fix is unaffected (25-29 fps).Still open: the deterministic scene-transition stall on the new-base LLVM arms (plain and
--state-in-memory; the SIM arm is the upstream base + state-in-memory, so this is one bug, not two). AOT, old and C all pass. State of knowledge:- Stall positions are game-state-deterministic (frame 176641 from
part1-in-town.sav, 94212/93763/93395 on the Phenac scene). - The stalled state is persistent: a savestate taken in it stalls even the C backend on load (
states/pre-stall.sav), so a guest-visible structure goes wrong on the approach and nothing recovers. - In the stalled state interrupts go quiet, not stormy: <4k raises/20 s vs 20k healthy (SI/DSP/VI/PE all ~stopped), zero deliveries, GPU never finishes a frame, while the guest spins through the interrupt-callback walker, the AX voice list and the yield loop at full speed.
- Ruled out by direct experiment: the inline paired-single FMA path (disabled -- still stalls), all vectorized paired-single arithmetic (disabled -- still stalls), the interrupt-callback list itself (dumped in the stalled state: 4 nodes, well-formed, identical to healthy), delivery eagerness (scoped -- C/AOT cured, new/SIM not), edge-gated delivery (no effect).
- Lockstep float divergences at the sound-engine entries (
0x8016F474cluster) persist even with all vector paired paths disabled, so their origin is elsewhere (psq quantized load/store or scalar float lowering are the remaining candidates); they may still be symptoms rather than cause.
Repro assets:
states/pre-stall.sav(stalls any backend instantly),part1-in-town.sav(new/SIM stall ~15 frames after load, capped or uncapped). Next probe queued: run the C arm frompart1-in-town.savsaving a state every second, find the last state from which the new arm still passes, then RAM-diff consecutive states across the new arm's execution to catch the first corrupted structure and its writer.- Stall positions are game-state-deterministic (frame 176641 from
RAM-level investigation: the stall is a phantom thread, and the goal metric is met by the AOT arm
LLVM is now faster than C on Colosseum wherever it runs correctly: the AOT arm measures 56.5 vs C's 53.5 uncapped on
part1-in-town.savand 43-47 vs 42 on the Phenac scene, clean in every run. The remaining defect is specific to the new-base emitter (plain and state-in-memory), and this session pinned down what actually breaks.What the corruption is
Thread-state dumps in the stalled state:
healthy stalled OSCurrentThread(0x800000E4)803FBCB0(main, state RUNNING)8048E010-- an address inside the main thread's stackreal threads 803FBCB0/803FB998RUNNING / READY both READY, never scheduled again A change tracker on
0x800000E4caught the write:803FBCB0 -> 8048E010, dispatch site = 0x800A1930-- the scheduler's own thread-switch store (stw r30, OSCurrentThread). The scheduler selected the phantom from a run queue, meaning something upstream enqueued a stack address as a thread -- the signature ofOSWakeupThreadrunning against a wait object that lived on a stack frame which had already been unwound.The phantom's garbage context then executes: depending on module timing the visible result is the hard wedge (frame 176641), or a 0.35-0.5 fps crawl where interrupts sit pending and unmasked (
pi_cause=0x548underpi_mask=0xFFC, EE set) while the guest runs garbage state -- the same crawl signature as the original retrace-timeout bug, which is why the presentations looked related.Ruled out this session, each by direct experiment
- OS interrupt masks (
0x800000C4/C8) and PI mask: byte-identical healthy vs stalled. - Run-queue heads and both threads' link fields: polled every dispatch from 2.5M on -- never hold a stack-range value at a dispatch boundary; the enqueue-and-consume happens inside a single native burst.
- Timebase staleness (volatile, uncache
mftbreads): no effect. - Inline paired-single FMA, and all vectorized paired-single arithmetic: disabled wholesale, still stalls.
- Delivery policy (eager / edge-gated / post-mtmsr-scoped): changes which failure mode appears, never prevents the corruption.
Tooling landed along the way
- ModernGekko PR LLVM-backend module stalls in the audio interrupt path; C backend unaffected #30:
DOLRECOMP_LLVM_WRITE_JOURNALwas missing from the module reuse identity, so toggling it silently reused a non-journaled module. - The module RAM write journal works for watchpoints, but the journaled module's timing shifts the race into the crawl mode instead of the phantom, so the write of the phantom pointer has not been caught red-handed yet.
The remaining hunt, precisely scoped
Find the guest code that, under new-base codegen only, wakes a stale on-stack wait object at the scene transition. Best next steps: (1) a burst-boundary ring of the last N dispatch sites, dumped when the tracker fires, to identify the burst that performed the enqueue; (2) diff the new emitter against the AOT emitter for the specific functions in that burst. Repro assets:
states/pre-stall.sav(poisoned state, stalls every backend),part1-in-town.sav(+~15 frames to failure), tracker/probes in this thread's history.- OS interrupt masks (
The historical 60 fps work, and what the ring probe changed about the stall
The dispatch-ring result: the "phantom" was never a phantom
The ring capture at the corruption moment, walked backward, names the whole flow:
- An init sequence (
801BDCC4bl-table calling800B85B4repeatedly) -- scene loading. - FP-unavailable exceptions (vector 0x800) around the music/event engine (
801A6F78,801B4B40). - Twelve dispatches inside
8000550C-- which disassembles to memcpy's inner loop, called
from0x80006410copying a 712-byte (OSContext-sized) block to0x803A1700. 0x80006450+then does:OSCreateThread(&thread_on_stack /* r1+120 */, ..., prio 1),
OSResumeThread,OSJoinThread-- and installs a custom handler for exception Emitter fix series + Dolphin-differential test suite from a downstream runtime project #8
(FP-unavailable) at0x800064A0, which is justOSLoadContext(saved)-- a crash-recovery
longjmp.0x800064C4is the game's crash-report printer.
So
OSCurrentThread = 0x8048E010is the game's own scene-loader worker thread, whose OSThread
legitimately lives on the caller's stack -- not corruption. The stall is Colosseum's crash
guard catching a fault in that worker under new-base codegen: the ring's final entry shows the
switch to the worker ending inside the crash-report path. C, AOT and old-LLVM run the worker
without faulting. The open question is now precisely "what does the worker execute differently
under the new emitter that trips the crash guard" -- with the crash window fully mapped
(0x80006378-0x80006520 plus the worker entry).The 60 fps / 6-scenes work, located
Three layers, in different states:
- The matching decompilation (
C:\Users\douglaswhittingham\pkmn-colosseum): 88.05% fuzzy,
6,786/8,603 functions matched, with hundreds of per-function promotion worktrees
(phase3/4/5-close-* covering HSD graphics, MusyX audio, SDK, fight, field...). This is the
scene-by-scene campaign the "targeted 60 fps on 6 scenes" memory refers to. - The profiling distillate:
games/pkmn-colo-recomp/patches/GC6E01-native-substitutions.csv
-- the seven hottest guest functions from scene traces: PSMTXConcat (6.1M calls in one traced
scene), PSMTXCopy/Identity/MultVec, HSD_MtxScaledAdd, memset, and memcpy (3.1M calls, most
through the L1 locked cache). Verified today: no consumer of this CSV exists anywhere in the
current toolchain -- nodolrecomp_native_*symbol appears in any built module. It is a
worklist, not a shipped feature. (The original C:...\pkmn-colo-recomp project dir was emptied
on Aug 20 during the move to G:.) - The widescreen code patches (
GC6E01-native-widescreen.csv): these ARE live -- verified
baked intopatched-main.dolinside every module we have benchmarked (word at 0x80005300 is
the patched value).--no-modsnever excluded them.
Why the substitution list matters now
It attacks exactly the costs the backend benchmarks measure and cannot remove:
- The C arm's 7.7M hook-fallback instructions per run are dominated by the locked-cache memcpy
path the CSV names. - PSMTX*/HSD are paired-single-heavy guest code -- the biggest per-scene CPU item.
- Uncapped Colosseum today: C 0.90x realtime, AOT 0.94x. Native substitution of these seven
functions stacks with any backend and is the credible route past 1.0x -- and the decomp already
holds matched C sources for most of them, so the substitutes largely exist. - Bonus: substituting the FP-heavy PSMTX*/HSD functions removes guest FP from exactly the code the
scene-loader worker runs, which may sidestep the new-emitter crash window entirely.
Attempting to run with the runner's mod layer enabled (dropping
--no-mods) crawls both backends
-- the runtime mod path is not how these were meant to be applied; the substitution mechanism
needs to be built (link native implementations over the seven guest addresses at module build).- An init sequence (
Correction: Colosseum freezes on every LLVM arm; Skyward Sword now works
Correction to the earlier "Colosseum fixed at 28 fps" claim
The mtmsr/EE-delivery fix did cure the retrace-interrupt starvation (0.25 fps -> the
game renders and cold-boots past the THP intros). But from a savestate, every LLVM
arm still freezes after a few seconds, and the fps averages reported earlier were
measuring "how long until it froze", not sustained throughput.Evidence: uncapped 25 s runs,
part1-phenac-nw.sav, repeated back to back.arm run 1 end frame run 2 end frame speed C backend 278262 278285 1.22 LLVM, upstream main 277678 277678 0.06 LLVM, pre-upstream fix build 277678 277678 0.06 C reaches a different frame each run because it keeps running. Every LLVM arm halts
at a byte-identical frame with the guest CPU at 6% of realtime. Reproduced on three
states so far:part1-phenac-nw(277678),willie-battle(67679),part1-in-town
(176641). Not upstream-specific -- the pre-upstream build freezes at the same frame --
and not caused by native substitutions, which were not present in the control arms.This is the same phantom-thread / crash-guard defect tracked above, but its scope is
wider than previously recorded: it is not limited to scene transitions, it happens
from ordinary gameplay states. Colosseum LLVM performance numbers should not be
quoted until this is fixed, because none of them represent sustained execution.Cold boot is the exception: a 150 s cold-boot run rendered throughout at ~55 fps, so
the trigger is tied to savestate-restored state rather than to any scene reached
by playing forward.Skyward Sword (SOUE01, Wii/Broadway) now runs on LLVM
Previously recorded as producing no usable LLVM number at all. On current upstream
main it renders correctly (verified by screenshot: full HUD, hearts, control
prompts, textures) and runs freely -- frame counts advance differently every run,
no freeze.arm run 1 run 2 speed LLVM 33.82 fps 35.08 fps 1.18 C 43.87 fps 43.07 fps 1.50 LLVM is at 0.79x of C here, so there is headroom to chase, but the title is no
longer broken on the LLVM backend.Native substitutions: implemented, not yet measurable
Seven profile-identified hot guest functions (PSMTXConcat/Copy/Identity/MultVec,
HSD_MtxScaledAdd, memcpy, memset) can now be replaced with native implementations
viaDOLRECOMP_SUBSTITUTIONS=addr=symbol,.... The mechanism works end to end and is
semantically correct -- the substituted Colosseum module renders the reference scene
pixel-identically to the unsubstituted one. Its performance effect cannot be measured
until the freeze above is resolved.Two bugs found while building it, both worth noting for anyone extending the IR:
src/ir/dolir_builder.chad no<stdlib.h>, so a newly addedgetenvwas
implicitly declared as returningint; the truncated pointer segfaulted the
recompiler. MSVC reports this only as warning C4013.control_flow.cppdereferencesvalues_[term.condition]unconditionally for
DOLIR_TERM_INDIRECT, so an unconditional indirect edge still needs a
materialized constant-true condition.DOLIR_NO_VALUEcrashes the emitter.
Decrementer starvation (real, unresolved) + Apple Silicon results
The decrementer is never delivered from native code
Colosseum's frozen guest spins at
800feb9c+8009df80-- the same pair that dominated
the earlier lockstep run (728k divergences atentry=0x800FEB9C end=0x8009DF80, native
MSR0x9932vs interpreter0x1932). Disassembling0x800FEA74-0x800FEB9Cshows an
infinite OS dispatcher loop: walk a callback list, call each ready callback, then
under OSDisableInterrupts merge a pending list into it by priority, OSRestoreInterrupts,
and branch unconditionally back to the top. It never returns; it only stops hogging the
CPU when an interrupt preempts it.The runtime's asynchronous exception set is
SYNC_EXCEPTION_MASK = ~(EXCEPTION_EXTERNAL_INT | EXCEPTION_DECREMENTER | EXCEPTION_PERFORMANCE_MONITOR)but the delivery gate added for the retrace fix tests
EXCEPTION_EXTERNAL_INTonly. So
the decrementer -- which drives GameCube OS alarms and thread timeslicing -- is never
delivered while the guest runs native code. Titles driven by VI retrace (Mario Kart,
Luigi's Mansion, Skyward Sword) are unaffected; Colosseum's dispatcher waits on
alarm-driven preemption that never arrives. That is a concrete answer to "why only
Colosseum".Widening the gate to
(ppc.Exceptions & ~SYNC_EXCEPTION_MASK)measurably changes the
failure: the previously deterministic halt (three states, always the identical end frame)
became intermittent, and one run reached speed 1.05 -- full realtime with a fresh end
frame.variant end frames across runs verdict external-int only (shipped) 277678, 277678, 277678 always frozen + decrementer, post-mtmsr gate 277865, 277454, 277678 intermittent + decrementer, any EE=1 boundary 277505 (speed 1.05), 277678, 277678 intermittent Not a fix on its own, and currently reverted in the tree, but it is a genuine gap and the
most specific lead so far.Correction: the revert was based on a contaminated signal
The widened gate was reverted because Skyward Sword appeared to drop to 0 fps. That
reading was wrong. A background agent was regenerating SOUE01 modules for a separate arm
matrix at the same time --active-module.txthad been repointed from8a49f3e1...to
759fcc8d...underneath the benchmark, and a second emulator was competing for the GPU.
Skyward Sword was fine throughout. The decrementer change must be re-tested on a quiet
machine before any conclusion about it is drawn.Method rule this produces: never benchmark a title while another agent or session is
generating or running that same title, and check the module hash inactive-module.txt
before and after a benchmark series.Regression check (concurrent load present, numbers not quotable)
Both titles run and advance frames -- no freeze, no regression to the broken state:
title LLVM C Mario Kart DD 41.3 / 44.3 fps 52.1 / 19.9 fps Luigi's Mansion 33.4 / 37.3 fps 52.8 / 48.6 fps The C 19.9 fps outlier shows the load contamination; these are health checks, not scores.
Apple Silicon, upstream main (M5, 10 cores)
Luigi's Mansion
foyer.sav, Null graphics, headless, uncapped, interleaved blocks with a
discarded warm run per arm.- LLVM median 81.18 fps (79.95, 81.15, 81.21, 81.45)
- C median 79.10 fps (74.48, 77.68, 79.10, 79.27, 79.56)
- Ratio 1.026x (1.016x restricting to the quietest campaign)
Versus the previous Apple Silicon campaign's baseline of ~71.5 fps on this state, that is
about +13.5%, and it replaces the earlier 0.93x regression with parity-to-slightly-ahead.
No freeze: end frames varied every run (6678...7018), ~2400 frames per 30 s window.Caveats: the C control is a pre-existing module from an older dolrecomp revision, not a
same-commit rebuild, because the C backend at 1bec355 emits calls to
ppc_fp_available_inlinewhich does not exist in that checkout's GXRuntime. The Mac is a
shared, noisy host -- one campaign was discarded entirely at load average 15.5 -- so
anything under ~5% there needs many more pairs.macOS build blocker found
The LLVM backend does not build a module on macOS at 1bec355. Generation dies at
chunk 57/4164:LLVM ERROR: Global variable 'fix_pair_nan_f64' has an invalid section specifier '.text.unlikely.fix_pair_nan_f64': mach-o section specifier requires a segment and section separated by a comma.markColdFixup()insrc/backend/llvm/fp_fixups.cpp:21sets an ELF-style section name
unconditionally. Skipping the cold-section hint on Mach-O targets lets all 4164 chunks
generate. Worth fixing upstream.Colosseum: the freeze is a guest crash dump, not a scheduling bug
What the guest is actually doing
Instrumenting the stuck state and dumping the buffer the guest is writing settles it.
The LLVM arm parks in the C runtime's line-buffered stdio write loop
(0x800C76F0-0x800C77E0, which callsmemchr(buf, '\n', n)at0x800C811C), and the
bytes it is emitting are:r26=80313900 r29=0000001A r30=80310A14 r31=00000001 msr=00001030 text=Non-recoverable Exception %d....Unhandled Exception %d... DSISR = 0x%08x DAR = 0x%08x.....TB = 0x%016llx... Instruction at 0x%x (read from SRR0) attempted to acce[ss]...r29 = 0x1A= 26 bytes, exactly the length of"Non-recoverable Exception ". That is the
GameCube SDK's unhandled-exception dump, andmsr = 0x1030(EE clear, FP clear, ME/IR/DR
set) is exception context.So the LLVM arm takes a guest exception the OS treats as unhandled, and then hangs
inside its own crash dump, whose output never drains, looping forever with interrupts
disabled. Everything previously described as "the freeze" is that hang. The apparent
scheduling symptoms were downstream of it.What this retires
- The decrementer theory is dead. Counters over a stuck run:
dec_seen=30,
dec_del=22againstext_seen=393828. The decrementer is essentially never pending, so
it cannot be what starves the guest. The static check that motivated it was also wrong:
Colosseum, Mario Kart and Luigi's Mansion all have identical decrementer usage
(2mtspr DEC, 1mfspr DEC-- the same SDK leaves), so "only Colosseum uses alarms" was
never true. Branchfix/async-exception-deliveryholds the widening; it is a dead end
for this bug and should not be merged on these grounds. - The phantom-thread reading is superseded.
OSCurrentThreadpointing into a stack was
the game's own scene-loader worker; the crash guard around it is the SDK error path, not
corruption.
The live signal
Blocked-interrupt PC sampling, same stuck run, is unambiguous:
arm ext_seen ext_del blocked (EE=0) top EE=0 site LLVM 393,828 8,933 384,903 (98%) 800c77e0-- 1476 of 1503 samplesC 79,255 30,353 48,970 (62%) 80164f1c-- 88 samples;800c77e0absentC never blocks at
800c77e0at all. The site is the crash-dump loop, so this is simply
another view of the same hang.What is still unknown
Which exception, and at what SRR0. Three probes failed to catch it and the reason is
instructive: the exception vectors live in low RAM, copied there at boot, so they are not
part of the recompiled DOL, are not natively dispatchable, and never appear either as a
native dispatch target or asppc.pcat the sampling points tried. A probe that catches
this has to hook the fallback/interpreter path, or the runtime'sppc_take_exception
inside the module (module-side, so it needs a regenerated module).That is the next step, and it is now a narrow one: name the exception and its faulting
instruction, and the defect in the generated code follows directly.Method notes worth keeping
- The exception vectors are not in the DOL. Any probe that samples only native dispatch
will miss all guest exception entry. frame_countequality across runs is good evidence of a hang but not proof: on Skyward
Sword two healthy runs matched by coincidence. Confirm withpresent_countandstate=.- A module's directory hash is derived from
dolrecomp_binary_sha256, not from emitter
options, so "the hash changed" does not prove a codegen flag took effect. Diff an actual
emitted chunk instead.
- The decrementer theory is dead. Counters over a stuck run:
Module-side exception probe: every exception taken is benign
Instrumenting
ppc_take_exceptioninside the module (the only place that sees guest
exception entry, since the vectors are not natively dispatchable) and regenerating the
Colosseum LLVM module gives the first 10 exceptions of a stuck run:vector=0xC00 exception=0x08 srr0=0x80098038 (system call -- sc at 0x80098034) vector=0x800 exception=0x20 srr0=0x80165330 (FP unavailable, lazy FPU) vector=0x800 exception=0x20 srr0=0x8016CBB0 (FP unavailable) vector=0xC00 exception=0x08 srr0=0x8009B2FC (system call -- DCFlushRange) ... same three kinds repeatingNo DSI (0x300), no ISI (0x400), no alignment (0x600), no program (0x700). Only system
calls and lazy-FPU traps, both entirely normal. So within the observation window the guest
takes no faulting exception at all, yet it is simultaneously emitting
"Non-recoverable Exception "(26 bytes, matchingr29=0x1Aexactly).Two possibilities remain, and they are cheap to separate:
- The crash-dump write begins before the first dispatch of the window -- i.e. immediately
on savestate restore -- so the faulting exception is never observed. Against this: the C
backend restores the same state and runs normally, so the state itself is healthy. - Both arms enter this print path, but the C arm's write drains and the LLVM arm's does
not. The loop retries viabl 0x800C7454at0x800C77B0and only exits when that
returns non-zero, so a flush that never reports progress would hang exactly here while
the same code on C completes. This would make the hang a stdio/flush-path defect rather
than a crash at all, and would explain the absence of any faulting exception.
Distinguishing them needs the buffer probe run against the C module as well: if C also
writes"Non-recoverable Exception "and proceeds, possibility 2 is confirmed and the
investigation moves to0x800C7454and what it depends on.- The crash-dump write begins before the first dispatch of the window -- i.e. immediately
Colosseum: the divergence is pinned to one point
C never enters the crash path at all
A host-side probe counting entries into the stdio write loop (
0x800C76F0-0x800C77E4),
run on both arms frompart1-phenac-nw.sav:arm write-loop entries text being written LLVM many; parks there "Non-recoverable Exception "(26 bytes,r29=0x1A)C zero -- So the earlier "maybe both arms print it and only C drains" theory is dead. Only the LLVM
arm reaches the unhandled-exception path.The path into it
Dispatch ring captured at the first write-loop entry, oldest first:
80164F1C 800C4824 801653AC 80164478 8015D714 80162254 8009E758 800AE9FC 80163F98 800AEBD8 8009E724 8009DF5C 801643C8 8009A190 8009A1E0 80006378 80005530 8000642C 8009C6F0 8009A190 8009A1E0 800C8864MusyX audio code (
0x8015xxxx-0x8016xxxx) -> context save (8009A190/8009A1E0) ->
the game's error handler (80006378->80005530memcpy ->8000642C) -> crash printer.The exceptions are legitimate
Module-side probe on
ppc_take_exception(the only place that sees guest exception entry,
since the vectors live in low RAM outside the recompiled DOL). Both FP-unavailable SRR0
values point at real FPU instructions, so the traps and their reported addresses are
correct, not corrupted:0x80165330 lfs r0,-25008(r2) 0x8016CBB0 fsubs f2,f23,f26The divergence point
First twelve exceptions, same savestate, same runtime binary:
# C arm LLVM arm 1 0xC00 srr0=80098038 0xC00 srr0=80098038 2 0x800 srr0=80165330 0x800 srr0=80165330 3 0x800 srr0=800A3250 0x800 srr0=8016CBB0 4 0xC00 srr0=80098038 0xC00 srr0=80098038 5 0xC00 srr0=80098038 0xC00 srr0=80098038 6 0xC00 srr0=8009B2FC 0xC00 srr0=8009B2FC 7-12 identical identical The two arms are in lockstep through exception #2 and diverge at #3: C's next FP trap is in
the SDK matrix library (0x800A3250, inside thePSMTX*range), the LLVM arm's is in MusyX
(0x8016CBB0). Everything before that point matches instruction-for-instruction as far as
this probe can see.That is a narrow, well-defined window: between the FP-unavailable trap at
0x80165330
and the following trap. Whatever the LLVM backend does differently happens inside it. The
next step is lockstep restricted to that window --STATICRECOMP_LOCKSTEP_STARTset just
before exception #2 with_NODEDUP=1-- rather than the whole-run lockstep that previously
drowned in step-cap artifacts.Housekeeping
GXRuntime/src/core/cpu_exception.ccarries a temporary probe; both GC6E01 modules were
regenerated with it and print to stderr on every guest exception. Revert the file and
regenerate both modules before quoting any benchmark from them.Windowed lockstep, and a direct test of the block it names
Divergence pinned to a 244-dispatch window (LLVM exc#2 at dispatch 18992, exc#3 at 19236;
the C arm takes 2074 dispatches over the same stretch). Lockstep aimed at it
(_START=18900 _LIMIT=19400 _NODEDUP=1 _STEPCAP=200000) gave 258 divergences dominated by
one repeating pair:entry=0x8016CB7C end=0x8016C0B0: r5:N=0xcc010000,I=0x4 r6:N=0x0,I=0x803fc860 r7:N=0x0,I=0xff0000 r8:N=0x1,I=0x9000580 mmio#:N=27,I=18 writes at 0xCC0080000x8016CB7Cis write-gather-pipe code, and0x8016CBB0-- where LLVM's exception #3 fires
-- is inside it:8016CBB0 fsubs f2,f23,f26 8016CBB4 lis r3,0xcc01 8016CBC0 stfs f2,-32768(r3) ; 0xCC008000, the GX write-gather pipeThe MMIO count mismatch is probably not real:
emitGuestStoreissues exactly one write per
path (MEM1 / MEM2 / external) before joining, and the lockstep shadow interpreter is known
to re-run blocks and double-issue side effects.Direct test, independent of lockstep -- forcing that range to the interpreter:
configuration speed end frame LLVM, unmodified 0.06 277678 every run LLVM, 0x8016c000-0x8016d000 interpreted 0.84 - 0.88 277433 every run CPU throughput goes from 6% of realtime to ~87%, so the block is implicated, but frames
still do not advance -- the failure moves rather than clearing. Caveat: forcing ranges to
the interpreter has hung this title on its own before, so mixed-mode is not a clean
instrument.Resolved: Colosseum runs at full speed on the LLVM backend. The failure was a modified DOL.
The extracted game this investigation was run against,
games/pkmn-colo-recomp/extracted/Pokemon-Colosseum-USA, has a widescreen mod baked intomain.dol:address clean what we were benchmarking 0x800A3930FFA01090(fmr f29,f2)4BF619D0(b 0x80005300)0x8000530000000000C3A2B084(lfs f29,-20348(r2))0x8000530400000000EFBD00B2(fp single op)0x80005308000000004809E62C(b 0x800A3934)A 16:9 hook that diverts the projection-matrix path into hand-injected FP code placed in former zero padding. It renders visibly wrong -- the boot logo is squashed and menus collapse into a vertical strip -- and it sits in exactly the FP/matrix region where this issue kept finding "divergences". Two other extractions in the same tree (
Colo-Fresh,Colo-PiDOL) are clean.Measured against a clean DOL
Colo-Fresh, with fresh savestates captured on that DOL (the previous states were captured on the modded binary, so their RAM carried the patched code and were never valid against a clean module). Interleaved, warm, uncapped:savestate plain LLVM LLVM + state-in-memory C backend title 49.26 61.22 86.45 early 46.93 61.01 92.16 mid 49.76 64.48 86.64 state-in-memory is worth 1.24x - 1.30x on this title, and its uncapped
speedis above realtime in every state (1.006 / 1.024 / 1.078). Capped, it holds the frame rate exactly:savestate fps vps speed title 60.07 60.05 1.001 early 60.12 60.11 1.002 mid 59.94 59.94 1.000 Rendering verified by screenshot: correct aspect, correct text, no distortion. No freeze in any state, on any arm.
C is still ahead uncapped (LLVM/C ~0.66 - 0.74 with state-in-memory), so there is optimization headroom, but it no longer affects playability -- the title is 60 Hz content and the LLVM backend sustains it.
What this retracts
Everything in this issue measured against
Pokemon-Colosseum-USAis unreliable. Specifically retracted: the deterministic freeze and its "identical end frame" evidence; the0.25 fpsretrace-timeout analysis; the interrupt starvation counters (blocked_eeratios, raised/delivered); the__OSUnhandledExceptioncrash-dump finding; the phantom-thread reading; the decrementer theory; and the0x8016CB7Cgather-pipe divergence. Some of those may describe real effects of the injected code, but none of them describe the shipping game.ExpansionPak/RecompCore#18 has been withdrawn for the same reason -- a clean DOL runs at full speed with no delivery fix present at all.
What remains worth doing
- Low priority, real: on the modded DOL the C backend copes with the injected code at
0x80005300and the LLVM backend does not. That is a genuine backend difference in handling code placed in former padding, and deserves its own issue if anyone wants the widescreen mod to work. - Optimization: closing the remaining LLVM/C gap on this title.
Method changes adopted
- Benchmark only from cleanly extracted, unmodified DOLs; verify
main.dolagainst the untouched extraction first. - Savestates are only valid against a module built from the same DOL.
moderngekko-portappliesData/Sys/GameSettings/<id>.ini[OnFrame_Enabled]patches with no opt-out flag; for GC6E01 that is the memory-card patch. The widescreen hook came from the extraction itself, not from that mechanism.
- Low priority, real: on the modded DOL the C backend copes with the injected code at
sorry doug but i aint reading allat, good luck on your recomp lmao
You're good. This is just tracking. The recomp is getting closer.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDone
Correction (updated): I originally filed this as AArch64-specific. That was wrong. The failure reproduces on x86-64 with the same DOL and the same harness, so this is a Colosseum + LLVM-backend defect, not an ARM one. Details and measurements below.
Building Pokémon Colosseum (GC6E01) with the LLVM backend produces a module that does not get through the boot sequence. The C backend, same DOL, same runtime, same automation harness, boots normally.
Measurements
Cold boot, A-press driver, no savestate, identical harness on each platform:
The guest is not deadlocked — it executes (
speedis nonzero, the VI keeps ticking), it just never makes forward progress. Note the LLVM arm reports a higherspeedthan C on the Pi while producing no frames: it is spinning cheaply instead of doing work.Luigi's Mansion (GLME01) with the LLVM backend runs fine on the same Pi (769 frames, 6.11 fps), so the AArch64 LLVM backend is not broken in general.
Where it ends up
perfon the wedged aarch64 run puts ~58% of samples in three guest functions, which the DOL disassembly identifies as the thread scheduler:0x8009BBE0—OSSaveContext(saves GQR1–5 viamfspr 0x393..0x397, CR, LR, MSR, CTR, XER)0x8009BC50—OSLoadContext, which resumes a thread withrfi0x800A17E0— run-queue enqueue / rescheduleNone of these appear in the C backend's profile, which is dominated by real game work (
func_800FD5E0, paired-single ops, FP helpers).What is NOT the cause
I instrumented the shared runtime to rule these out rather than guess:
ppc_rfifires ~150M times in 70s (vs ~150k for C, ~1000x more), and at every sampled resumeCPUStateis correct:srr0=0x800A183C,r3=0x00000001.0x800A183Cis the instruction afterbl OSSaveContext, andr3=1is exactly the "resumed" return value the scheduler tests. The C backend resumes at the same address with the same value, so this churn is a normal idle loop — the LLVM arm is just doing ~1000x more of it.guest_800A183C_b23), reached fromnative_entry, and the block loads r3 fromCPUStateand branches onicmp eq i32 %state2.7, 0correctly.embedded-data; there are no unmodeled opcodes.Still open
I have not isolated the primitive that diverges. The dispatch-address stream diverges from the C backend within the first handful of dispatches at boot, and the LLVM arm reaches far less distinct guest code over a run (909 vs 4000+ distinct addresses in the same window).
Reproduction is easy on x86-64 now, so this no longer needs ARM hardware to chase. Happy to keep digging or hand over what I have.