Repository navigation
LLVM backend does not sustain execution on a DOL with code injected into former padding #27
Description
Activity
Still reproduces on a current tree, and the diagnostics point somewhere other than the injected code — which may change what this issue is about.
Setup: modded DOL (
Pokemon-Colosseum-USA), a savestate captured on that same DOL (part1-bar-cutscene.sav), runtime at currentmoderngekko-vendorplus the post-mtmsrexternal-interrupt delivery that landed since this was filed. Frames advanced over a fixed 20s window, input pinned, one emulator at a time.arm frames/20s fps native_exc hook_fb smc_mismatch c 605 29.97 17,737 7,519,763 0 llvm-new 16 0.25 343 37,481 0Three things that narrow it:
It is not SMC fallback.
smc_mismatch=0on both arms — the module was built from the modded DOL, so the chunk hashes match and nothing is being forced to the interpreter. The stall happens with the injected code fully native.It is not stuck in the injected region. Dispatch sampling shows nothing at or near
0x80005300. The hot sites are:800a128c samples=11719 800a183c samples=10586 8000550c samples=7 800d41e8 samples=4The guest reaches an OS wait loop at
0x800a128c/0x800a183cand never leaves it.The signature is interrupt starvation, not a codegen fault.
native_excis 343 against C's 17,737 on the same state. The guest is executing (174.8M native dispatches) but almost no exceptions are being delivered — it is spinning on an event that never arrives.For what it is worth,
0x800a128cis the same wait loop that showed up while diagnosing the interpreter-fallback hotspot inOSRestoreInterruptson the clean DOL. In that case the guest was likewise burning cycles in that loop whilenative_excsat at ~300 against a healthy ~12,000, and the cause was that the address the runtime returned to aftermtmsrwas not enterable, so the interrupt was never delivered there.That suggests a reframing worth testing: the patched branch at
0x800A3930may leave the guest re-entering at an address the LLVM backend did not make enterable, with the visible failure being the wait loop rather than anything wrong with the three injected instructions. If so this is closer to an entry-point/dispatch-boundary problem than a "code discovered in former padding" problem, and the C backend copes because it labels every address.I have not confirmed that mechanism here — the above is what the counters and dispatch profile show, not a root cause. Happy to dig further if useful; the repro is cheap now that the modded DOL and matching states are validated.
Ran a control that I think changes what this issue is about. The LLVM module I reproduced with stalls on the clean DOL too, so the injected code is not the variable.
The control
Same module vintage, same recompiler binary, only the DOL and state differ:
arm / DOL frames/20s native_exc top dispatch pc llvm CLEAN dol, clean-field2 7 93 800a128c C CLEAN dol, clean-field2 653 18,971 8009df3c llvm MODDED dol, part1-bar-cutscene 16 343 800a128c C MODDED dol, part1-bar-cutscene 684 18,412 800feb9cThe LLVM arm dies the same way on both DOLs — same wait loop at
0x800a128c, same interrupt starvation. The modded DOL is not what breaks it.The confound is module vintage. The modules in question were built 2026-08-21 with dolrecomp
6bdb3831ef75. A current unfixed LLVM build (3166222f85ec) runsclean-field2at 554 frames, so whatever broke that August build has since been addressed.Rebuilt from the modded DOL with a current recompiler
Same modded DOL, module regenerated with a current build (includes the post-
mtmsrregion-leader change from #28):state current LLVM Aug-21 LLVM C part1-bar-cutscene 208 f / 28.4 fps 15 f / 0.26 734 f part1-in-town 0 f / 34.7 fps 7 f / 0.27 823 f willie-battle 0 f / 19.0 fps 10 f / 0.40 363 fLarge improvement over the old build —
part1-bar-cutscenegoes 15 → 208 frames withnative_exc343 → 8,873 — but it is not fully fixed: still well short of C, and two states show a static guestframe_countalongside a healthy hostfps, which I cannot yet account for. I would not claim this issue is resolved.The trampoline, for reference
0x800A3930: b 0x80005300 (was: fmr f29,f2) 0x80005300: lfs f29, -20348(r2) 0x80005304: fmuls f29, f29, f0 0x80005308: b 0x800A3934 <- returns mid-functionThe return target
0x800A3934is an address nothing branched to in the original image. That is the same "must be enterable" shape as #28, so it may still matter for this title even though it is not what caused the stall I reproduced.Suggested reframing
The reproduction in the original report may have been measuring a general LLVM-backend defect of that period rather than anything mod-specific. Worth re-running against a current build before treating "code injected into former padding" as the cause — on current builds the modded DOL runs, just not yet at parity with C.
Correcting my previous comment. I said the modded DOL was not the variable, based on an old module stalling on both DOLs. That reasoning was right about that module and wrong about the issue. With a current build there is a real modded-DOL defect, and it is sharper than the original report describes.
What actually happens
Module rebuilt from the modded DOL with a current recompiler (includes the post-
mtmsrregion-leader change from #28). Run frompart1-in-town.sav, captured on that same DOL:state=running booted=1 fps=28.6795 speed=0.324565 frame_count=176641 <- after ~90 secondsIt boots, loads the state, and the guest executes — 174.8M native dispatches, non-zero
speed— butframe_countnever leaves the savestate's own value. Zero video frames completed, ever. C on the identical DOL and state was past 176953 on first sample and kept climbing.Dispatch sampling puts it in an OS wait loop:
800a128c samples=11719 800a183c samples=10586with
native_excat 343 against C's 17,737 on the same state. It is starved of the interrupt that would let it leave the loop, not stuck executing bad code.Reproducibility
part1-in-townandwillie-battlefail every time (3 separate sessions).part1-bar-cutsceneandpart1-phenac-citysometimes run. Controls interleaved immediately before and after each failure — clean-DOL LLVM, and C on the modded DOL — passed each time, so this is not drift.What it is not
- Not the injected code. Forcing
80005300-80005320and800A3900-800A3960to the interpreter viaSTATICRECOMP_FALLBACK_RANGESdoes not clear it. - Not SMC.
smc_mismatch=0; the module was built from the modded DOL so hashes match. - Not the backend generally. Mario Kart 1052 frames, Luigi's Mansion 419, clean-DOL Colosseum 639, all with the same LLVM lineage.
- Not I/O or the host. Module is page-cache warm (0s read), GPU healthy, 81 GB RAM free, CPU 12%.
The trampoline, for reference
0x800A3930: b 0x80005300 (was: fmr f29,f2) 0x80005300: lfs f29, -20348(r2) 0x80005304: fmuls f29, f29, f0 0x80005308: b 0x800A3934 <- returns mid-function0x800A3934is an address nothing branched to in the original image. Since forcing those ranges to the interpreter does not help, the trampoline may not be the trigger directly — but the presence of a branch target the original image never had is the kind of thing that changes region formation elsewhere, which would fit an entry-boundary explanation rather than a codegen one.Two measurement notes for anyone reproducing
booted=1can be written after a 90-120s poll gives up, so an apparent "fails to boot" here is often "slow to publish status" — checkstatus.txtdirectly rather than trusting a poll timeout. Andfpsinstatus.txtis a last-written value, not a heartbeat: a stalled run keeps reporting its finalfpsindefinitely. Use theframe_countdelta plus whether automation commands are still being consumed.- Not the injected code. Forcing
More elimination. Summary first: both LLVM configurations fail, only C works, and it is the DOL rather than the savestates.
state-in-memory fails too
I had only tested the default LLVM configuration.
--state-in-memorybehaves the same:arm frames/20s cmds consumed C modded in-town 265 yes llvm modded in-town 0 no HUNG llvm --state-in-memory 0 no HUNG llvm --state-in-memory willie 0 no HUNG llvm --state-in-memory bar-cut 0 no HUNG--state-in-memoryalso failspart1-bar-cutscene, which the default configuration sometimes survives. So this is not specific to one codegen mode.It is not the savestates
The
part1-*states were captured in August on an older build, so I captured a fresh one: modded DOL, C module, saved through the automation protocol minutes before the test.C fresh state 279 frames OK llvm fresh state 0 frames HUNG sim fresh state 0 frames HUNG sim old state 0 frames HUNGA brand new state on the modded DOL hangs exactly the same way. Stale savestates are excluded.
It is not the modified regions
Forcing the changed code to the interpreter via
STATICRECOMP_FALLBACK_RANGESdoes not help:80005300-80005320 HUNG (the three injected instructions) 80005300-80005500 HUNG (the whole chunk, as named by the SMC diagnostic) 80005000-80005600 HUNG (wider) 800A3800-800A3A00 HUNG (the patch-site chunk) 80005300-80005500,800A3800-800A3A00 HUNG (both)It is not chunk structure
Modded and clean builds emit an identical 4889 chunks. The injected code lands inside a chunk that was previously zero padding; it does not create a new code range.
The guest never reaches the patched code
Dispatch sampling on a hung run:
800a128c samples=11719 800a183c samples=10586 8000550c samples=7 800d41e8 samples=4Nothing in
800A39xxor800053xx. The guest parks in an OS wait loop withnative_excat 343 against C's 17,737 on the same state — starved of the interrupt that would let it proceed. It never executes the trampoline at all.What that leaves
Four changed words in the DOL break the LLVM backend somewhere the guest reaches before touching them, and the failure mode is interrupt starvation rather than bad code at the patch site. Forcing either modified region to the interpreter does not help, so the defect is not the codegen for those instructions.
Also worth noting the priority framing: for anyone using the widescreen mod, this is not a corner case — it is the difference between running the mod on the C backend and being unable to use the LLVM backend at all.
Reproduction notes
- Controls interleaved immediately before and after every failure (C on the modded DOL, LLVM on the clean DOL); both pass each time, so this is not host drift.
fpsinstatus.txtis a last-written value, not a heartbeat — a hung run reports its finalfpsindefinitely. Use theframe_countdelta plus whether.cmdfiles are still being consumed from the automation directory.booted=1can be published later than a short poll allows; an apparent boot failure is often a slow status write.save_stateneeds a path relative to the automation directory; an absolute path is rejected withpath=<file> is required.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsTodo
Splitting this out of #20, at low priority.
The Colosseum extraction we had been benchmarking carries a widescreen mod in
main.dol:0x800A3930hasfmr f29,f2replaced withb 0x80005300, and three instructions are injected at0x80005300-0x80005308, an address range that is zero padding in a clean DOL.The C backend copes with this. The LLVM backend does not.
On the modded DOL the LLVM arm does not sustain execution -- from a savestate it stops advancing frames at a deterministic point with the guest CPU at ~6% of realtime, while the C backend runs the same state normally. On a clean DOL neither backend has any problem, and the LLVM backend runs the title at a locked 60 fps.
This does not affect the shipping game, so it is not urgent. It matters if the widescreen mod is meant to be supported, and it may point at something generally useful: how the LLVM backend treats code discovered in regions that were zero-filled in the original image, and whether such code is reached only through a patched branch it may not have analysed as a call target.
Reproduction, if wanted:
games/pkmn-colo-recomp/extracted/Pokemon-Colosseum-USAgames/pkmn-colo-recomp/extracted/Colo-Fresh0x80005300,0x80005304,0x80005308,0x800A3930Note that savestates are only valid against a module built from the same DOL, so a modded-DOL repro needs states captured on the modded DOL.