Skip to content

LLVM backend does not sustain execution on a DOL with code injected into former padding #27

Description

@dougchansan

Splitting this out of #20, at low priority.

The Colosseum extraction we had been benchmarking carries a widescreen mod in main.dol: 0x800A3930 has fmr f29,f2 replaced with b 0x80005300, and three instructions are injected at 0x80005300-0x80005308, an address range that is zero padding in a clean DOL.

The C backend copes with this. The LLVM backend does not.

On the modded DOL the LLVM arm does not sustain execution -- from a savestate it stops advancing frames at a deterministic point with the guest CPU at ~6% of realtime, while the C backend runs the same state normally. On a clean DOL neither backend has any problem, and the LLVM backend runs the title at a locked 60 fps.

This does not affect the shipping game, so it is not urgent. It matters if the widescreen mod is meant to be supported, and it may point at something generally useful: how the LLVM backend treats code discovered in regions that were zero-filled in the original image, and whether such code is reached only through a patched branch it may not have analysed as a call target.

Reproduction, if wanted:

  • modded: games/pkmn-colo-recomp/extracted/Pokemon-Colosseum-USA
  • clean: games/pkmn-colo-recomp/extracted/Colo-Fresh
  • the two differ at 0x80005300, 0x80005304, 0x80005308, 0x800A3930

Note that savestates are only valid against a module built from the same DOL, so a modded-DOL repro needs states captured on the modded DOL.

Activity

  1. dougchansan commented on Sep 1, 2026

    @dougchansan
    ContributorAuthor

    Still reproduces on a current tree, and the diagnostics point somewhere other than the injected code — which may change what this issue is about.

    Setup: modded DOL (Pokemon-Colosseum-USA), a savestate captured on that same DOL (part1-bar-cutscene.sav), runtime at current moderngekko-vendor plus the post-mtmsr external-interrupt delivery that landed since this was filed. Frames advanced over a fixed 20s window, input pinned, one emulator at a time.

    arm       frames/20s   fps      native_exc   hook_fb      smc_mismatch
    c              605     29.97       17,737    7,519,763         0
    llvm-new        16      0.25          343       37,481         0
    

    Three things that narrow it:

    It is not SMC fallback. smc_mismatch=0 on both arms — the module was built from the modded DOL, so the chunk hashes match and nothing is being forced to the interpreter. The stall happens with the injected code fully native.

    It is not stuck in the injected region. Dispatch sampling shows nothing at or near 0x80005300. The hot sites are:

    800a128c  samples=11719
    800a183c  samples=10586
    8000550c  samples=7
    800d41e8  samples=4
    

    The guest reaches an OS wait loop at 0x800a128c/0x800a183c and never leaves it.

    The signature is interrupt starvation, not a codegen fault. native_exc is 343 against C's 17,737 on the same state. The guest is executing (174.8M native dispatches) but almost no exceptions are being delivered — it is spinning on an event that never arrives.

    For what it is worth, 0x800a128c is the same wait loop that showed up while diagnosing the interpreter-fallback hotspot in OSRestoreInterrupts on the clean DOL. In that case the guest was likewise burning cycles in that loop while native_exc sat at ~300 against a healthy ~12,000, and the cause was that the address the runtime returned to after mtmsr was not enterable, so the interrupt was never delivered there.

    That suggests a reframing worth testing: the patched branch at 0x800A3930 may leave the guest re-entering at an address the LLVM backend did not make enterable, with the visible failure being the wait loop rather than anything wrong with the three injected instructions. If so this is closer to an entry-point/dispatch-boundary problem than a "code discovered in former padding" problem, and the C backend copes because it labels every address.

    I have not confirmed that mechanism here — the above is what the counters and dispatch profile show, not a root cause. Happy to dig further if useful; the repro is cheap now that the modded DOL and matching states are validated.

  2. dougchansan commented on Sep 1, 2026

    @dougchansan
    ContributorAuthor

    Ran a control that I think changes what this issue is about. The LLVM module I reproduced with stalls on the clean DOL too, so the injected code is not the variable.

    The control

    Same module vintage, same recompiler binary, only the DOL and state differ:

    arm / DOL                        frames/20s   native_exc   top dispatch pc
    llvm  CLEAN dol,  clean-field2         7           93        800a128c
    C     CLEAN dol,  clean-field2       653       18,971        8009df3c
    llvm  MODDED dol, part1-bar-cutscene   16          343        800a128c
    C     MODDED dol, part1-bar-cutscene  684       18,412        800feb9c
    

    The LLVM arm dies the same way on both DOLs — same wait loop at 0x800a128c, same interrupt starvation. The modded DOL is not what breaks it.

    The confound is module vintage. The modules in question were built 2026-08-21 with dolrecomp 6bdb3831ef75. A current unfixed LLVM build (3166222f85ec) runs clean-field2 at 554 frames, so whatever broke that August build has since been addressed.

    Rebuilt from the modded DOL with a current recompiler

    Same modded DOL, module regenerated with a current build (includes the post-mtmsr region-leader change from #28):

    state                  current LLVM        Aug-21 LLVM      C
    part1-bar-cutscene     208 f / 28.4 fps    15 f / 0.26      734 f
    part1-in-town            0 f / 34.7 fps     7 f / 0.27      823 f
    willie-battle            0 f / 19.0 fps    10 f / 0.40      363 f
    

    Large improvement over the old build — part1-bar-cutscene goes 15 → 208 frames with native_exc 343 → 8,873 — but it is not fully fixed: still well short of C, and two states show a static guest frame_count alongside a healthy host fps, which I cannot yet account for. I would not claim this issue is resolved.

    The trampoline, for reference

    0x800A3930:  b     0x80005300        (was: fmr f29,f2)
    0x80005300:  lfs   f29, -20348(r2)
    0x80005304:  fmuls f29, f29, f0
    0x80005308:  b     0x800A3934        <- returns mid-function
    

    The return target 0x800A3934 is an address nothing branched to in the original image. That is the same "must be enterable" shape as #28, so it may still matter for this title even though it is not what caused the stall I reproduced.

    Suggested reframing

    The reproduction in the original report may have been measuring a general LLVM-backend defect of that period rather than anything mod-specific. Worth re-running against a current build before treating "code injected into former padding" as the cause — on current builds the modded DOL runs, just not yet at parity with C.

  3. dougchansan commented on Sep 2, 2026

    @dougchansan
    ContributorAuthor

    Correcting my previous comment. I said the modded DOL was not the variable, based on an old module stalling on both DOLs. That reasoning was right about that module and wrong about the issue. With a current build there is a real modded-DOL defect, and it is sharper than the original report describes.

    What actually happens

    Module rebuilt from the modded DOL with a current recompiler (includes the post-mtmsr region-leader change from #28). Run from part1-in-town.sav, captured on that same DOL:

    state=running
    booted=1
    fps=28.6795
    speed=0.324565
    frame_count=176641      <- after ~90 seconds
    

    It boots, loads the state, and the guest executes — 174.8M native dispatches, non-zero speed — but frame_count never leaves the savestate's own value. Zero video frames completed, ever. C on the identical DOL and state was past 176953 on first sample and kept climbing.

    Dispatch sampling puts it in an OS wait loop:

    800a128c  samples=11719
    800a183c  samples=10586
    

    with native_exc at 343 against C's 17,737 on the same state. It is starved of the interrupt that would let it leave the loop, not stuck executing bad code.

    Reproducibility

    part1-in-town and willie-battle fail every time (3 separate sessions). part1-bar-cutscene and part1-phenac-city sometimes run. Controls interleaved immediately before and after each failure — clean-DOL LLVM, and C on the modded DOL — passed each time, so this is not drift.

    What it is not

    • Not the injected code. Forcing 80005300-80005320 and 800A3900-800A3960 to the interpreter via STATICRECOMP_FALLBACK_RANGES does not clear it.
    • Not SMC. smc_mismatch=0; the module was built from the modded DOL so hashes match.
    • Not the backend generally. Mario Kart 1052 frames, Luigi's Mansion 419, clean-DOL Colosseum 639, all with the same LLVM lineage.
    • Not I/O or the host. Module is page-cache warm (0s read), GPU healthy, 81 GB RAM free, CPU 12%.

    The trampoline, for reference

    0x800A3930:  b     0x80005300        (was: fmr f29,f2)
    0x80005300:  lfs   f29, -20348(r2)
    0x80005304:  fmuls f29, f29, f0
    0x80005308:  b     0x800A3934        <- returns mid-function
    

    0x800A3934 is an address nothing branched to in the original image. Since forcing those ranges to the interpreter does not help, the trampoline may not be the trigger directly — but the presence of a branch target the original image never had is the kind of thing that changes region formation elsewhere, which would fit an entry-boundary explanation rather than a codegen one.

    Two measurement notes for anyone reproducing

    booted=1 can be written after a 90-120s poll gives up, so an apparent "fails to boot" here is often "slow to publish status" — check status.txt directly rather than trusting a poll timeout. And fps in status.txt is a last-written value, not a heartbeat: a stalled run keeps reporting its final fps indefinitely. Use the frame_count delta plus whether automation commands are still being consumed.

  4. dougchansan commented on Sep 2, 2026

    @dougchansan
    ContributorAuthor

    More elimination. Summary first: both LLVM configurations fail, only C works, and it is the DOL rather than the savestates.

    state-in-memory fails too

    I had only tested the default LLVM configuration. --state-in-memory behaves the same:

    arm                          frames/20s   cmds consumed
    C          modded in-town        265           yes
    llvm       modded in-town          0           no    HUNG
    llvm --state-in-memory            0           no    HUNG
    llvm --state-in-memory  willie     0           no    HUNG
    llvm --state-in-memory  bar-cut    0           no    HUNG
    

    --state-in-memory also fails part1-bar-cutscene, which the default configuration sometimes survives. So this is not specific to one codegen mode.

    It is not the savestates

    The part1-* states were captured in August on an older build, so I captured a fresh one: modded DOL, C module, saved through the automation protocol minutes before the test.

    C     fresh state    279 frames   OK
    llvm  fresh state      0 frames   HUNG
    sim   fresh state      0 frames   HUNG
    sim   old state        0 frames   HUNG
    

    A brand new state on the modded DOL hangs exactly the same way. Stale savestates are excluded.

    It is not the modified regions

    Forcing the changed code to the interpreter via STATICRECOMP_FALLBACK_RANGES does not help:

    80005300-80005320                    HUNG   (the three injected instructions)
    80005300-80005500                    HUNG   (the whole chunk, as named by the SMC diagnostic)
    80005000-80005600                    HUNG   (wider)
    800A3800-800A3A00                    HUNG   (the patch-site chunk)
    80005300-80005500,800A3800-800A3A00  HUNG   (both)
    

    It is not chunk structure

    Modded and clean builds emit an identical 4889 chunks. The injected code lands inside a chunk that was previously zero padding; it does not create a new code range.

    The guest never reaches the patched code

    Dispatch sampling on a hung run:

    800a128c  samples=11719
    800a183c  samples=10586
    8000550c  samples=7
    800d41e8  samples=4
    

    Nothing in 800A39xx or 800053xx. The guest parks in an OS wait loop with native_exc at 343 against C's 17,737 on the same state — starved of the interrupt that would let it proceed. It never executes the trampoline at all.

    What that leaves

    Four changed words in the DOL break the LLVM backend somewhere the guest reaches before touching them, and the failure mode is interrupt starvation rather than bad code at the patch site. Forcing either modified region to the interpreter does not help, so the defect is not the codegen for those instructions.

    Also worth noting the priority framing: for anyone using the widescreen mod, this is not a corner case — it is the difference between running the mod on the C backend and being unable to use the LLVM backend at all.

    Reproduction notes

    • Controls interleaved immediately before and after every failure (C on the modded DOL, LLVM on the clean DOL); both pass each time, so this is not host drift.
    • fps in status.txt is a last-written value, not a heartbeat — a hung run reports its final fps indefinitely. Use the frame_count delta plus whether .cmd files are still being consumed from the automation directory.
    • booted=1 can be published later than a short poll allows; an apparent boot failure is often a slow status write.
    • save_state needs a path relative to the automation directory; an absolute path is rejected with path=<file> is required.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions