Skip to content

Native register ABI regresses generated-code throughput ~5-10x (MKDD 51 -> 10 fps) #24

Description

@dougchansan

1bec355 ("improve llvm backend") and 0137523 ("Add native register ABI") incorporate the changes from #21, #22 and #23 -- thanks for folding those in. Measuring current main against the PR branches on the same runtime build and pinned savestates, though, shows a large performance regression:

title / scene PR #23 branch current main runtime
Mario Kart DD, bench.sav, capped 51.0 fps 10.0 fps same binary
Colosseum, Phenac scene, capped 25-29 fps 2.4 fps (vps 4.9) same binary

Counters for the Colosseum runs show the same per-cycle hook-fallback rate on both builds, but guest throughput collapses ~12x on main (1.09G cycles retired vs 13.3G in an equal window), so the cost is in the generated code itself rather than in fallback traffic. The native ABI applies automatically per function via nativeABIFlags() with no opt-out flag, so there is no way to A/B it from a built module -- an env or CLI gate would make this easy to bisect.

Setup notes for reproducing: Windows x86-64, clang toolchain modules, runtime with ExpansionPak/RecompCore#18 applied (that PR is the delivery half of #23's fix -- without it the mtmsr exit in emitStateWrite creates the boundary but pending external interrupts still wait for slice granularity, and Colosseum's retrace wait starves back to ~0.25 fps).

Activity

  1. dougchansan commented on Aug 24, 2026

    @dougchansan
    ContributorAuthor

    Correction: the headline numbers in this issue do not hold up. Closing.

    The original measurements were single first-runs on freshly generated modules, which include shader-compilation warmup, on a machine that shows large run-to-run variance. Re-measured warm and interleaved today:

    title / scene current main PR-era build verdict
    MKDD bench.sav (interleaved 2+2) 45.8, 47.1 fps 39.1, 52.3 fps no signal; box variance dominates
    Colosseum Phenac, warm 24.2, 27.1 fps 25-29 fps equal
    Luigi's Mansion foyer, uncapped 45.0 fps (1.51x) 44.1 fps (1.50x) equal

    The "51 -> 10" and "28 -> 2.4" figures were cold-run artifacts, not the native ABI. Apologies for the noise -- the ABI gate suggestion stands only as a nice-to-have for future A/B work, not as a regression need.

    Method note for anyone benchmarking these builds: never take the first run on a newly generated module as a measurement; it compiles shaders.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions