diff --git a/compiler/src/PureScript/Backend/Wasm/MiddleEnd/Optimize/DictElim.purs b/compiler/src/PureScript/Backend/Wasm/MiddleEnd/Optimize/DictElim.purs index 3ddfbf2..f2f0208 100644 --- a/compiler/src/PureScript/Backend/Wasm/MiddleEnd/Optimize/DictElim.purs +++ b/compiler/src/PureScript/Backend/Wasm/MiddleEnd/Optimize/DictElim.purs @@ -18,6 +18,7 @@ module PureScript.Backend.Wasm.MiddleEnd.Optimize.DictElim ( buildCtx , simplifyModule , summarize + , normalFormSizeCap ) where import Prelude @@ -132,8 +133,30 @@ summarize keepKeys m = m { decls = Array.filter keep m.decls } useNbE :: Boolean useNbE = true +-- | A reduced declaration larger than this many IR nodes is re-reduced with the inline +-- | context emptied (ADR 0035, the Layer C size guard). Inlining exists to *shrink* code by +-- | firing redexes; when a binding instead inlines into a normal form orders of magnitude +-- | larger than any real declaration — the canonical case is the `genericShow` dictionary of a +-- | large derived-`Generic` ADT inlined into a `show`, which produces no redex, only bulk — it is +-- | pure code-size blow-up. Falling back to the un-inlined form keeps that dictionary an ordinary +-- | call, the same (correct) shape `--no-opt` emits, and is what bounds NbE when a program is +-- | itself `show`-heavy (notably the compiler compiling itself). The threshold sits far above any +-- | genuine declaration (tens of thousands of nodes) and far below the observed blow-ups +-- | (5×10⁵–2×10⁶), so only pathological declarations fall back; the guard measures the *actual* +-- | reduced size, so a large-but-shared normal form (which quote CSEs back down) is kept inlined. +normalFormSizeCap :: Int +normalFormSizeCap = 200_000 + reduce :: Ctx -> M.Expr -> M.Expr -reduce ctx = if useNbE then normalize ctx else simplifyExpr ctx +reduce ctx e = + let + r = reduce1 ctx e + in + if exprSize r > normalFormSizeCap && ctxInlines ctx then reduce1 (ctx { inline = Map.empty, instanceFields = Map.empty }) e + else r + where + reduce1 c = if useNbE then normalize c else simplifyExpr c + ctxInlines c = not (Map.isEmpty c.inline) || not (Map.isEmpty c.instanceFields) simplifyModule :: Ctx -> M.Module -> M.Module simplifyModule ctx m = m { decls = map go m.decls } diff --git a/compiler/src/PureScript/Backend/Wasm/MiddleEnd/Optimize/Semantics.purs b/compiler/src/PureScript/Backend/Wasm/MiddleEnd/Optimize/Semantics.purs index 0bbc888..1663db9 100644 --- a/compiler/src/PureScript/Backend/Wasm/MiddleEnd/Optimize/Semantics.purs +++ b/compiler/src/PureScript/Backend/Wasm/MiddleEnd/Optimize/Semantics.purs @@ -19,9 +19,12 @@ module PureScript.Backend.Wasm.MiddleEnd.Optimize.Semantics import Prelude -import Control.Monad.State (State, evalState, get, put) +import Control.Monad.State (State, get, modify_, put, runState) import Data.Array as Array import Data.Either (Either(..)) +import Data.Lazy (Lazy, defer, force) +import Data.List (List(..), (:)) +import Data.List as List import Data.Map (Map) import Data.Map as Map import Data.Maybe (Maybe(..), fromMaybe) @@ -58,6 +61,12 @@ data Sem -- a recursive group, never unfolded: bindings (bodies already evaluated with the group -- names opaque) plus the continuation semantic value. | SLetRec (Array RecB) Sem + -- a value shared across use sites, tagged with the memo key it was unfolded from (ADR 0035 + -- Layer B). Transparent to *reduction* — `unShared` strips it wherever a value is consumed + -- (β / projection / match / perform) so it never blocks a redex — but opaque to `quote`, which + -- CSEs it: the first occurrence is bound once to a hoisted `let` and the rest reference it, + -- killing the M2 re-quote explosion. The key is the inline binding's name (already unique). + | SShared String Sem | SNeu Neu type RecB = { meta :: Maybe Meta, ident :: String, expr :: Sem } @@ -85,7 +94,27 @@ data Match -- | Normalise an expression: evaluate to `Sem`, then quote back to IR. normalize :: Ctx -> M.Expr -> M.Expr -normalize ctx e = evalState (quote (pctxOf ctx) (eval ctx Set.empty Map.empty e)) 0 +normalize ctx e = + let + Tuple body st = runState (quote (pctxOf ctx) (eval ctx memo Set.empty Map.empty e)) initialQ + -- the CSE'd shared values, hoisted as a `let` at the declaration top. `defs` is most-recent- + -- first and a binding may reference an earlier one (quoting a value records its dependencies + -- first), so reverse to dependency order. They are closed, so a single non-recursive group works. + defs = Array.fromFoldable (List.reverse st.defs) + in + if Array.null defs then body + else M.Let (map (\(Tuple n ex) -> M.NonRec Nothing n ex) defs) body + where + initialQ = { counter: 0, shared: Map.empty, defs: Nil } + + -- ADR 0035 Layer A: evaluate each inline-set binding **once** and share the resulting `Sem` at + -- every use site, rather than re-evaluating its body per reference (the M1 exponential). The memo + -- is keyed by binding name; `evalVar` forces it. The inline set is a DAG — it is built from + -- `NonRec` infos in `DictElim.buildCtx`, and mutual recursion is a `Rec` group `infoOf` skips — + -- so the lazy thunks force without re-entrancy, and a diamond returns the same shared `Sem` on + -- every path (sound because, being acyclic, a binding never references one of its own ancestors). + memo :: Map String (Lazy Sem) + memo = map (\body -> defer \_ -> eval ctx memo Set.empty Map.empty body) ctx.inline pctxOf :: Ctx -> PCtx pctxOf ctx = { eff: ctx.effectfulForeigns, impure: ctx.impureBindings, memEff: ctx.memEffBindings } @@ -96,13 +125,13 @@ pctxOf ctx = { eff: ctx.effectfulForeigns, impure: ctx.impureBindings, memEff: c -- | path; a reference to one is left as a call (`NTop`) rather than re-unfolded, which -- | breaks cycles in the inline set and bounds unfolding depth (the rule-based engine -- | tolerated cycles only via its `maxPasses` fuel; NbE has no fuel, so it needs this). -eval :: Ctx -> Set String -> Env -> M.Expr -> Sem -eval ctx = go +eval :: Ctx -> Map String (Lazy Sem) -> Set String -> Env -> M.Expr -> Sem +eval ctx memo = go where pctx = pctxOf ctx - go visited env = case _ of - M.Var q -> evalVar visited env q + go visited env e = case e of + M.Var q -> evalVar env q -- a record literal is the one literal with a reduction (field projection / update), -- so it gets its own semantic value; all other literals are inert `SLit` M.Lit (LitObject kvs) -> SRecord (map (map (go visited env)) kvs) @@ -110,11 +139,11 @@ eval ctx = go e@(M.Constructor _ _ _) -> SNeu (NCtorDecl e) M.Abs ps body -> SLam ps (\args -> go visited (bindParams env ps args) body) M.App f args -> evalApp visited env f args - M.Accessor l e -> accessor ctx visited l (go visited env e) + M.Accessor l e -> accessor ctx memo visited l (go visited env e) M.Update e mb kvs -> update (go visited env e) mb (map (map (go visited env)) kvs) M.Perform e -> performSem pctx (go visited env e) e - M.Case scruts alts -> evalCase ctx visited env scruts alts - M.Let binds body -> evalLet ctx visited env binds body + M.Case scruts alts -> evalCase ctx memo visited env scruts alts + M.Let binds body -> evalLet ctx memo visited env binds body -- short-circuit the boolean operators into control flow (matching PureScript/JS -- semantics and avoiding the strict `i32.or`/`i32.and`), before they apply as foreigns @@ -124,14 +153,14 @@ eval ctx = go | qkey q == Just boolConjKey -> go visited env (boolCase a b (M.Lit (LitBoolean false))) _, _ -> apply (go visited env f) (map (go visited env) args) - evalVar visited env = case _ of + evalVar env = case _ of Qualified Nothing x -> fromMaybe (SNeu (NLocal x)) (Map.lookup x env) q@(Qualified (Just _) _) -> case qkey q of Just k - -- unfold a binding in the inline set, unless it is already being unfolded on - -- this path (a cycle / self-reference) — then leave it as a call (ADR 0020). - -- Re-evaluated at each use site, reproducing the current inline policy. - | Just body <- Map.lookup k ctx.inline -> if Set.member k visited then SNeu (NTop q) else go (Set.insert k visited) Map.empty body + -- unfold a binding in the inline set (ADR 0020), but **once** per `normalize`: every + -- reference shares the memoized `Sem` (ADR 0035 Layer A) instead of re-evaluating the + -- body. `memo` has an entry for exactly the inline-set keys (`ctx.inline`). + | Just lz <- Map.lookup k memo -> SShared k (force lz) | Set.member k ctx.dataCtors -> SCtorApp q [] | otherwise -> SNeu (NTop q) Nothing -> SNeu (NTop q) @@ -139,10 +168,18 @@ eval ctx = go -- | Apply a semantic value to arguments — β when the head is a known lambda (with -- | arity handling), accumulation for a known constructor, otherwise a stuck `NApp` -- | (flattened so a curried spine becomes one n-ary application). +-- | Strip the `SShared` tag wherever a value is *consumed* by a reduction, so sharing never +-- | blocks a redex (the tag exists only to let `quote` CSE an unconsumed shared value). +unShared :: Sem -> Sem +unShared = case _ of + SShared _ s -> unShared s + s -> s + apply :: Sem -> Array Sem -> Sem apply head args | Array.null args = head | otherwise = case head of + SShared _ s -> apply s args SLam ps fn -> let np = Array.length ps @@ -158,12 +195,13 @@ apply head args bindParams :: Env -> Array String -> Array Sem -> Env bindParams env ps args = Array.foldl (\e (Tuple p a) -> Map.insert p a e) env (Array.zip ps args) -accessor :: Ctx -> Set String -> String -> Sem -> Sem -accessor ctx visited l = case _ of +accessor :: Ctx -> Map String (Lazy Sem) -> Set String -> String -> Sem -> Sem +accessor ctx memo visited l = case _ of + SShared _ s -> accessor ctx memo visited l s SRecord fs -> fromMaybe (SNeu (NAccessor l (SRecord fs))) (lookupSem l fs) -- a transparent (newtype / dictionary) constructor is the identity on its payload, so -- a field read sees through it: `Dict({…}).l` is `{…}.l` - SCtorApp q [ payload ] | Just k <- qkey q, Set.member k ctx.newtypeCtors -> accessor ctx visited l payload + SCtorApp q [ payload ] | Just k <- qkey q, Set.member k ctx.newtypeCtors -> accessor ctx memo visited l payload s@(SNeu (NTop q)) | Just k <- qkey q , Just fields <- Map.lookup k ctx.instanceFields @@ -173,12 +211,13 @@ accessor ctx visited l = case _ of -- `heytingAlgebraBoolean.implies` calls its own `.disj`): mark it visited so the -- projected field's own back-reference stays a call rather than looping. case Array.find (\(Tuple fl _) -> fl == l) fields of - Just (Tuple _ fieldExpr) -> eval ctx (Set.insert k visited) Map.empty fieldExpr + Just (Tuple _ fieldExpr) -> eval ctx memo (Set.insert k visited) Map.empty fieldExpr Nothing -> SNeu (NAccessor l s) s -> SNeu (NAccessor l s) update :: Sem -> Maybe (Array String) -> Array (Tuple String Sem) -> Sem update e mb kvs = case e of + SShared _ s -> update s mb kvs SRecord fs -> SRecord (Array.foldl overwrite fs kvs) _ -> SNeu (NUpdate e mb kvs) where @@ -187,53 +226,57 @@ update e mb kvs = case e of else Array.snoc fs (Tuple k v) performSem :: PCtx -> Sem -> M.Expr -> Sem -performSem pctx se orig = case se of - -- performing a literal thunk runs its body (apply to the unit); always sound, even - -- for an effectful body — this is the pure-`Effect` collapse's entry point - SLam _ _ -> apply se [ unit_ ] - _ - | runPure pctx orig -> apply se [ unit_ ] - | otherwise -> SNeu (NPerform se) +performSem pctx se orig = + let + s = unShared se + in + case s of + -- performing a literal thunk runs its body (apply to the unit); always sound, even + -- for an effectful body — this is the pure-`Effect` collapse's entry point + SLam _ _ -> apply s [ unit_ ] + _ + | runPure pctx orig -> apply s [ unit_ ] + | otherwise -> SNeu (NPerform s) where unit_ = SLit (LitInt 0) -- case ------------------------------------------------------------------------ -evalCase :: Ctx -> Set String -> Env -> Array M.Expr -> Array M.Alt -> Sem -evalCase ctx visited env scrutsE alts = +evalCase :: Ctx -> Map String (Lazy Sem) -> Set String -> Env -> Array M.Expr -> Array M.Alt -> Sem +evalCase ctx memo visited env scrutsE alts = let - scruts = map (eval ctx visited env) scrutsE + scruts = map (eval ctx memo visited env) scrutsE in - case selectAlt ctx visited env scruts alts of + case selectAlt ctx memo visited env scruts alts of Just sem -> sem - Nothing -> SNeu (NCase scruts (map (evalAlt ctx visited env) alts)) + Nothing -> SNeu (NCase scruts (map (evalAlt ctx memo visited env) alts)) -- | Select the first alternative that definitely matches (binding its sub-patterns), -- | stopping — and leaving the whole `case` — at the first undecidable or guarded -- | alternative. Mirrors `Simplify.caseOfKnown`. -selectAlt :: Ctx -> Set String -> Env -> Array Sem -> Array M.Alt -> Maybe Sem -selectAlt ctx visited env scruts = go +selectAlt :: Ctx -> Map String (Lazy Sem) -> Set String -> Env -> Array Sem -> Array M.Alt -> Maybe Sem +selectAlt ctx memo visited env scruts = go where go alts = case Array.uncons alts of Nothing -> Nothing Just { head: alt, tail } -> case alt.result of Right body -> case matchAllSem ctx alt.binders scruts of - MYes subs -> Just (eval ctx visited (Array.foldl (\e (Tuple x s) -> Map.insert x s e) env subs) body) + MYes subs -> Just (eval ctx memo visited (Array.foldl (\e (Tuple x s) -> Map.insert x s e) env subs) body) MNo -> go tail MUnknown -> Nothing Left _ -> Nothing -- | Evaluate an alternative's body(ies) with its binder variables bound to opaque -- | locals (the scrutinee is not known, so the bound values are not either). -evalAlt :: Ctx -> Set String -> Env -> M.Alt -> NAlt -evalAlt ctx visited env alt = +evalAlt :: Ctx -> Map String (Lazy Sem) -> Set String -> Env -> M.Alt -> NAlt +evalAlt ctx memo visited env alt = let env' = Array.foldl (\e v -> Map.insert v (SNeu (NLocal v)) e) env (alt.binders >>= binderVars) in { binders: alt.binders , result: case alt.result of - Right e -> Right (eval ctx visited env' e) - Left gs -> Left (map (\g -> { guard: eval ctx visited env' g.guard, expression: eval ctx visited env' g.expression }) gs) + Right e -> Right (eval ctx memo visited env' e) + Left gs -> Left (map (\g -> { guard: eval ctx memo visited env' g.guard, expression: eval ctx memo visited env' g.expression }) gs) } matchAllSem :: Ctx -> Array Binder -> Array Sem -> Match @@ -253,14 +296,15 @@ matchSem ctx = case _, _ of | Just k <- qkey ctor, Set.member k ctx.newtypeCtors -> case subs of [ sub ] -> matchSem ctx sub s -- transparent newtype: the value is its payload _ -> MUnknown - | Just k <- qkey ctor -> case s of + | Just k <- qkey ctor -> case unShared s of SCtorApp q cargs | qkey q == Just k -> matchAllSem ctx subs cargs | otherwise -> MNo _ -> MUnknown | otherwise -> MUnknown - LiteralBinder _ lit, SLit slit -> matchLitSem lit slit - LiteralBinder _ _, _ -> MUnknown + LiteralBinder _ lit, s -> case unShared s of + SLit slit -> matchLitSem lit slit + _ -> MUnknown matchLitSem :: Literal Binder -> Literal Sem -> Match matchLitSem = case _, _ of @@ -287,22 +331,22 @@ combine = case _, _ of -- | single bindings (order preserved), each inlined or retained per the current gates -- | (single-use/dead pure, trivial record, small lambda); a recursive group is retained, -- | its bound variables opaque (never unfolded — that is the infinite-loop hazard). -evalLet :: Ctx -> Set String -> Env -> Array M.Bind -> M.Expr -> Sem -evalLet ctx visited env binds body = case Array.uncons binds of - Nothing -> eval ctx visited env body +evalLet :: Ctx -> Map String (Lazy Sem) -> Set String -> Env -> Array M.Bind -> M.Expr -> Sem +evalLet ctx memo visited env binds body = case Array.uncons binds of + Nothing -> eval ctx memo visited env body Just { head: M.NonRec _ x rhs, tail } -> let - rhsSem = eval ctx visited env rhs - rest e = evalLet ctx visited e tail body + rhsSem = eval ctx memo visited env rhs + rest e = evalLet ctx memo visited e tail body in if inlineLet ctx x rhs body then rest (Map.insert x rhsSem env) else SLet x rhsSem (\xv -> rest (Map.insert x xv env)) Just { head: M.Rec rs, tail } -> let env' = Array.foldl (\e r -> Map.insert r.ident (SNeu (NLocal r.ident)) e) env rs - rs' = map (\r -> { meta: r.meta, ident: r.ident, expr: eval ctx visited env' r.expr }) rs + rs' = map (\r -> { meta: r.meta, ident: r.ident, expr: eval ctx memo visited env' r.expr }) rs in - SLetRec rs' (evalLet ctx visited env' tail body) + SLetRec rs' (evalLet ctx memo visited env' tail body) -- | Whether a `let` binding should be inlined rather than retained — the *current* -- | policy (ADR 0020 stage 2): single-use or dead and pure, or a trivial record, or a @@ -315,10 +359,17 @@ inlineLet ctx x rhs body = -- quote ----------------------------------------------------------------------- -type Q = State Int +-- | `quote`'s state (ADR 0035 Layer B): the fresh-binder counter, plus the CSE table for shared +-- | values — `shared` maps a memo key to the `let` name its quoted value was bound to, and `defs` +-- | accumulates those `(name, expr)` bindings (most-recent-first) to hoist at the top of the +-- | normalized declaration. The shared values are inline bindings, evaluated with an empty +-- | environment, so their quoted expressions are closed and hoist without capture. +type QState = { counter :: Int, shared :: Map String String, defs :: List (Tuple String M.Expr) } + +type Q = State QState quote :: PCtx -> Sem -> Q M.Expr -quote pctx = case _ of +quote pctx sem = case sem of -- merge a directly-nested lambda into one parameter list (disjoint params), so a -- curried worker `\n -> \s -> …` becomes the arity-2 `\n s -> …` whose saturated -- self-call is a direct, tail-callable call (constant-stack TCE; ADR 0015). Binders @@ -343,6 +394,20 @@ quote pctx = case _ of rs' <- traverse (\r -> (\e -> { meta: r.meta, ident: r.ident, expr: e }) <$> quote pctx r.expr) rs body' <- quote pctx body pure (M.Let [ M.Rec rs' ] body') + -- CSE a shared value (ADR 0035 Layer B): quote it **once** into a hoisted `let` (recorded in + -- `defs`) and emit a reference; every later occurrence of the same key reuses the binding. The + -- name is reserved before quoting the body, so a (defensive) self/mutual reference reuses it + -- rather than recursing forever. + SShared k s -> do + st <- get + case Map.lookup k st.shared of + Just name -> pure (M.Var (Qualified Nothing name)) + Nothing -> do + name <- fresh "$shared" + modify_ \st' -> st' { shared = Map.insert k name st'.shared } + e <- quote pctx s + modify_ \st' -> st' { defs = Tuple name e : st'.defs } + pure (M.Var (Qualified Nothing name)) SNeu n -> quoteNeu pctx n quoteNeu :: PCtx -> Neu -> Q M.Expr @@ -395,9 +460,9 @@ fresh base0 = do base = case String.indexOf (Pattern "$q") base0 of Just i -> String.take i base0 Nothing -> base0 - n <- get - put (n + 1) - pure (base <> "$q" <> show n) + st <- get + put st { counter = st.counter + 1 } + pure (base <> "$q" <> show st.counter) mergeAbs :: Array String -> M.Expr -> M.Expr mergeAbs ps = case _ of diff --git a/compiler/src/PureScript/Backend/Wasm/MiddleEnd/Optimize/Specialize.purs b/compiler/src/PureScript/Backend/Wasm/MiddleEnd/Optimize/Specialize.purs index 9ead761..8ffabd6 100644 --- a/compiler/src/PureScript/Backend/Wasm/MiddleEnd/Optimize/Specialize.purs +++ b/compiler/src/PureScript/Backend/Wasm/MiddleEnd/Optimize/Specialize.purs @@ -36,7 +36,6 @@ import Prelude import Control.Monad.State (State, gets, modify_, runState) import Data.Array as Array import Data.Either (Either(..)) -import Data.Foldable (foldl) import Data.Map (Map) import Data.Map as Map import Data.Maybe (Maybe(..), fromMaybe) @@ -48,6 +47,7 @@ import Data.Tuple (Tuple(..), snd) import PureScript.Backend.Wasm.MiddleEnd.FreeVars (binderVars, freeVars) import PureScript.Backend.Wasm.MiddleEnd.IR as M import PureScript.Backend.Wasm.MiddleEnd.Optimize.Simplify (substMany) +import PureScript.Backend.Wasm.MiddleEnd.Serialize.Hash (hashString) import PureScript.CoreFn (Literal(..), ModuleName, Qualified(..)) -- | A top-level function that is a candidate callee: its parameter list, body, and @@ -327,42 +327,25 @@ lambdaFrees = case _ of e -> freeVars [] e -- a structural key for the lambda with its free variables abstracted, so two --- lambdas equal up to their captures share a specialization +-- lambdas equal up to their captures share a specialization. The free vars are +-- abstracted to positional markers in a single capture-avoiding pass (rather than one +-- `substVar` per free), and the canonical form is then **hashed** rather than used +-- verbatim: the raw `show` of a large lambda body is a multi-kilobyte string, and it was +-- both built (genericShow) and *retained as a `Map` key* for every specialization +-- candidate — quadratic-ish key comparisons plus heavy GC that dominated the optimization of +-- large higher-order modules. A 16-hex digest keeps the dedup `Map` keys tiny and fixed-size. +-- A clash would mis-share two distinct specializations, but the digest is the same one already +-- trusted for the build-cache identity, so an accidental clash is astronomically unlikely. canonicalKey :: Array String -> M.Expr -> String canonicalKey frees lam = - show (foldlWithIndexArr (\i e f -> substVar f (M.Var (Qualified Nothing ("#" <> show i))) e) lam frees) + hashString (show (substMany markers lam)) + where + markers = + Map.fromFoldable + (Array.mapWithIndex (\i f -> Tuple f (M.Var (Qualified Nothing ("#" <> show i)))) frees) -- substitution / helpers ------------------------------------------------------ --- replace free occurrences of local `name` with `repl`, stopping at shadowing -substVar :: String -> M.Expr -> M.Expr -> M.Expr -substVar name repl = go - where - go = case _ of - e@(M.Var (Qualified Nothing n)) -> if n == name then repl else e - e@(M.Var _) -> e - M.Lit lit -> M.Lit (mapLit go lit) - e@(M.Constructor _ _ _) -> e - M.Accessor l e -> M.Accessor l (go e) - M.Update e cf kvs -> M.Update (go e) cf (map (map go) kvs) - M.Abs ps b -> if Array.elem name ps then M.Abs ps b else M.Abs ps (go b) - M.App f a -> mkApp (go f) (map go a) - M.Perform e -> M.Perform (go e) - M.Case ss alts -> M.Case (map go ss) (map goAlt alts) - M.Let bs body -> - if Array.elem name (bs >>= boundNames) then M.Let bs body - else M.Let (map goBind bs) (go body) - goAlt alt = - if Array.elem name (alt.binders >>= binderVars) then alt - else alt - { result = case alt.result of - Right e -> Right (go e) - Left gs -> Left (map (\g -> { guard: go g.guard, expression: go g.expression }) gs) - } - goBind = case _ of - M.NonRec meta i e -> M.NonRec meta i (go e) - M.Rec rs -> M.Rec (map (\r -> r { expr = go r.expr }) rs) - -- keep `App` heads flat (never an `App` of an `App`) mkApp :: M.Expr -> Array M.Expr -> M.Expr mkApp head args @@ -404,9 +387,6 @@ litExprs = case _ of LitObject kvs -> map snd kvs _ -> [] -foldlWithIndexArr :: forall a b. (Int -> b -> a -> b) -> b -> Array a -> b -foldlWithIndexArr f z arr = foldl (\acc (Tuple i a) -> f i acc a) z (Array.mapWithIndex Tuple arr) - key :: ModuleName -> String -> String key modName ident = joinWith "." modName <> "." <> ident diff --git a/compiler/test/NbeStress.purs b/compiler/test/NbeStress.purs new file mode 100644 index 0000000..1975125 --- /dev/null +++ b/compiler/test/NbeStress.purs @@ -0,0 +1,118 @@ +-- | The NbE exponential regression guard (bug A, ADR 0035 §8). +-- | +-- | It drives `Semantics.normalize` directly against a synthetic **diamond inline DAG**: an +-- | inline set `b0 … bd` where each `bᵢ = f b₍ᵢ₋₁₎ b₍ᵢ₋₁₎` references the previous binding twice. +-- | Normalizing `bd` re-evaluates the shared leaf `b0` once per path = Θ(2ᵈ) on the *unfixed* +-- | reducer (M1 in eval, M2 in quote), so the result size **doubles every depth** and the wall +-- | clock explodes around d≈22-25. After the ADR-0035 sharing fix (Layer A eval memo + Layer B +-- | quote CSE) the result stays **O(d)** (≈ 4d) and the loop runs instantly — verified without +-- | building the whole compiler to wasm. +-- | +-- | `spec` is the routine `test:unit` guard: the linear-bound diamond check above, plus a guard for +-- | the **Layer C size cap** (`DictElim.simplifyModule` falls back to the un-inlined form when a +-- | declaration inlines genuine, unshareable bulk — the `genericShow`-into-`show` blow-up that hung +-- | self-compilation). `main` sweeps deeper diamond depths for manual inspection: +-- | +-- | spago test -p compiler -m Test.NbeStress +module Test.NbeStress where + +import Prelude + +import Data.Array as Array +import Data.Foldable (for_) +import Data.Map (Map) +import Data.Map as Map +import Data.Maybe (Maybe(..)) +import Data.Set as Set +import Data.Tuple (Tuple(..)) +import Effect (Effect) +import Effect.Console (error) as Console +import PureScript.Backend.Wasm.MiddleEnd.IR as M +import PureScript.Backend.Wasm.MiddleEnd.Optimize.Analysis (exprSize) +import PureScript.Backend.Wasm.MiddleEnd.Optimize.DictElim (normalFormSizeCap, simplifyModule) +import PureScript.Backend.Wasm.MiddleEnd.Optimize.Semantics (normalize) +import PureScript.CoreFn (Literal(..), Qualified(..)) +import Test.Spec (Spec, describe, it) +import Test.Spec.Assertions (shouldEqual) + +maxDepth :: Int +maxDepth = 20 + +-- A reference to inline binding `bᵢ` (module `M`), keyed so `qkey` = `"M.bᵢ"`. +ref :: Int -> M.Expr +ref i = M.Var (Qualified (Just [ "M" ]) ("b" <> show i)) + +-- An opaque (non-inline) head, so the application stays a neutral whose two operands are each +-- re-evaluated — the binary fan-out that makes the DAG a diamond rather than a chain. +opaque :: M.Expr +opaque = M.Var (Qualified (Just [ "Ext" ]) "f") + +-- The inline set for depth `d`: `b0` is a neutral leaf; `bᵢ = f b₍ᵢ₋₁₎ b₍ᵢ₋₁₎`. +diamondInline :: Int -> Map String M.Expr +diamondInline d = + Map.fromFoldable + ( Array.cons (Tuple "M.b0" (M.Var (Qualified (Just [ "Ext" ]) "leaf"))) + (map (\i -> Tuple ("M.b" <> show i) (M.App opaque [ ref (i - 1), ref (i - 1) ])) (Array.range 1 d)) + ) + +ctxFor :: Int -> { newtypeCtors :: Set.Set String, dataCtors :: Set.Set String, inline :: Map String M.Expr, instanceFields :: Map String (Array (Tuple String M.Expr)), effectfulForeigns :: Set.Set String, impureBindings :: Set.Set String, memEffBindings :: Set.Set String } +ctxFor d = + { newtypeCtors: Set.empty + , dataCtors: Set.empty + , inline: diamondInline d + , instanceFields: Map.empty + , effectfulForeigns: Set.empty + , impureBindings: Set.empty + , memEffBindings: Set.empty + } + +-- An **un-shareable** term of `n` distinct-position leaves (a flat array of int literals): unlike a +-- diamond, `quote` cannot CSE it, so its normal form really is ~`n` nodes. Inlining a binding bound +-- to this produces genuine bulk with no reduction — the `genericShow`-into-`show` pathology in +-- miniature — which is exactly what the size cap must catch. +flatBig :: Int -> M.Expr +flatBig n = M.Lit (LitArray (Array.replicate n (M.Lit (LitInt 0)))) + +-- a one-declaration module `U.user = `, the unit `simplifyModule` reduces. +userModule :: M.Expr -> M.Module +userModule body = { name: [ "U" ], decls: [ M.NonRec Nothing "user" body ] } + +-- the reduced size of `U.user` after `simplifyModule`. +userSize :: M.Module -> Int +userSize m = case Array.head m.decls of + Just (M.NonRec _ _ e) -> exprSize e + _ -> -1 + +spec :: Spec Unit +spec = do + -- | Routine guard (ADR 0035 §8): a depth-20 diamond normalizes to a *linear*-size term. On the + -- | pre-sharing reducer this is 2^20 ≈ a million nodes (seconds of work + a huge tree); the + -- | Layer A memo + Layer B quote CSE keep it ≈ 4·20 + 3 = 83. The generous `< 1000` bound passes + -- | instantly when sharing holds and fails (or times out) the moment the exponential returns. + describe "Semantics.normalize — NbE exponential guard (ADR 0035)" do + it "normalizes a depth-20 diamond inline DAG to linear size, not O(2^d)" do + let d = 20 + (exprSize (normalize (ctxFor d) (ref d)) < 1000) `shouldEqual` true + + -- | The Layer C size cap: inlining that blows the normal form past `normalFormSizeCap` falls back + -- | to the un-inlined form (the binding stays a call). This is what bounds NbE when a declaration + -- | inlines genuine, unshareable bulk — the `genericShow` dictionary of a large derived-`Generic` + -- | ADT inlined into `show`, the case that hung the compiler compiling itself. The companion case + -- | proves the cap discriminates: a reduced form *under* the cap is still inlined. + describe "DictElim.simplifyModule — code-size cap (ADR 0035 Layer C lite)" do + it "falls back to the un-inlined call when inlining blows the size cap" do + let big = M.Var (Qualified (Just [ "M" ]) "big") + let ctx = (ctxFor 0) { inline = Map.singleton "M.big" (flatBig (normalFormSizeCap + 10)) } + -- inlined it would be > cap; the cap forces the un-inlined form, so `user` stays a small call. + (userSize (simplifyModule ctx (userModule big)) < normalFormSizeCap) `shouldEqual` true + it "still inlines a binding whose reduced form is under the cap (the cap discriminates)" do + let small = M.Var (Qualified (Just [ "M" ]) "small") + let ctx = (ctxFor 0) { inline = Map.singleton "M.small" (flatBig 50) } + -- well under the cap: inlining proceeds, so `user` grows to the inlined array (≫ a bare call). + (userSize (simplifyModule ctx (userModule small)) > 50) `shouldEqual` true + +main :: Effect Unit +main = do + Console.error "NbE diamond stress — result size should be O(d); on the unfixed reducer it doubles per depth:" + for_ (Array.range 1 maxDepth) \d -> + Console.error (" d=" <> show d <> " normalized-size=" <> show (exprSize (normalize (ctxFor d) (ref d)))) diff --git a/compiler/test/Unit/Compiler.purs b/compiler/test/Unit/Compiler.purs index 8e2fd58..bf2143f 100644 --- a/compiler/test/Unit/Compiler.purs +++ b/compiler/test/Unit/Compiler.purs @@ -29,6 +29,7 @@ import Test.Unit.PureScript.Backend.Wasm.Ulib.Interface as UlibInterface import Test.Unit.PureScript.Backend.Wasm.MiddleEnd.Transl as Transl import Test.Unit.PureScript.CoreFn as CoreFn import Test.Unit.PureScript.ExternsFile as ExternsFile +import Test.NbeStress as NbeStress main :: Effect Unit main = runSpecAndExitProcess [ consoleReporter ] do @@ -37,6 +38,7 @@ main = runSpecAndExitProcess [ consoleReporter ] do Externs.spec Caf.spec Codegen.spec + NbeStress.spec Lower.spec Match.spec Transl.spec diff --git a/docs/design-decisions/0020-reduction-aware-inliner.md b/docs/design-decisions/0020-reduction-aware-inliner.md index 51565ee..8cd869d 100644 --- a/docs/design-decisions/0020-reduction-aware-inliner.md +++ b/docs/design-decisions/0020-reduction-aware-inliner.md @@ -22,6 +22,15 @@ > (it recomputes values on every `eval`/`quote` traversal), independent of the fusion motivation > here. ADR 0035 sequences the fix: a behaviour-neutral **sharing/memo** pass (the scalability gate) > first, then this ADR's reduction-aware inline-or-share **policy** on top. +> +> **Update (2026-06-17):** [ADR 0035](0035-sharing-nbe-reduction-aware-inlining.md)'s **sharing pass +> (Layers A + B) landed**, removing the NbE recomputation exponential. ~~and opened the scalability gate — the optimized self-compile now reduces +> through `Optimize.Specialize` instead of spinning.~~ _(A+B alone did **not** clear `Optimize.Specialize`; +> completing the optimized self-compile also required a Layer-C-lite `normalFormSizeCap` code-size cap +> and the `Optimize.Specialize` dedup-key fix — see [ADR 0035](0035-sharing-nbe-reduction-aware-inlining.md).)_ +> **This ADR's reduction-aware *decision* is +> ADR 0035 Layer C, now deferred**: with the exponential gone, it is an optimization-quality +> improvement (the fusion win) rather than a scalability blocker. ## Context diff --git a/docs/design-decisions/0035-sharing-nbe-reduction-aware-inlining.md b/docs/design-decisions/0035-sharing-nbe-reduction-aware-inlining.md index 3e1482b..efd28c1 100644 --- a/docs/design-decisions/0035-sharing-nbe-reduction-aware-inlining.md +++ b/docs/design-decisions/0035-sharing-nbe-reduction-aware-inlining.md @@ -1,6 +1,6 @@ # 0035. Sharing/memoizing the NbE reducer, then reduction-aware inlining -- Status: Proposed +- Status: Accepted — **Layers A + B landed 2026-06-17** (the sharing/scalability gate); Layer C (reduction-aware policy) deferred - Date: 2026-06-16 > Realizes **stage 3** of [ADR 0020](0020-reduction-aware-inliner.md) (whose NbE core landed as @@ -9,6 +9,42 @@ > self-compilation — the NbE reducer is **itself exponential** for lack of sharing — and sequences > the fix so the scalability gate is opened *before* the inline-policy rewrite. +> **Progress (2026-06-17, branch `feat/reduction-aware-inlining`).** Layers **A and B are +> implemented** in `MiddleEnd.Optimize.Semantics`. ~~The scalability gate is open: the +> optimized self-compile now compiles *through* `Optimize.Specialize` (the old hang point) — the +> exponential is gone (it then OOMs in a later module, which is the independent whole-program +> *memory* floor, not this defect).~~ **(Corrected below, 2026-06-17: A+B removed the NbE +> exponential but did *not* by themselves clear `Optimize.Specialize`; the "OOM" read was premature — +> the real remaining blocker there was *code size*, fixed by the canonicalKey + size-cap work.)** +> Layer A is a `Data.Lazy` memo (`Map String (Lazy Sem)`) keyed +> by inline-binding name; Layer B is an `SShared k Sem` tag that `unShared` strips at every +> reduction site (so it never blocks a redex) while `quote` CSEs it into a hoisted `let` (the Q +> state carries the CSE table). The identity mechanism (§Decision, "settled at build time") was +> chosen as **value-tagging**, *not* `unsafeRefEq` (rejected: not referentially transparent) nor +> ids-assigned-during-eval (rejected: would force the whole HOAS evaluator into a monad). **Byte +> equality was *not* required** — the current compiler is not yet a correct self-host reference, so +> the gate is *no regression in the test suite + benchmarks + `examples/`*, which holds: unit +> ~~160/160~~ **162/162** (the diamond O(d) guard plus the two Layer-C-lite cap guards, below), +> e2e 150/150, the `test:bin` examples, bench no +> regression (most benches 1.05–1.19× *faster* from the CSE sharing; baseline untouched), and +> `examples/metatheory` compiles and runs correctly. The exponential regression guard is +> `compiler/test/NbeStress.purs` (`Test.NbeStress.spec`, wired into `test:unit`): a depth-20 diamond +> inline DAG normalizes to **linear** size (≈ 4·d + 3) rather than 2^d. + +> **Correction (2026-06-17).** Layers A+B were **necessary but not sufficient** to compile +> `Optimize.Specialize` / complete the optimized self-compile. With the NbE *recomputation* +> exponential gone, that module still hung on two **code-size** problems orthogonal to A/B's fix +> (A+B bound recomputation, not output size): **(1)** `Optimize.Specialize.canonicalKey` built its +> de-dup `Map` key by `show`-ing the substituted lambda body — multi-KB strings, *built and retained +> as keys* — now **hashed** (`Serialize.Hash.hashString`) over a single `substMany` pass; **(2)** NbE +> inlines the derived `genericShow` dictionary of the large IR ADTs into multi-million-node normal +> forms, now bounded by **`DictElim.normalFormSizeCap`** (a *Layer C lite* size guard — distinct from +> the deferred reduction-aware *policy* of §Layer C): a reduced declaration exceeding the cap is +> re-reduced with the inline context emptied (the binding stays a call — the `--no-opt`-correct +> shape). With **all three** (A, B, and (1)+(2)) the optimized self-compile of `PursWasm.CLI.Main` +> now **completes**, writing a valid 8 MB `index.wasm`. Two cap regression guards were added to +> `Test.NbeStress` (unit 160 → 162). + ## Context Compiling `purs-wasm` with itself (806 modules in `output/`, 286 reachable from `Main` — vs the @@ -132,13 +168,21 @@ here: ## Consequences - **The default-path self-host scalability gate opens at the end of Layer B** — `normalize` - becomes polynomial, so the `Optimize.Specialize`-class modules compile. This is the gating fix + becomes polynomial. ~~so the `Optimize.Specialize`-class modules compile.~~ _(Correction + 2026-06-17: polynomial `normalize` is necessary but not sufficient — `Optimize.Specialize` + additionally needed the canonicalKey hashing + `normalFormSizeCap` code-size cap; see the + Correction note up top.)_ This is the gating fix among the self-host blockers (the others — `Impurify` stack-safety, the whole-program memory floor, and lowering's super-linear passes — are independent and tracked separately). -- **Sharing is separable from policy.** Layers A+B are behaviour-neutral scalability fixes - verifiable against the strict byte-equal gate; only Layer C deliberately changes output (and is - guarded by the fusion-converges + collapses-intact criteria of - [ADR 0020](0020-reduction-aware-inliner.md)). +- **Sharing is separable from policy.** Layers A+B are behaviour-neutral scalability fixes; only + Layer C deliberately changes the inline *policy* (and is guarded by the fusion-converges + + collapses-intact criteria of [ADR 0020](0020-reduction-aware-inliner.md)). + - **Correction (2026-06-17):** this originally said A+B are "verifiable against the strict + byte-equal gate". Only **Layer A** turned out byte-identical; **Layer B's CSE necessarily + changes the IR** (it hoists shared values into `let`s — the "more `let`s" the next bullet + predicts), so it is *behaviour*-neutral, not *byte*-neutral. The realized gate is therefore + **test-suite + benchmark + `examples/` no-regression** (the current compiler is not yet a + correct self-host reference, so byte-matching it has limited value), which A+B meet. - **Quote may introduce more `let`s** (shared values that were previously copied). This is the intended contraction; it interacts with lowering's own sharing and must not regress the tuned collapses, which the bench gate checks. @@ -155,16 +199,29 @@ renaming (the [ADR 0032](0032-caller-homed-specialization-for-incremental-builds `output/` build (286 modules from `Main`) reaches and **completes** the `Optimize.Specialize` module in polynomial time, and a synthetic depth-*d* diamond inline DAG normalizes in O(*d*), not O(2^*d*) (a new unit test, the exponential regression guard). + - ✅ **Landed 2026-06-17.** `Data.Lazy` memo in `Semantics.normalize`, byte-IDENTICAL (bench + differential). **Correction to this step's gate:** Layer A *alone* does not make `normalize` + O(*d*) — quote (M2) still re-walks the shared DAG, so the diamond stays exponential until + Layer B. The O(*d*) guard (`Test.NbeStress`) therefore exercises **A + B together**, and the + `Optimize.Specialize` completion likewise needs both. 2. **Layer B** — memoized/sharing `quote` + join points. `normalize` is polynomial; the default self-host path no longer explodes. + - ✅ **Landed 2026-06-17.** `SShared k Sem` + `unShared` (strip at every reduction site) + a + `quote` CSE table hoisting shared values to a top-of-decl `let`. *No `NCase`-continuation join + point was needed* for the gate — value-level CSE alone made `normalize` linear on the diamond + and on `Optimize.Specialize`; the join-point refinement remains available if a future case + surfaces a shared *continuation* the value CSE misses. Gate met without byte-equality (see the + Progress note up top): tests + bench + `examples/metatheory` no regression. 3. **Layer C** — reduction-aware inline/share. Fusion converges and shrinks; the State / dictionary - / comparison / Effect collapses stay intact. + / comparison / Effect collapses stay intact. *(Deferred — the scalability gate is already open + after A + B, so C is now an optimization-quality improvement, not a scalability blocker.)* 4. Demote the round/pass caps to a pure backstop (largely already gone under [ADR 0021](0021-streaming-dependency-ordered-wpo.md)'s single-pass loop). An interim **total-size budget backstop** (stop unfolding once a term exceeds N× its input) can be landed before step 1 if the gate must open immediately; it is a safety net, not the fix, and is -removed once Layer B lands. +removed once Layer B lands. *(Not needed — A + B landed together and opened the gate directly; no +interim backstop was introduced.)* ## Alternatives considered diff --git a/docs/design-decisions/0037-separate-per-module-codegen-and-linking.md b/docs/design-decisions/0037-separate-per-module-codegen-and-linking.md new file mode 100644 index 0000000..7da993d --- /dev/null +++ b/docs/design-decisions/0037-separate-per-module-codegen-and-linking.md @@ -0,0 +1,149 @@ +# 0037. Separate per-module codegen and linking (per-module wasm + `wasm-merge`) + +- Status: Accepted (decision accepted; implementation phased — not yet started) +- Date: 2026-06-17 + +## Context + +Building a large program is slow. Self-compiling the `purs-wasm` CLI (324 reachable +modules → an 8 MB wasm) takes ~45 min. Measured breakdown: + +- module select + decode: ~35 s; +- per-module **optimization** (DictElim/NbE/Specialize/…): ~10–15 min cold, but + **already cached** — the `.pmi`/`.pmo` cache (ADR 0034) serves it in ~25 s on a warm + rebuild; +- **lower + codegen (`buildModule`)**: **~30 min, and it re-runs whole-program on every + build** — never cached; +- Binaryen `-O` + validate + `wasm-merge`: the remainder. + +Profiling the 30-min `buildModule`: ~73 % is inside `binaryen.js` (the Emscripten-compiled +Binaryen, one FFI call per IR node) and ~27 % GC; the PureScript side is spread thin (no +single quadratic — `reachableFunctions` is a clean BFS, `assignProgramReps` a bounded +fixpoint). The cost is the *volume* of whole-program codegen, which the `.pmo` cache does +not touch (it caches optimized MIR, not emitted wasm). + +So the path to fast **incremental** rebuilds (the dev-iteration case; a cold 8 MB build +cannot be ~seconds while we build the module node-by-node through `binaryen.js`) is to make +**codegen per-module and cacheable**: emit each module's wasm once, reuse it when the module +is unchanged, and `wasm-merge` the set — extending the cache from "optimized MIR" to +"emitted wasm". This is the "batch compiler" shape that ADR 0009 (single-wasm output) and +ADR 0021 (which keeps "one Binaryen module, populated incrementally") deliberately avoided, +because for wasm-GC it means **sharing GC struct types across separately-built modules** and +**turning every cross-module call into an import/export boundary** — judged too hard at the +time. + +Three throwaway spikes (hand-written `.wat`, assembled with `wasm-as`, linked with +`wasm-merge --all-features`, run on Node 24) re-tested those fears against Binaryen 123 and +found them solvable: + +1. **GC type sharing (①).** A struct value built in module A is read in module B after merge + — `wasm-merge` **canonicalises structurally-identical types**, including subtyped and + recursive (`$Data` base + `$Cons` subtype) shapes. Constraint: wasm-GC is *isorecursive*, + so identity is per **rec group** — a type grouped with different neighbours stays a + distinct type (a cross-module `ref.cast` then traps). Mitigation: emit each type as its + **own singleton rec group**; singletons canonicalise regardless of a module's other types. + This is directly applicable because our codegen types ADT fields as `i32`/`f64`/`eqref`, + never a specific subtype, so the GC type graph **has no mutual-recursion cycles** — every + type is a base or a base-subtype and can be a singleton. +2. **Cross-module calls + closures (②).** A closure created in module A (a `funcref` to A's + lifted body + a captured env) is applied in module B via the **shared runtime + `applyClo`/`call_ref $Code`** after merge — the `funcref` survives the merge and dispatches + correctly. Cross-module calls become import/export pairs that `wasm-merge` resolves, the + same mechanism already used for the runtime and foreign providers. The closure ABI was + already engineered for this (`$Clo` holds a generic `funcref`, not `(ref $Code)`, keeping + `$Code` out of `$Clo`'s rec group — RuntimeTypes.purs). +3. **Representation ABI (③).** Today `Lower.Unbox.assignProgramReps` is a **whole-program** + fixpoint: a parameter's unboxed rep is the join over *every* call site (sound only if all + callers agree, since unboxing a non-`$Int` `eqref` traps). This was an opportunistic + optimisation enabled by ADR 0009's whole-program availability — an earlier *local* version + (the U2-era `assignReps`) existed and worked; U3 made it whole-program to also unbox + function parameters/results across calls (so a tail loop runs entirely in `i32`). It was + **not** chosen because per-module was impossible. + +## Decision + +Adopt **separate per-module compilation to wasm, linked by `wasm-merge`** (the model Grain +uses: each source file → an object file holding *signature + lowered IR*, then a link/merge +step; Grain likewise codegens via Binaryen). Concretely: + +- **Module-boundary ABI is fixed boxed (`eqref`).** Exported/imported functions take and + return boxed values. The representation (unbox) analysis becomes **module-local**: it keeps + the U3 fixpoint *within* a module (so intra-module tail loops still run unboxed — the main + performance driver) but pins cross-module-visible parameters/results to `Boxed`. We + therefore **lose only the cross-module increment U3 added over U2**, not intra-module + unboxing. A worker/wrapper split can later recover the loop of an *exported* recursive + function if measurement ever warrants it. +- **The lowered ANF stays representation-free at module boundaries** so it is a stable, + cacheable per-module artifact: `.pmi` carries the interface (export signatures, the + deterministic type/label ids), `.pmo` carries the per-module **lowered ANF** (the + codegen input), and codegen emits per-module wasm — `wasm-merge` links them with the + runtime and foreign providers into the single wasm ADR 0009 still mandates as the *output*. +- **GC types are emitted as singleton rec groups**, deterministically from the field-rep + signature, so each module's copy canonicalises under merge. +- **Cross-module calls are import/export pairs**; closures dispatch through the shared runtime. + +This is staged so each step is independently verifiable against the current whole-program +output before the build is actually split: + +- **Phase 0 — groundwork (behaviour-neutral on the current build):** emit data types as + singleton rec groups; assign record-label / intern ids by a **deterministic global scheme** + (so per-module emission agrees without a whole-program pass — barrier ④). +- **Phase 1 — module-local representation:** restrict `assignProgramReps` to the boxed + boundary (behind a flag; A/B against the suite + bench). +- **Phase 2 — per-module codegen:** per-module lower → ANF in `.pmo`; per-module codegen → + per-module wasm; `wasm-merge` link; cross-module CAF-init ordering at link time (barrier ⑤, + partly designed in ADR 0021's link/emit split); whole-program DCE deferred to Binaryen's + merge-time DCE (barrier ⑥). +- **Phase 3 — per-module wasm cache:** reuse unchanged modules' wasm; re-emit only changed + modules; incremental rebuild → ~seconds. + +## Consequences + +- **Incremental rebuilds approach ~seconds** (re-emit the changed modules + merge); a cold + build stays codegen-bound (this does not by itself make the first build fast). +- **Lose cross-module unboxing — measured worst case ~2.6×.** The existing bench corpus + cannot measure this (the `fib`/`sumLoop`/… benches live in one `Bench.Main` module and + cross a boundary only a handful of times per run — their hot loops are intra-module, so the + boxed-boundary simulation never touches them; it only confirmed intra-module unboxing is + preserved). A dedicated microbenchmark (`Bench.Main.crossModule`: a tight intra-module loop + whose every iteration calls `Bench.Helper.step`, a **self-recursive — so never inlined — + O(1)** `Int → Int`) measured the per-crossing cost directly: `crossModule(20M)` ran in + **~34 ms unboxed vs ~87 ms with the boundary pinned to boxed (~2.6×)**, i.e. **~2.6 ns per + crossing** (one `$Int` alloc + unbox). This is the **worst case** — a trivial callee crossed + every iteration, where boxing dominates the ~1.75 ns iteration. The penalty scales as + `~2.6 ns / (callee work + iteration)`, so it shrinks toward negligible as the callee does + real work, and it only applies where a hot scalar cross-module call **survives inlining** + (cheap functions inline away; the per-module model keeps cross-module inlining via summaries + — dict-elim / general inline / caller-homed specialization). The residual hot case (an + *exported, recursive, scalar* function) can be recovered later with a worker/wrapper split. + How often real programs hit the un-recovered case is the open question; for the self-compile + and the bench corpus it appears rare, but that is not yet quantified. +- **Cache coupling stays minimal.** Because the boundary ABI is fixed (no reps in the + interface), a module's `.pmi` does not carry representation decisions, so a dependency's + internal rep change does not invalidate dependents. +- **New required work:** deterministic label/intern ids (④), link-time CAF-init ordering (⑤), + merge-time DCE (⑥). +- **Relationship to prior records.** ADR 0009's *single-wasm output* is preserved — the merge + produces one module; we compile per-module *internally*. This **revisits ADR 0021's + rejected "per-module separate wasm + link" alternative** with the spike evidence above, and + *builds on* ADR 0021's per-module optimisation + summaries and ADR 0032/0034's caches rather + than replacing them. + +## Alternatives considered + +- **Keep whole-program codegen, shrink the constant** (cut the 27 % GC / per-node `Effect` + and array allocation, thin the Binaryen FFI). Helps every build ~1.5–2×, but is bounded and + never reaches ~seconds; orthogonal and can still be done. +- **Parallelise codegen across worker threads** (each its own `binaryen.js`, then merge). + Speeds the *cold* build up to ~#cores×, but does nothing for incremental and adds + per-worker Binaryen memory + coordination. A possible later addition, not the incremental win. +- **Carry representations in `.pmi` to keep cross-module unboxing.** Recovers the U3 + increment, but: parameter reps flow caller→callee while compilation flows callee→caller, so + it needs either a link-time fixpoint over the interface reps (a whole-program step, not + purely incremental) or a heuristic commit (residual boxing where it misses the join); and it + reintroduces rep-driven cache coupling. Rejected for now because the measured boxed-boundary + cost is small; revisit (result reps first — they flow naturally callee→caller) only if a + real workload shows a hot cross-module numeric path regressing. +- **Cache the lowered ANF only, keeping whole-program codegen.** Skips lowering on a hit, but + lowering is not the ~30-min bottleneck (codegen is), so the win is small; the ANF cache is + worthwhile only as the per-module codegen *input* (Phase 2), not on its own. diff --git a/docs/design-decisions/README.md b/docs/design-decisions/README.md index d8e8e8f..a3b5610 100644 --- a/docs/design-decisions/README.md +++ b/docs/design-decisions/README.md @@ -82,8 +82,9 @@ language and are kept out of version control.) | 0032 | [Caller-homed specialization for per-module, incremental builds](0032-caller-homed-specialization-for-incremental-builds.md) | Accepted | | 0033 | [Shipping `ulib` as precompiled MIR (`.pmo`) artifacts](0033-precompiled-ulib-pmo-artifacts.md) | Proposed | | 0034 | [Split the module cache into `.pmi` interface and `.pmo` object](0034-pmi-interface-pmo-object-split.md) | Accepted | -| 0035 | [Sharing/memoizing the NbE reducer, then reduction-aware inlining](0035-sharing-nbe-reduction-aware-inlining.md) | Proposed | +| 0035 | [Sharing/memoizing the NbE reducer, then reduction-aware inlining](0035-sharing-nbe-reduction-aware-inlining.md) | Accepted (Layers A+B + a Layer-C-lite size cap + the Specialize dedup-key fix landed 2026-06-17 — the optimized self-compile completes; full reduction-aware Layer C policy deferred) | | 0036 | [Parameterized join points for decision-tree leaves](0036-join-points-for-decision-tree-leaves.md) | Proposed (de-prioritized — measured duplication ~1.16×, not the `--no-opt` floor) | +| 0037 | [Separate per-module codegen and linking (per-module wasm + `wasm-merge`)](0037-separate-per-module-codegen-and-linking.md) | Accepted (barriers ①②③ validated by spikes; boxed module boundary chosen; implementation phased — not started) | ## Scope @@ -100,10 +101,15 @@ runs on wasm. See [`docs/developers-guide/supported-features.md`](../developers- Current frontiers, tracked by the records above: streaming / incremental codegen (ADR 0021 Phase 2 — reachability pruning and dependency-ordered single-pass optimization shipped; the **`.pmi`/`.pmo` incremental build cache** — default-on, decode-free for unchanged modules — -shipped too, ADR 0032 phase 4 / ADR 0034); a sharing/memoizing NbE reducer and the reduction-aware -inline-or-share selection (ADR 0020's NbE core is implemented; ADR 0035 sequences the sharing fix -that makes the reducer non-exponential — the self-compilation scalability *time* gate — ahead of the -reduction-driven decision, neither of which has landed); the `--no-opt` self-compilation *space* +shipped too, ADR 0032 phase 4 / ADR 0034); the reduction-aware inline-or-share selection (ADR +0020's NbE core is implemented; **ADR 0035 Layers A+B (NbE sharing — makes the reducer +non-exponential) + a Layer-C-lite `normalFormSizeCap` (bounds the `genericShow` code-size blow-up) ++ the `Optimize.Specialize` dedup-key fix (hash, not `show`) landed 2026-06-17, and the optimized +self-compile now *completes*** (writes an 8 MB wasm), where it used to hang at `Optimize.Specialize`; +A+B alone removed the recomputation exponential but did not clear that module — the code-size cap + +dedup-key fix were also required. The full reduction-driven inline *decision* (ADR 0035 Layer C +policy) is deferred — an optimization-quality improvement, not a remaining scalability blocker); the +`--no-opt` self-compilation *space* gate — the front-half whole-program memory floor (decode + translate + lambda-lift holding all MIR at once, ADR 0009) — addressed by **copy-reduction (landed: translate + lambda-lift are fused per module and each module's CoreFn is dropped before the next, so the program is never resident diff --git a/docs/developers-guide/optimizations.md b/docs/developers-guide/optimizations.md index 6d1837d..00067a2 100644 --- a/docs/developers-guide/optimizations.md +++ b/docs/developers-guide/optimizations.md @@ -196,6 +196,14 @@ intrinsic `Data.Eq.eqIntImpl`; `compare` to `ordIntImpl(LT, EQ, GT)`. Two guards it cheap and terminating: a **size cap** (large instances such as the `Generic` `to`/`from` records are not inlined) and **acyclicity**. +A second, *post-reduction* guard catches blow-up the candidate cap cannot predict: if a +declaration's **reduced** form exceeds `DictElim.normalFormSizeCap`, it is re-reduced with the +inline context emptied — the binding stays an ordinary call, the same shape `--no-opt` emits. The +canonical case is the derived `genericShow` dictionary of a large ADT inlined into a `show`, which +produces no redex, only bulk; the cap bounds NbE when a program is itself `show`-heavy (notably the +compiler compiling itself). It measures the *actual* reduced size, so a large-but-shared normal form +(which `quote` CSEs back down — ADR 0035 Layer B) stays inlined. (ADR 0035, "Layer C lite".) + ### General known-function inlining `Optimize/Inline.purs` extends the inline set beyond dictionary plumbing to *ordinary*