Skip to content

SMT: state ground typing facts for refined pure applications (#4591) - #4602

Open
gebner wants to merge 1 commit into
masterfrom
smt-ground-typing
Open

gebner wants to merge 1 commit into
masterfrom
smt-ground-typing

Conversation

@gebner

@gebner gebner commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

Fixes #4591 (alternative to #4597).

Problem

The result refinement of a pure application f a1 .. an reached the solver only through f's typing axiom. To use it, the solver has to instantiate the axiom, prove the argument typings (i.e. f's preconditions), and then unfold the result refinement. Since #4515 stopped restating argument refinements, Z3 often fails to find that chain under nonlinear arithmetic. For example, it may not learn that eval (slice s k n) : nat holds, so sign facts disappear (#4591).

Change

split_goals (FStarC.SMTEncoding.ErrorReporting) now states HasType e t for each application whose result type is refined and that is not under a binder. The fact is placed next to the formula the application occurs in: at an implication's hypothesis, at an if guard, and before a goal. It is valid there because the typechecker established that the application is well typed at that position, much like the typing hypothesis we already emit for a let-bound variable.

  • Guards: the later operands of ==>, \/, /\, &&, || and ite are typechecked assuming the earlier ones, so facts are only taken from unguarded positions. In a hypothesis known to hold, both sides of /\ and && count. Without this, x >= 0 ==> p (g x) with g : x:int{x >= 0} -> r:nat{r <= x} would prove x >= 0. See tests/micro-benchmarks/TypingFactsGuards.fst.
  • Let-bindings: for a hypothesis x == e where x's type already is e's result type, the fact on e is skipped because it duplicates x's typing. The duplicates caused the blowup that Bug4405 guards against.
  • Deduplication: facts are not restated along a path.
  • No fresh names: type lookups don't draw fresh names, so printed names in outputs don't change.
  • Option: --ext typing_facts=off|hastype|refinement, default hastype. refinement asserts the refinement formulas instead of HasType.

Proof and test changes

  • FStar.OrdSet.lemma_as_set_disjoint_left: added eq_lemma (intersect s1 s2) empty. Mentioning the ground term sorted f (intersect s1 s2) makes Z3 give up immediately with "incomplete quantifiers", independent of rlimit. This is pre-existing: it fails the same way on master when the term is written explicitly (Adding a true ground hypothesis makes Z3 give up instantly with 'incomplete quantifiers' (rlimit-independent) #4601).
  • PulseCore.IndirectionTheorySep.read_inv_age: added set_loc__age1 w l; assert on l (later p') (age1_ w).
  • tests/micro-benchmarks/SquashSubtypingDivergence: added UI.one_to_vec_lemma #32 31. The proof was seed-fragile on master too (3–16 rlimit); it now needs 0.15 on every seed.
  • TestErrorLocations: eliminate exists (n:nat). n = 0 with assert (n = 0) now reports one error instead of two. Once the precondition obligation (the existential) is assumed, the assertion follows from the witness's typing.

Testing

The result refinement of a pure application `f a1 .. an` was reachable by
the solver only through f's typing axiom: instantiate it, prove the argument
typings (f's preconditions), then unfold the result refinement. Since #4515
stopped restating argument refinements, that chain is often not found under
nonlinear arithmetic, e.g. `eval (slice s k n) : nat` (#4591).

split_goals now states `HasType e t` for each such application, not under a
binder, at implication hypotheses, `if` guards and before goals. Facts are
only taken from unguarded positions (the later operands of ==>, \/, /\, &&,
|| and ite are typechecked under the earlier ones), and not for `e` in a
hypothesis `x == e` where x already has e's type (let-bindings).
Controlled by `--ext typing_facts=off|hastype|refinement` (default hastype).

Proof fixes: FStar.OrdSet.lemma_as_set_disjoint_left (see #4601),
PulseCore.IndirectionTheorySep.read_inv_age, and a fragile proof in
tests/micro-benchmarks/SquashSubtypingDivergence. TestErrorLocations now
reports one error instead of two.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@gebner

gebner commented Sep 25, 2026

Copy link
Copy Markdown
Contributor Author

!bench

@github-actions

Copy link
Copy Markdown
Contributor

🔬 F* Performance Comparison

Baseline: Base (master @ c0c03b9883e25b19807230d17af06558d1bf0052) | Patched: PR #4602 (smt-ground-typing @ 5fe51bab6355bfccbc97fc649f07db33fadd8cf5) | Matched: 2909 tests

Summary

Metric Median Mean Geo Mean Std Dev P5 → P95 Min → Max
Memory +0.0% +0.0% 1.000× 0.3% -0.2% → +0.5% -2.0% → +4.6%
Time -1.6% -1.6% 0.983× 4.5% -7.5% → +4.3% -82.8% → +56.2%

Totals: Memory: 62.4 MiB | Time: +7.1s

Change Distribution

Metric 🟢 Improved (>5%) ⚪ Unchanged (±5%) 🔴 Regressed (>5%)
Memory 0 (0%) 2909 (100%) 0 (0%)
Time 392 (15%) 2206 (82%) 92 (3%)

Heavy Tests (baseline > 100 MiB, n=776)

Metric Median Mean Geo Mean Std Dev P5 → P95
Memory +0.0% +0.0% 1.000× 0.4% -0.6% → +0.6%
Time -1.4% -1.3% 0.986× 4.1% -6.7% → +4.3%
📉 Top 20 Memory Improvements
File Mem (base) Mem (patch) Mem Δ Time (base) Time (patch) Time Δ
…ples/by-example/_cache/PulseTutorial.Algorithms.fst.checked 345.0 MiB 338.0 MiB -2.0% 6.29s 5.41s -14.0%
examples/metatheory/_cache/MiniValeSemantics.fst.checked 991.2 MiB 984.9 MiB -0.6% 20.26s 19.45s -4.0%
…/examples/by-example/_cache/PulseTutorial.Loops.fst.checked 342.0 MiB 338.0 MiB -1.2% 5.01s 4.61s -7.9%
…lse/share/pulse/examples/dice/_cache/CBOR.Pulse.fst.checked 783.0 MiB 779.0 MiB -0.5% 31.70s 30.81s -2.8%
tests/tactics/_cache/Prettify.fst.checked 547.0 MiB 544.0 MiB -0.6% 16.35s 15.91s -2.7%
tests/custard/pulse/_cache/CborBoundarySlice.fst.checked 689.0 MiB 686.0 MiB -0.4% 36.32s 34.31s -5.5%
pulse/share/pulse/examples/_cache/CustomSyntax.fst.checked 289.0 MiB 287.0 MiB -0.7% 3.73s 3.51s -5.9%
pulse/share/pulse/examples/dice/_cache/DPE.fst.checked 678.0 MiB 676.0 MiB -0.3% 21.86s 23.07s +5.5%
pulse/test/_cache/ExtractUninit.fst.checked 212.0 MiB 210.0 MiB -0.9% 1.20s 1.03s -14.6%
pulse/test/_output/PrintCheck.fst.output 154.0 MiB 152.0 MiB -1.3% 0.56s 0.57s +1.4%
pulse/test/bug-reports/_cache/DependentTuples.fst.checked 155.0 MiB 153.0 MiB -1.3% 0.59s 0.58s -2.4%
pulse/test/bug-reports/_cache/GhostAdmit.fst.checked 151.0 MiB 149.0 MiB -1.3% 0.52s 0.53s +1.0%
pulse/test/bug-reports/_output/Bug166.krml 221.0 MiB 219.0 MiB -0.9% 1.18s 1.16s -1.5%
pulse/test/error_messages/_cache/IntroExistsFail.fst.checked 180.0 MiB 178.0 MiB -1.1% 0.61s 0.60s -1.5%
pulse/test/error_messages/_cache/LetBindingType.fst.checked 151.0 MiB 149.0 MiB -1.3% 0.52s 0.49s -4.8%
…lse/test/error_messages/_output/InvariantPayload.fst.output 180.0 MiB 178.0 MiB -1.1% 0.61s 0.62s +2.0%
pulse/test/nolib/_cache/Annots.fst.checked 139.0 MiB 137.0 MiB -1.4% 0.48s 0.47s -2.3%
tests/custard/pulse/_cache/PulseGlobalArrayEmpty.fst.checked 151.0 MiB 149.0 MiB -1.3% 0.48s 0.49s +3.3%
tests/extraction/backends/_cache/ExtUIntMask.fst.checked 96.7 MiB 95.5 MiB -1.3% 56.07s 71.56s +27.6%
doc/book/code/_cache/Vec.fst.checked 112.2 MiB 111.0 MiB -1.1% 1.12s 1.08s -3.5%
📈 Top 20 Memory Regressions
File Mem (base) Mem (patch) Mem Δ Time (base) Time (patch) Time Δ
examples/metatheory/_cache/LambdaOmega.fst.checked 252.5 MiB 264.2 MiB +4.6% 7.52s 7.42s -1.3%
pulse/share/pulse/examples/_cache/Quicksort.Base.fst.checked 355.0 MiB 363.0 MiB +2.3% 8.09s 8.43s +4.2%
tests/hacl/_cache/Lib.Sequence.fsti.checked 195.1 MiB 199.5 MiB +2.2% 4.15s 4.44s +7.1%
…ples/dsls/bool_refinement/_cache/BoolRefinement.fst.checked 482.7 MiB 486.6 MiB +0.8% 75.39s 74.82s -0.8%
tests/custard/pulse/_output/PulseComment.dc 283.0 MiB 286.0 MiB +1.1% 1.36s 1.38s +1.4%
tests/custard/pulse/_cache/PulseSlice.fst.checked 235.0 MiB 238.0 MiB +1.3% 2.12s 2.35s +11.0%
…pulse/examples/_cache/PulseExample.BinarySearch.fst.checked 233.0 MiB 236.0 MiB +1.3% 2.38s 2.26s -5.1%
…e/pulse/examples/_cache/Example.ImplicitBinders.fst.checked 151.0 MiB 154.0 MiB +2.0% 0.57s 0.62s +8.3%
tests/error-messages/_output/Bug2899.fst.output 90.5 MiB 92.9 MiB +2.7% 0.40s 0.41s +2.5%
tests/error-messages/_output/Bug2899.fst.json_output 90.5 MiB 92.9 MiB +2.7% 0.39s 0.41s +2.8%
tests/error-messages/_cache/Bug2899.fst.checked 90.5 MiB 92.9 MiB +2.7% 0.40s 0.40s -0.2%
tests/custard/pulse/_output/ArrTup.dc 217.0 MiB 219.0 MiB +0.9% 1.15s 1.17s +1.6%
tests/custard/pulse/_cache/PulseBasic.fst.checked 193.0 MiB 195.0 MiB +1.0% 0.95s 0.94s -0.8%
pulse/test/nolib/_output/Test.Matcher.fst.output 166.0 MiB 168.0 MiB +1.2% 0.66s 0.63s -4.1%
…se/test/error_messages/_output/LeftoverResources.fst.output 150.0 MiB 152.0 MiB +1.3% 0.52s 0.53s +1.2%
…se/test/error_messages/_output/IllTypedInvariant.fst.output 183.0 MiB 185.0 MiB +1.1% 0.67s 0.66s -0.5%
pulse/test/bug-reports/_cache/RecordOfArrays.fst.checked 199.0 MiB 201.0 MiB +1.0% 0.99s 0.99s +0.2%
…st/bug-reports/_cache/BugUnificationUnderBinder.fst.checked 180.0 MiB 182.0 MiB +1.1% 0.65s 0.65s -0.2%
pulse/test/_cache/UnfoldMetaArg.fst.checked 150.0 MiB 152.0 MiB +1.3% 0.52s 0.52s +0.6%
pulse/test/_cache/Goto.fst.checked 233.0 MiB 235.0 MiB +0.9% 1.98s 1.87s -5.4%
⬇️ Top 20 Time Improvements
File Mem (base) Mem (patch) Mem Δ Time (base) Time (patch) Time Δ
tests/extraction/backends/_cache/ExtUIntUnsigned.fst.checked 112.2 MiB 112.3 MiB +0.1% 111.40s 97.57s -12.4%
…sts/extraction/backends/_cache/ExtIntShiftArith.fst.checked 78.2 MiB 78.6 MiB +0.5% 37.89s 34.77s -8.2%
tests/extraction/backends/_cache/ExtIntSigned.fst.checked 86.4 MiB 87.8 MiB +1.6% 12.80s 10.53s -17.8%
tests/extraction/cmi/3/_output/B.ml 3.6 GiB 3.6 GiB +0.0% 36.72s 34.61s -5.7%
tests/custard/pulse/_cache/CborBoundarySlice.fst.checked 689.0 MiB 686.0 MiB -0.4% 36.32s 34.31s -5.5%
tests/extraction/backends/_cache/ExtIntNe.fst.checked 76.6 MiB 76.4 MiB -0.4% 2.08s 0.36s -82.8%
tests/extraction/backends/_cache/ExtUInt8Lognot.fst.checked 78.9 MiB 78.8 MiB -0.1% 20.01s 18.31s -8.5%
…cro-benchmarks/_cache/SquashSubtypingDivergence.fst.checked 45.6 MiB 45.9 MiB +0.7% 4.80s 3.17s -33.9%
tests/extraction/backends/_cache/ExtIntDivRem.fst.checked 76.9 MiB 78.7 MiB +2.3% 15.10s 13.72s -9.1%
pulse/share/pulse/examples/_cache/GhostBag.fst.checked 266.0 MiB 267.0 MiB +0.4% 6.34s 5.13s -19.1%
tests/calc/_cache/Long.fst.checked 1.1 GiB 1.1 GiB +0.0% 35.82s 34.70s -3.1%
tests/extraction/_cache/Division.fst.checked 75.3 MiB 76.6 MiB +1.7% 3.08s 2.18s -29.4%
…lse/share/pulse/examples/dice/_cache/CBOR.Pulse.fst.checked 783.0 MiB 779.0 MiB -0.5% 31.70s 30.81s -2.8%
…ples/by-example/_cache/PulseTutorial.Algorithms.fst.checked 345.0 MiB 338.0 MiB -2.0% 6.29s 5.41s -14.0%
tests/extraction/backends/_cache/ExtUIntRotate.fst.checked 81.5 MiB 81.6 MiB +0.1% 24.49s 23.64s -3.5%
examples/metatheory/_cache/MiniValeSemantics.fst.checked 991.2 MiB 984.9 MiB -0.6% 20.26s 19.45s -4.0%
tests/semiring/_cache/CanonCommSemiring.fst.checked 417.8 MiB 418.0 MiB +0.0% 12.74s 11.97s -6.1%
tests/vale/_cache/X64.Poly1305.Math_i.fst.checked 155.4 MiB 155.5 MiB +0.1% 3.60s 2.87s -20.1%
tests/micro-benchmarks/_output/Strict.ml 196.7 MiB 196.8 MiB +0.1% 15.12s 14.48s -4.2%
tests/micro-benchmarks/_cache/Test.IrrationalPow.fst.checked 42.4 MiB 42.5 MiB +0.2% 1.52s 0.91s -40.0%
⬆️ Top 20 Time Regressions
File Mem (base) Mem (patch) Mem Δ Time (base) Time (patch) Time Δ
tests/custard/_output/NormBudget.rejected 4.8 GiB 4.8 GiB +0.0% 238.73s 267.97s +12.2%
tests/vale/_cache/X64.Vale.Decls.fst.checked 273.8 MiB 275.1 MiB +0.5% 34.06s 53.20s +56.2%
tests/extraction/backends/_cache/ExtUIntMask.fst.checked 96.7 MiB 95.5 MiB -1.3% 56.07s 71.56s +27.6%
…re/pulse/examples/dice/_cache/DPE.Messages.Spec.fst.checked 323.0 MiB 324.0 MiB +0.3% 85.99s 88.74s +3.2%
pulse/share/pulse/examples/dice/_cache/DPE.fst.checked 678.0 MiB 676.0 MiB -0.3% 21.86s 23.07s +5.5%
pulse/share/pulse/examples/dice/_cache/DPE_CBOR.fst.checked 392.0 MiB 392.0 MiB +0.0% 13.93s 15.05s +8.0%
examples/algorithms/_cache/StringMatching.fst.checked 135.5 MiB 136.2 MiB +0.5% 6.44s 7.46s +15.9%
…e/pulse/examples/dice/_cache/DPE.Messages.Parse.fst.checked 406.0 MiB 405.0 MiB -0.2% 37.15s 37.73s +1.6%
tests/vale/_cache/X64.Poly1305.Bitvectors_i.fst.checked 137.8 MiB 139.6 MiB +1.3% 5.37s 5.87s +9.5%
examples/data_structures/_cache/BinomialQueue.fst.checked 117.1 MiB 117.9 MiB +0.7% 5.85s 6.30s +7.7%
…/share/pulse/examples/dice/_cache/CBOR.Spec.Map.fst.checked 506.0 MiB 506.0 MiB +0.0% 8.60s 8.99s +4.5%
pulse/share/pulse/examples/_cache/MSort.Base.fst.checked 402.0 MiB 402.0 MiB +0.0% 7.59s 7.96s +4.9%
…ples/parallel/_cache/Example.RingBufferTransfer.fst.checked 367.0 MiB 366.0 MiB -0.3% 5.18s 5.53s +6.7%
pulse/share/pulse/examples/_cache/Quicksort.Base.fst.checked 355.0 MiB 363.0 MiB +2.3% 8.09s 8.43s +4.2%
tests/hacl/_cache/Lib.Sequence.fsti.checked 195.1 MiB 199.5 MiB +2.2% 4.15s 4.44s +7.1%
…reports/closed/_cache/SquashSubtypingDivergence.fst.checked 67.4 MiB 67.5 MiB +0.1% 7.33s 7.60s +3.6%
…lse/examples/parallel/_cache/Promises.Examples3.fst.checked 253.0 MiB 252.0 MiB -0.4% 2.26s 2.50s +10.6%
tests/custard/pulse/_cache/PulseSlice.fst.checked 235.0 MiB 238.0 MiB +1.3% 2.12s 2.35s +11.0%
tests/bug-reports/closed/_cache/Bug312.fst.checked 49.1 MiB 49.0 MiB -0.2% 0.92s 1.07s +16.4%
…share/pulse/examples/dice/_cache/DPE.TestClient.fst.checked 276.0 MiB 276.0 MiB +0.0% 2.10s 2.25s +7.0%

ℹ️ The full comparison table (2909 tests) was omitted to keep this report within GitHub's comment size limit. See the attached HTML report for all tests.

Note: Memory values report the peak OCaml heap of the F* process (excluding Z3 subprocesses). Time measurements are from single runs and may be noisy, especially for fast tests. Tests with baseline time < 0.1s are excluded from time statistics. The geometric mean (Geo Mean) of the patched/baseline ratio is the most robust summary statistic for performance comparisons (1.0× = no change, <1× = improvement).


📎 Download HTML report and raw .ramon files

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Large SMT regression in nightly-2026-09-23: an unused Lemma in scope makes a later proof 20x more expensive

1 participant