Skip to content

Dispatch p/q functions once per vector when the parameters are scalar - #202

Draft
joethorley wants to merge 1 commit into
devfrom
joethorley/multi-quantile-dispatch
Draft

joethorley wants to merge 1 commit into
devfrom
joethorley/multi-quantile-dispatch

Conversation

@joethorley

Copy link
Copy Markdown
Member

Closes #197

Summary

.pdist() and .qdist() dispatched the distribution function once per element of q or p through mapply(). For the model-averaged distribution every one of those scalar calls rebuilt the ssd_emulti() skeleton, normalised the weights, constructed the CDF closure and (since #196) derived the support limits, and the closure returned by pmulti_fun() resolved each component's p*_ssd() by name and re-ran purrr::imap() on every one of the roughly twelve evaluations inside each uniroot() solve.

Two changes, both behaviour preserving:

  • When every distribution parameter is a single value, .pdist()/.qdist() call the distribution function once over the non-missing elements. NA, NaN and invalid-parameter semantics match the elementwise .pd()/.qd() exactly, and the function is never called for missing elements, as before. Parameter vectors, which the bootstrap passes through ssd_qmulti(), still take the elementwise path so they recycle against the probabilities.
  • pmulti_fun() resolves the component functions and separates weights from parameters once, outside the returned closure, which now does only the weighted sum.

Verification

No snapshot changed anywhere in the suite, which is the direct evidence the refactor is behaviour preserving. New tests in test-pqr.R compare the vectorised and elementwise paths for every distribution in ssd_dists_all() over NA, NaN, -Inf, 0, 1, Inf and out-of-range inputs, check that invalid scalar parameters give NaN for every element, and check that parameter vectors still recycle. Full suite 1427 passing, 0 failures; R CMD check 0 errors, 0 warnings, 0 notes; air clean.

Timing on the ccme_boron six-distribution fit (before / after):

Call Before After
ssd_qmulti(1:99 / 100, ...) 0.266 s 0.051 s
ssd_hc(ci = TRUE, nboot = 100, ci_method = "multi_fixed"), four proportions 9.17 s 4.66 s

The remaining bootstrap cost is in refitting, not in quantile dispatch.

`.pdist()` and `.qdist()` called the distribution function once per
element of `q` or `p` through `mapply()`, so `ssd_qmulti(1:99 / 100)`
rebuilt the mixture, normalised the weights and constructed the CDF
closure 99 times, and `pmulti_fun()` resolved each component's function
by name and re-ran `purrr::imap()` on every one of the dozen evaluations
inside each `uniroot()` solve.

When every parameter is a single value, call the distribution function
once over the non-missing elements, reproducing the elementwise NA, NaN
and invalid-parameter semantics exactly. Parameter vectors, as the
bootstrap passes, still dispatch elementwise so that they recycle
against the probabilities. Resolve the component functions and separate
the weights from the parameters once in `pmulti_fun()`, outside the
returned closure.

No snapshot changes. `ssd_qmulti(1:99 / 100)` on the `ccme_boron` fit
drops from 0.27 s to 0.05 s and a four-proportion, 100-bootstrap
`ssd_hc(ci_method = "multi_fixed")` from 9.2 s to 4.7 s. Equivalence
tests compare the vectorised and elementwise paths for every
distribution over missing, boundary and invalid inputs.

Closes #197.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant