This repository contains the simulation code that generates the data
used by the plan_sample_size() function in the
sprtt package on CRAN.
The simulations produce a comprehensive dataset of Sequential
Probability Ratio Test (SPRT) sample size scenarios across various
parameter combinations, enabling sample size recommendations without
real-time computation.
Sample size planning for sequential tests requires extensive Monte Carlo simulations to characterize the sampling behavior under various scenarios of interest. Rather than forcing users to run these computationally intensive simulations every time they need sample size recommendations, we pre-compute a comprehensive dataset covering a wide range of realistic scenarios. This approach offers several advantages:
- CRAN guidelines: The goal is to keeping the main sprtt package lightweight while simulation data remains external and downloadable on-demand.
- Speed: Sample size recommendations are returned instantly by rendering a report based on pre-computed results
- Reproducibility: All users access the same simulation results
- Transparency: Full simulation code is publicly available for inspection and verification
- Computational Efficiency: Eliminates redundant computation across research groups
The sprtt package automatically downloads the simulation data attached to releases in this repository when needed. Users typically don’t need to interact with this repository directly unless they want to:
- Examine the simulation scripts in detail
- Reproduce the simulation results for verification
- Extend the simulations with additional parameter combinations
Automatic Download via sprtt Package
The plan_sample_size() function in sprtt handles data access
automatically. However, you can manually manage the cached data:
# Download the sample size data manually
sprtt::download_sample_size_data()
# Check if data is downloaded and view cache location
sprtt::cache_info()
# Clear the cached data
sprtt::cache_clear()For complete documentation on using the plan_sample_size() function,
visit the sprtt package
website.
The primary output of this repository is
sprtt_external_data_plan_sample_size.rds, attached to the latest
release.
This compressed dataset contains pre-computed SPRT sample size
characteristics for:
- Multiple effect sizes of interest (Cohen’s f)
- Various group sizes
- Range of power and decision-rate levels
File Size and Storage
Due to the comprehensive nature of the simulations, the final dataset (~150 MB) was too large to be included in the sprtt package directly. To maintain CRAN package size limits while providing full functionality:
- The dataset is compressed using R’s
.rdsformat with maximum compression - It is hosted externally via GitHub releases rather than bundled with the package
- The sprtt package includes lightweight download utilities to fetch the data only when needed
- Users cache the data locally after first download, avoiding repeated downloads
Dataset versioning
Each release of this dataset is tagged (e.g., v0.1.0-data) and the
.rds file includes version metadata directly:
loaded <- load_sample_size_data()
loaded$description # short description
loaded$version # e.g. "v0.1.0-data"
loaded$created # date the dataset was created
loaded$n_rep # number of simulation iterations per condition
loaded$data # the dataThis means every HTML report generated by plan_sample_size()
permanently records which dataset version it was based on, including a
direct link to the corresponding GitHub release. Reports generated from
older cached data can always be traced back to the exact simulation that
produced them.
sprtt_plan_sample_size/
├── R/ # Core R functions for simulations
├── cluster/ # Cluster computing scripts
│ └── tool_sprt_sample/ # SLURM job scripts and local R runners
├── analysis/ # Data combination and compression script
├── raw_data/ # Generated raw simulation data
├── data/ # Intermediate processed data
├── meta_data/ # Final combined datasets
└── output/ # Log files and SLURM outputs
Key Directories:
R/: Contains the core simulation functions used to generate raw data and apply sequential testscluster/tool_sprt_sample/: Includes both SLURM bash scripts (.sh) for HPC clusters and R scripts (.R) for local executionanalysis/: Scripts for combining batched simulation results into the final datasetraw_data/,data/,meta_data/: Simulation data at different processing stages
The simulation pipeline consists of three main stages, each building upon the previous to create the final dataset:
Script: tool_sprt_simulate_data.R
Output: Raw simulation datasets saved in raw_data/
Script: tool_sprt_apply.R (or cluster equivalent)
Output: Test results saved in data/
This stage processes the raw simulated data by applying the SPRT. For each simulated dataset:
- Applies the SPRT after each observation
- Calculates the likelihood ratio
- Determines stopping decision (accept H0, accept H1, or continue sampling)
- Records sample sizes at which decisions are reached
Script: analysis/sprt_tool_samples_analyze_batches.R
Output: sprtt_external_data_plan_sample_size.rds in meta_data/
The final stage aggregates results from all simulation batches into a single comprehensive dataset:
- Combines individual batch files into the final dataset
- Applies compression to minimize file size while maintaining fast access
- Validates completeness of the combined dataset
This compressed dataset is what gets attached to repository releases and downloaded by the sprtt package.
For small-scale testing or parameter exploration:
# Navigate to repository root
source("cluster/tool_sprt_sample/run_tool_sprt_simulate_data.R")
source("cluster/tool_sprt_sample/run_tool_sprt_apply.R")For full-scale simulations covering multiple parameter combinations, the repository includes hierarchical SLURM job scripts designed for high-performance computing clusters:
Mother Scripts (orchestrate multiple jobs):
data_mother_jobscript.shLaunches parallel raw data generation jobsapply_mother_jobscript.shLaunches parallel sequential test application jobs
Daughter Scripts (execute individual parameter combinations):
data_daughter_jobscript.shRuns a single raw data generation job for specific parametersapply_daughter_jobscript.shRuns a single test application job for specific parameters and batches of raw data
The hierarchical design works as follows:
- Mother script defines the full parameter grid (effect sizes, group configurations, etc.)
- For each parameter combination, the mother script submits a separate daughter job to the cluster queue
- Daughter jobs run independently in parallel across available compute nodes
- Each daughter job handles a set of parameter combinations and a batch of raw data
- Results are saved to shared storage for later aggregation
The mother scripts also include options for:
- splitting up the raw data into batches to run even more nodes in parallel
- Excluding problematic compute nodes
- Configuring output directories for logs and results
- Managing job submission limits to avoid overwhelming the scheduler
The simulations systematically vary multiple parameters to create a comprehensive lookup table covering realistic research scenarios:
- True Effect Size (
f_simulated): Range of Cohen’s f values from small to large effects - Effect Size of Interest (
f_expected): Range of Cohen’s f values from small to large effects - Sample size (
max_n): Maximum sample size per group - Group configuration (
k_groups): Number of groups determined by standard deviation and sample ratio patternssd: Standard deviation patterns across groups (e.g., “11” for 2 groups, “1111” for 4 groups with equal variances)sample_ratio: Relative sample sizes across groups (e.g., “11” for equal allocation, “111” for balanced three-group designs)
- Distribution: Data generating mechanisms which is in this case a normal distribution
- Replications (
n_rep): 10,000 replications per parameter combination to ensure stable operating characteristics - Alpha and Power: Type I and Type II error rates
This is a research repository primarily for documenting and archiving simulation code. For bug reports or feature requests related to the sprtt package, please visit the main sprtt repository.