SPADE
We developed SPADE, an R-based command-line tool for allele frequency quality control and preliminary detection of sex-differentiated genetic signals. SPADE is designed to assess sex-specific allele frequency differences while accounting for disease prevalence and study design.
The tool accepts PLINK-formatted genotype count files as input and supports both pseudo-autosomal (PAR) and non-PAR variants, automatically applying appropriate statistical procedures for each genomic context. SPADE performs input validation, filtering, statistical computation, and comprehensive logging to ensure reproducibility.
Through integration with the argparse R package, SPADE provides a user-friendly command-line interface, allowing users to specify parameters such as case/control sample sizes, population prevalence, and output configurations. The software supports standard command-line options (e.g., --help) and records all runtime arguments.
To facilitate genome-wide analyses, SPADE can be combined with lightweight bash scripting to enable chromosome-level parallelization, allowing efficient large-scale processing.
Methodological details are available in our preprint:
Controls-Only Quality Control Metric from Early GWAS Can Attenuate Gene-Sex Interaction Signals in Contemporary Large-Scale Studies
https://doi.org/10.64898/2026.07.22.26358697
If you use SPADE in your research, please cite:
@article{chen2026spade,
author = {Chen, D. Z. and others},
title = {Controls-Only Quality Control Metric from Early GWAS Can Attenuate Gene-Sex Interaction Signals in Contemporary Large-Scale Studies},
journal = {medRxiv},
year = {2026},
doi = {10.64898/2026.07.22.26358697}
}
- R (>= 3.5)
- argparse
Install the required R package:
install.packages("argparse")SPADE accepts PLINK-formatted genotype count files (.gcount) as input. For each analysis, four input files are required:
- Female cases (
.gcount) - Male cases (
.gcount) - Female controls (
.gcount) - Male controls (
.gcount)
Additionally, a PLINK .bim file is required for variant annotation.
Rscript SPADE.r \
-fd female_cases.gcount \
-md male_cases.gcount \
-fh female_controls.gcount \
-mh male_controls.gcount \
-pf 0.04 \
-pm 0.08 \
-o output_prefix \
-b input.bim \
--mac 5| Option | Description |
|---|---|
-fd |
Female cases genotype count file (PLINK .gcount format) |
-md |
Male cases genotype count file |
-fh |
Female controls genotype count file |
-mh |
Male controls genotype count file |
-pf |
Disease prevalence in females (e.g., 0.04 for 4%) |
-pm |
Disease prevalence in males |
-o |
Output prefix for result files |
-b |
PLINK .bim file for variant annotation |
--mac |
Minimum allele count threshold for filtering |
--no-indels |
Indels filter |
--help |
Show help message |
Below is an example of the SPADE output format (first two rows shown):
| CHROM | BP | ID | A1 | A2 | hF_A1A1 | hF_A1A2 | hF_A2A2 | hM_A1A1.A1 | hM_A1A2 | hM_A2A2.A2 | dF_A1A1.A1 | dF_A1A2 | dF_A2A2.A2 | dM_A1A1.A1 | dM_A1A2 | dM_A2A2.A2 | hF_MISSING_CT | dF_MISSING_CT | hM_MISSING_CT | dM_MISSING_CT | EAF_FemaleControls | EAF_MaleControls | EAF_FemaleCases | EAF_MaleCases | EAFsd_FemaleControls | EAFsd_MaleControls | EAFsd_FemaleCases | EAFsd_MaleCases | N_FemaleControls | N_MaleControls | N_FemaleCases | N_MaleCases | EAFDiff_CaseOnly_Sex | EAFDiff_ControlOnly_Sex | EAFDiff_EstPopulation_Sex | EAFDiffsd_EstPopulation_Sex | log10P_EstPopulation_Sex | log10P_CaseOnly_Sex | log10P_ControlOnly_Sex | log10P_FisherCombinedCaseControl |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| X | 1 | rs1 | A | G | 18168 | 69943 | 11889 | 50741 | NA | 49259 | 24422 | 67189 | 8389 | 82292 | NA | 17708 | 0 | 0 | 0 | 0 | 0.468605 | 0.49259 | 0.419835 | 0.17708 | 0.000861143739308369 | 0.0015809651858912 | 0.000869488773791818 | 0.00120715646707459 | 1e+05 | 1e+05 | 1e+05 | 1e+05 | 0.242755 | -0.023985 | 0.000695000000000001 | 0.00167615652382169 | 0.16851023353268 | 5784.1054780858 | 39.7686584656196 | Inf |
| 2 | 2 | rs2 | A | G | 69342 | 28490 | 2168 | 68725 | 29061 | 2214 | 61334 | 34986 | 3680 | 69741 | 28182 | 2077 | 0 | 0 | 0 | 0 | 0.16413 | 0.167445 | 0.21173 | 0.16168 | 0.000812196670148308 | 0.000817035323440792 | 0.00089126543240496 | 0.00080675013232103 | 1e+05 | 1e+05 | 1e+05 | 1e+05 | 0.05005 | -0.00331500000000001 | 0.000949800000000001 | 0.0010855384567513 | 0.418395005039165 | 378.104713339255 | 2.3970140241462 | Inf |
SPADE outputs sex-prevalence-adjusted allele frequency estimates, variance measures, and multiple test statistics for case-only, control-only, and population-level inference.
| Column | Description |
|---|---|
| CHROM | Chromosome |
| BP | Base pair position |
| ID | SNP ID |
| A1 | Reference allele |
| A2 | Alternate allele |
| hF_A1A1, hF_A1A2, hF_A2A2 | Female control genotypes |
| hM_A1A1.A1, hM_A1A2, hM_A2A2.A2 | Male control genotypes |
| dF_A1A1, dF_A1A2, dF_A2A2 | Female case genotypes |
| dM_A1A1.A1, dM_A1A2, dM_A2A2.A2 | Male case genotypes |
| hF_MISSING_CT, hM_MISSING_CT, dF_MISSING_CT, dM_MISSING_CT | Missing counts per group |
| EAF_FemaleControls, EAF_MaleControls, EAF_FemaleCases, EAF_MaleCases | Estimated allele frequencies |
| EAFsd_FemaleControls, EAFsd_MaleControls, EAFsd_FemaleCases, EAFsd_MaleCases | SD of allele frequencies |
| N_FemaleControls, N_MaleControls, N_FemaleCases, N_MaleCases | Sample sizes |
| EAFDiff_CaseOnly_Sex | Case-only allele frequency difference |
| EAFDiff_ControlOnly_Sex | Control-only allele frequency difference |
| EAFDiff_EstPopulation_Sex | Sex-prevalence adjusted allele frequency difference |
| EAFDiffsd_EstPopulation_Sex | SD of adjusted difference |
| log10P_EstPopulation_Sex | -log10 p-value of adjusted test |
| log10P_CaseOnly_Sex | -log10 p-value for case-only comparison |
| log10P_ControlOnly_Sex | -log10 p-value for control-only comparison |
| log10P_FisherCombinedCaseControl | Fisher combined p-value |
SPADE automatically generates a log file with the suffix .log containing:
- Runtime arguments and parameter settings
- Input file paths and validation results
- Filtering statistics (e.g., variants removed by MAC threshold)
- Warning and error messages
- Processing timestamps This log ensures reproducibility and aids in debugging.
- Supports both pseudo-autosomal (PAR) and non-PAR variants with appropriate statistical procedures
- Performs input validation, filtering, and comprehensive logging
- Records all runtime arguments for reproducibility
- Can be combined with lightweight bash scripting for chromosome-level parallelization
for chr in {1..22}; do
Rscript SPADE.r \
-fd "chr${chr}_female_cases.gcount" \
-md "chr${chr}_male_cases.gcount" \
-fh "chr${chr}_female_controls.gcount" \
-mh "chr${chr}_male_controls.gcount" \
-pf 0.04 -pm 0.08 \
-o "output_chr${chr}" \
-b "chr${chr}.bim" \
--mac 500 &
done
waitRscript SPADE.r -fd toy-female_cases.gcount -md toy-male_cases.gcount -fh toy-female_controls.gcount -mh toy-male_controls.gcount -pf 0.04 -pm 0.08 -o toy_output -b toy.bim --mac 500A lightweight companion tool to assess differential missingness rates between sexes in case and control groups.
Rscript SPADEmissing.R input.txt output.txtSPADEmissing.R takes a single result file from SPADE (formatted as shown in the Sample Output section) and computes:
- OR_Missing_Case_FemaleVsMale: Odds ratio of missingness in female cases vs male cases
- log10P_Missing_Case_FemaleVsMale: -log10(p-value) for case group sex difference in missingness
- OR_Missing_Control_FemaleVsMale: Odds ratio of missingness in female controls vs male controls
- log10P_Missing_Control_FemaleVsMale: -log10(p-value) for control group sex difference in missingness These four columns are appended to the original input file.
- Missingness is defined as the proportion of samples failing genotype calling at a given variant
- Odds ratio > 1 indicates higher missingness in females; < 1 indicates higher missingness in males
- Fisher's exact test is used for p-value calculation
- For contingency tables with zero counts, a continuity correction of 0.5 is applied
