Multi-Platform Aggregation and Quantification of Transcripts
MPAQT is an R package for RNA-seq transcript quantification that integrates short-read and long-read sequencing data for improved accuracy. It supports both bulk and single-cell RNA-seq analysis.
Documentation: csglab.github.io/MPAQT
- Multi-platform integration: Combine Illumina short-reads with PacBio/ONT long-reads
- Bulk and single-cell: Unified API for both analysis types
- Positional bias correction: Account for 3' or 5' sequencing biases
- Prior integration: Incorporate custom transcript-specific priors into abundance estimation
- Flexible inputs: Start from FASTQ or pre-computed counts
Choose your preferred installation method:
| Method | Best For | Instructions |
|---|---|---|
| Local | Development, customization | Jump to section |
| Conda | Simple environment management | Jump to section |
| Apptainer | HPC clusters | Jump to section |
Install R (>= 4.0.0) from CRAN with compilation tools (make, zlib, curl).
Install the Bioconductor packages required by mpaqt_index():
if (!requireNamespace("BiocManager", quietly = TRUE))
install.packages("BiocManager")
BiocManager::install(c("Biostrings", "rtracklayer"))install.packages("pak")
pak::pak("csglab/MPAQT")The source repository is public. GitHub credentials are optional and can help avoid API rate limits.
Install kallisto (>= 0.50.1) and bustools (>= 0.43.1).
Verify:
kallisto version
bustools versionconda create -n mpaqt \
-c csglab -c conda-forge -c bioconda -c defaults \
r-mpaqt
conda activate mpaqt
R -e 'library(mpaqt)'The Conda package includes Biostrings, rtracklayer, kallisto, and
bustools, which are required to create an index with mpaqt_index().
The defaults channel supplies the r-gpboost dependency.
For HPC clusters:
# Pull image
apptainer pull mpaqt_2.4.0.sif \
library://csglab/mpaqt/mpaqt:2.4.0
# R API usage
apptainer exec mpaqt_2.4.0.sif R -e 'library(mpaqt)'
# Run interactively
apptainer shell mpaqt_2.4.0.siflibrary(mpaqt)
# 1. Create index (run once)
index <- mpaqt_index(
annotation = "gencode.v44.gtf",
transcriptome = "gencode.v44.transcripts.fa",
output_file = "mpaqt.index.rds"
)
# 2. Process short reads
sr_counts <- mpaqt_prepare_short_reads(
index = index,
fastq_1 = "sample_R1.fastq.gz",
fastq_2 = "sample_R2.fastq.gz",
output_dir = "results/"
)
# 3. Quantify
result <- mpaqt_quant(
index = index,
sr_counts = sr_counts,
positional_bias = "3p"
)
# 4. Get results
tpm_values <- tpm(result)Full documentation: https://csglab.github.io/MPAQT/
| Guide | Description |
|---|---|
| Installation | Detailed installation guide |
| Bulk Workflow | Complete bulk RNA-seq analysis |
| Single-Cell | Cluster-level quantification |
| API Reference | All functions |
If you use MPAQT in your research, please cite:
Apostolides, M., Choi, B., Navickas, A., Saberi, A., Soto, L. M., Goodarzi, H., & Najafabadi, H. S. (2024). Accurate isoform quantification by joint short- and long-read RNA sequencing. BioRxiv. https://doi.org/10.1101/2024.07.11.603067
We welcome contributions! See our Package Structure guide.
Report issues at: https://github.com/csglab/MPAQT/issues.
MIT License
