Bioinformatics Analyst focused on genomics and single-cell pipeline development. I build reproducible, HPC-scale workflows for NGS data and enjoy going from raw sequencing reads all the way to a biological answer.
- π¬ Currently building and maintaining a Nextflow SHARE-seq pipeline (scRNA + scATAC + sgRNA) for single-cell multiomics analysis
- π§° Comfortable across the full stack: pipeline orchestration (Nextflow), scripting (Python, Bash), and statistical analysis/reporting (R)
- π Interested in genomics, single-cell/multiomics, and building tools that make sequencing data easier to QC and trust
| Category | Tools |
|---|---|
| Languages | Python, R, Bash, SQL |
| Workflow / Job Management | Nextflow, Snakemake, SLURM, LSF, conda, Docker |
| NGS tools | STAR / STARsolo, BWA, GATK4, salmon, samtools, bcftools, cutadapt, Trimmomatic, FastQC, fastp, htseq-count, ArchR, MultiQC, DESeq2, Seurat |
| ML / Stats | scikit-learn, polygenic risk scoring, cross-validation, ROC/AUC, calibration |
| Visualization / Reporting | ggplot2, matplotlib, R Markdown |
- 𧬠SHARE_seq β end-to-end Nextflow pipeline for SHARE-seq multiome data: demultiplexing, barcode QC, STARsolo (RNA) + BWA/ArchR (ATAC) alignment, sgRNA guide assignment, and automated HTML QC reporting
- 𧬠germline-variant-calling-nf β Nextflow DSL2 GATK4 Best Practices pipeline (BWA-MEM β MarkDuplicates β BQSR β HaplotypeCaller β hard-filter) for NA12878, benchmarked against a GIAB truth set (precision/recall/F1 β 0.99) rather than just eyeballing the output VCF
- π rnaseq-diffexp β reproducible Snakemake pipeline for RNA-seq differential expression (
fastpβsalmonβtximport/DESeq2β R Markdown report), with CI that reruns the full pipeline on every push - π§« seurat-scrna-seq-docker β Dockerized Seurat scRNA-seq pipeline (QC β normalization β clustering β UMAP β marker genes) on the 10x pbmc3k dataset, with CI that builds the image and runs the container end to end on every push
- 𧬠colorectal-cancer-prs-ml β polygenic risk score (real PGS Catalog GWAS weights + real 1000 Genomes genotypes) used to demonstrate PRS ancestry-transferability limitations, then combined with ML classifiers (logistic regression, gradient boosting) on a calibrated simulated cohort
- π§ͺ QAA β RNA-seq QC and adapter-trimming benchmark, plus splice-aware alignment and
htseq-countanalysis used to empirically determine library strandedness - π Deduper β memory-efficient, UMI-aware PCR duplicate removal for sorted SAM files
- π Demultiplex β from-scratch FASTQ demultiplexer with index-hopping detection
- π¨ motif-mark β object-oriented tool for visualizing regulatory motifs (incl. degenerate IUPAC codes) on gene models
- π scripting-cookbook β runnable, CI-tested reference examples for common R/Python/Bash/SQL patterns (dplyr/pandas, regex, joins/CTEs/window functions, and more)
π« Connect with me on LinkedIn