Skip to content
View Junaid-Ahmed-Rupok's full-sized avatar

Block or report Junaid-Ahmed-Rupok

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Junaid-Ahmed-Rupok/README.md

Hi, I'm Junaid


🏆 Two "Best Paper" trophies in twelve months. 📊 Three research groups run their analysis on tools I built. 🎙️ Once had 10,000 people tuning in to hear me talk on the radio. Now I make machines talk sense instead.


⚡ The 10-Second Version

🎯 Doing Making ML statistically honest — causal inference, fairness, uncertainty, no p-hacking
🏆 Won 1st Best Paper @ SPECTRA 2026 · Best Paper @ APMEE 2025
🛠️ Built StatsPro · ReproHub · a citation-obsessed RAG chatbot
📝 Publishing 1 journal article, 1 invited book chapter, 7 conference papers, 2 more in review
🎓 Next Hunting for a PhD in Computational Science / ML / Governance Analytics

🧠 What I Actually Do

Most ML pipelines quietly skip the statistics part. I don't. I build systems where every claim — a model's accuracy, a policy's effect, a paper's result — has to survive a hypothesis test before I believe it.

That means: causal inference instead of correlation-shrugging, uncertainty quantification instead of point-estimate overconfidence, and reproducibility checks instead of "trust me, it ran on my machine."

Say it in one line: statistically rigorous, causally grounded, reproducible ML — for problems big enough to matter (governance, fairness, public policy).


🛠️ Stuff I Built (that people actually use)

Project Built with Why it's cool
♻️ ReproHub  [live] Python · SciPy Re-runs a paper's stats on the raw data and actually checks if the claims hold — not just "is p < 0.05"
📊 StatsPro  [live] Streamlit · scikit-learn CSV in, full report out — 10 tests, 6 ML models, AI-narrated insights. 3 research groups run on it
🤖 Smart RAG Chatbot  [live] LangChain · FAISS Refuses to answer without citing the exact passage. No hallucinated sources allowed
🤗 IMDb Transformer PyTorch · DistilBERT 87% accuracy, loss dropped 0.44 → 0.16 in 3 epochs
🌍 Economic Classification XGBoost 98.3% accuracy sorting 150+ countries by World Bank data
🔍 Crime Analytics XGBoost · SMOTE 93% accuracy across 6 crime types, spatial-temporal features
🏭 Enterprise Survey EDA Pandas · SciPy ANOVA across 55,620 records, 117 industries (p < 0.001, for real)

📚 Papers (the receipts)

🏆 Award-winning work
  • Ahmed, S.J., Islam Nahian, M.T., & Kwoshik, M.H.R. (2026). "CF-EGAT: A Causal Fairness-Aware Equity Graph Attention Network for Country-Level Environmental Livability Classification." Symposium on Photonics, Emerging Computational Technologies, Research & AI-Data Science (SPECTRA 2026). Oral Presentation. 🏆 1st Best Paper Award. DOI: 10.5281/zenodo.21195761
  • Ahmed, S.J. (2025). "Multi-Dimensional Statistical Similarity for Governance Classification: Beyond Arbitrary Thresholds in Comparative Politics." 6th Annual Paper Meet Electrical Engineering Division (APMEE 2025). Oral Presentation. 🏆 Best Research Paper Award.
📄 Journal & conference publications
  • Ahmed, S.J., Kwoshik, M.H.R., & Islam Nahian, M.T. (2026). "Machine Learning for Crime Classification: A Fairness-Aware Approach to Class Imbalance." Journal of Machine Learning and Applications, 2(1), 9–17. DOI: 10.61577/jmla.2026.100002
  • Ahmed, S.J., Islam Nahian, M.T., & Kwoshik, M.H.R. (2026). "RMA-BO: Regret-Minimizing Adaptive Bayesian Optimization." SPECTRA 2026. Oral Presentation. DOI: 10.5281/zenodo.21194394
  • Ahmed, S.J., Islam Nahian, M.T., & Kwoshik, M.H.R. (2026). "Environmental Livability Assessment via Adaptive Bootstrap-Retrained SHAP and Statistically-Constrained Pareto Counterfactuals: A Cross-National Analysis." 5th IEEE International Conference on Signal Processing, Information, Communication and Systems (SPICSCON 2026). BAUET, Bangladesh. August 13–14, 2026. Accepted for Presentation. IEEE Xplore.
  • Ahmed, S.J., Kwoshik, M.H.R., & Islam Nahian, M.T. (2026). "Machine Learning for Crime Classification: A Fairness-Aware Approach to Class Imbalance." SPICSCON 2026. BAUET, Bangladesh. August 13–14, 2026. Accepted for Presentation. IEEE Xplore. (conference presentation of the journal article above)
  • Ahmed, S.J. (2026). "DeepEnMap: Ordinal-Aware Multi-Modal Deep Learning for Energy Poverty Risk Mapping." IEMIS 2026 — 4th International Conference on Emerging Technologies in Data Mining and Information Security. UBC, Vancouver, Canada. August 10–12, 2026. Accepted for Presentation. Springer LNNS Series (Scopus, EI-Compendex, DBLP, ISI Proceedings).
  • Ahmed, S.J. (2026). "Density-Decoupled, Mask-Ablated Segmentation-Guided Diffusion for Controllable Mammography Synthesis: A Preliminary Study." IEMIS 2026 — 4th International Conference on Emerging Technologies in Data Mining and Information Security. UBC, Vancouver, Canada. August 10–12, 2026. Accepted for Presentation. Springer LNNS Series (Scopus, EI-Compendex, DBLP, ISI Proceedings).
📖 Book chapter (invited)
  • Ahmed, S.J. (2026). "Generative AI and Mathematical Optimization for Football Match Outcome Prediction: A Comparative Study of CatBoost, XGBoost, and TabNet with Kelly Index Stratification." In: Lahby, M. (ed.), Generative AI and Mathematical Optimization for Performance and Innovation in Football, Springer Optimization and Its Applications. Springer, Cham. (Invited Chapter, Under Review)
🔬 Manuscripts under review
  • Ahmed, S.J. (2026). "DemocracyGuard: Testing a Divergence-Index Reconciliation of Subjective and Objective Democracy Indicators for Forecasting Adverse Regime Transitions." Under review at Transactions on Machine Learning Research (TMLR). (Q1, Top-Tier Journal)
  • Ahmed, S.J. (2026). "FAI: Feature-Wise Adaptive Imputation via Downstream-Aware Method Selection." Under review at ICISET 2026 (IEEE Xplore).

🐍 Live from GitHub

A snake animation eating through my contribution graph

Updates daily · powered by a GitHub Action, not a static image

🧰 Tech I Reach For

Python PyTorch TensorFlow XGBoost HuggingFace LangChain scikit-learn Streamlit Docker Git

The math behind it: Hypothesis testing · Bayesian & causal inference · Bootstrap · FDR correction · Uncertainty quantification · Measure-theoretic probability


🔬 Day Job vs. Side Quests

Researcher, Royal Scientific Publications (2026–present) — chasing open problems in computational science and applied ML.

Senior Researcher, Young Learners' Research Lab, RUET (2022–2024) — ran the lab's hypothesis-testing pipeline, mentored 5+ junior researchers (one shipped their first conference paper because of it), and made "reproducible" the lab's default, not the exception.

On the side: tutored 50+ underprivileged students for free, mentored 4 into RUET/BUET/CUET, guided 5 to perfect board-exam GPAs — because good statistics should help real people, not just impress reviewers.


🎙️ Before the Math

Before the p-values, there was a radio mic. I hosted a weekly show during COVID that reached ~10,000 listeners, published a Bengali poetry collection ("Sob Odvuture" — 50+ poems, 200+ copies sold), and wrote/performed 15+ original songs. Turns out storytelling and statistical rigor want the same thing: something worth trusting.


🎓 B.Sc. CSE, RUET · HSC/SSC: perfect 5.00/5.00 · IELTS 7.0

🎯 Currently chasing: a PhD that lets me do this at scale


Pinned Loading

  1. statistical-analysis-app statistical-analysis-app Public

    AI-powered statistical analysis and AutoML platform for data preprocessing, hypothesis testing, visualization, predictive modeling, and report generation.

    Python

  2. Junaid-Ahmed-Rupok Junaid-Ahmed-Rupok Public

    Data Scientist | ML Engineer — 89.8% accuracy · 98.3% classification · Production ML

  3. continuum-rag-chatbot continuum-rag-chatbot Public

    🧠 A persistent memory RAG chatbot that never forgets. Uses Phi-3-mini LLM, ChromaDB vector database, and Ebbinghaus memory decay curves.

    Python 1

  4. ReproHub ReproHub Public

    Research Reproducibility Verification Platform

    Python 1

  5. DeepEnMap DeepEnMap Public

    Multi-modal deep learning framework for ordinal energy poverty risk mapping using satellite imagery and demographic data, with a custom ordinal-aware loss function.

    Jupyter Notebook