I build ML systems that hold up on messy, real-world data β documents, video, and text β and benchmark them honestly instead of trusting a demo.
π B.Tech CSE β SRM Institute of Science and Technology, KTR Β· CGPA 8.42
βοΈ Technical Manager for AI/ML β AWS Student Builder Group (AWS SBG), SRMIST
π Tech Co-Lead β IEEE GRSS Club
π¬ UROP researcher
πΌ Ex - JA Assure AI/ML Engineer Intern
πΌ IIT Ropar Intern working on GAR
π« hemishjain22@gmail.com
2 peer-reviewed publications (CVIP 2025, ICRAIS 2025) Β· 2 industry/research internships Β· ~70% estimated cut in manual data entry (on-prem document pipeline) Β· Engineered a three-layer end-to-end document intelligence pipeline.
| Role | What I did | When |
|---|---|---|
| AI/ML Research Intern β IIT Ropar | Built a benchmarking pipeline comparing FFmpeg, DeepStream, and GStreamer for real-time face detection on RTSP streams | May 2026 - Present |
| Intern β JA Assure | Built a three-layer, fully on-prem document intelligence pipeline: VLMs extract text/tables/checkboxes from PDF forms β open-source LLMs normalize to JSON β REST-based spreadsheet population. Zero external API calls; ~70% estimated reduction in manual entry | Jan - March 2026 |
| Project | What it is | Stack |
|---|---|---|
| Hiree | Built a multi-signal hiring platform that parses resumes and verifies claimed skills against live GitHub and LeetCode profiles, generating AI-powered job-fit scores and hiring verdicts via Groq LLMs with a Gemini fallback. | ext.js, React, FastAPI, Scikit-learn, Sentence-Transformers, Groq LLM, PostgreSQL |
| Savify | Built a productivity tool that ingests YouTube reels, videos, and blog URLs, extracting transcripts directly or transcribing audio via speech-to-text when captions are unavailable. | Python, Gemini API, Speech-to-Text, SQLite |
| LinguaLens | Snap or drop a medicine label, a government form, a signboard β and get a plain-language explanation tailored to you, with audio and follow-up questions. | Python, Groq API, Text-to-Speech |
| Site-sabha | One safety briefing, every worker's language, with proof that each worker understood it. | Sarvam AI: Saaras, Sarvam Vision, Sarvam-105B, Mayura, Bulbul, ffmpeg |
| Work | What it explores | Status |
|---|---|---|
| Multimodal Anemia Detection Using Convolution Neural Network and Ensemble Learning | Stacked ensemble over multiple modalities for non-invasive anemia screening | β Published β CVIP 2025 |
| Suicidal Text Detection Using Machine Learning and Large Language Model | NLP approach to classifying distress and emotion in text | β Published β ICRAIS 2025 |
| Robustness of Lightweight Retrieval on Metadata-Stripped Medical Images | A study on robust retrieval of metadata-stripped medical images using lightweight and deep learning approaches under various image degradations. | π Paper writting |
| Decoding Group Affect: A Systematic Review and Research Agenda | A study on group-level affect recognition from multi-person scenes using lightweight handcrafted and deep learning approaches under real-world degradations such as occlusion, crowding, and dynamic group dynamics. | π Paper writting |
| Water, Not Watts: Re-examining the Water Footprint of Liquid-Cooled AI Data Centres in India | A study on the true water footprint of liquid-cooled AI data centres in India using a three-axis taxonomy and disclosure-completeness formalization under a coal-intensive grid and warm-humid, water-stressed climate. | π Paper done |
| Battery-Evaluation | Remaining-useful-life prediction on NASA PCoE and randomized-usage Li-ion battery datasets, using a thermoelastic-stress feature stage validated against an independent finite-difference solver. | π¬ In progress |
| VLM benchmark on Indian government forms | Open-source VLMs (Qwen2.5-VL, MinerU2, DeepSeek-OCR, GLM-4.5V) on checkbox-state detection and multi-column table alignment | π¬ In progress |
| Architecture and Precision Leakage in Quantised Edge AI | Neural network architecture is recoverable from software-accessible timing, power and thermal side channels on embedded AI accelerators, and the degree of leakage is materially affected by the quantisation precision at which the model is deployed. | π¬ In progress |
- Extending HireSense with a retrieval layer and an evaluation harness
- Planning a Government Scheme Informer: deterministic rules engine for eligibility, LLM used only to explain the result
- Leading AI/ML tracks for the AWS SBG, SRM
- Leading AI/ML tracks for the IEEE GRSS, SRM
|
|
|
|
|
|
If a benchmark in one of my repos disagrees with the paper it's based on, that's usually the point.



