AI/ML Engineer — Computer Vision · Generative AI · Vision-Language Model Evaluation
Portfolio · LinkedIn · Medium · Email
I build and evaluate vision and multimodal systems — detection and OCR pipelines that run on real video, and evaluation harnesses that tell you whether a model is actually any good.
- 🔬 Currently at Snorkel AI, working on multimodal data and model evaluation
- 🧠 Also building AI features at Claw LegalTech
- 🎯 Focus: object detection, OCR, depth/3D from monocular video, VLM evaluation, agentic LLM pipelines
- 🎓 M.Sc. Information Technology (AI & ML), IIIT Lucknow
- 📍 India — open to remote and relocation
- 📫 saurabhsingh802213@gmail.com
| Project | What it does | Stack |
|---|---|---|
| YOLOv11 Object Detection | Detects objects in images, uploaded video or a live webcam feed, with an adjustable confidence threshold | Ultralytics YOLOv11, OpenCV, Streamlit |
| Multilingual OCR (Hindi + English) | Extracts and keyword-searches Devanagari and Latin text in uploaded images | EasyOCR, Flask |
| Industry Safety Detection | PPE and safety-gear detection for industrial footage, containerised for deployment | YOLOv7, Docker, Flask |
| 2D → 3D Video Conversion | Turns ordinary 2D video into stereoscopic 3D using monocular depth estimation | Depth estimation, OpenCV, Flask |
| Presentation Gesture Control | Drives presentation slides with hand gestures from a webcam — no clicker | OpenCV, cvzone hand tracking |
| Video Summarizer Agent | Agent that watches an uploaded video and answers questions about it | phidata agents, Gemini, Streamlit |
Languages Python · C++ · TypeScript / JavaScript · SQL
ML & Vision PyTorch · Hugging Face (transformers, datasets, accelerate) · PEFT / LoRA · timm · torchvision · OpenCV · albumentations · scikit-learn
Serving & Tooling FastAPI · Flask · Docker · Streamlit · pytest · ruff · MLflow / Weights & Biases