Skip to content
View saurabh23011's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report saurabh23011

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
saurabh23011/README.md

Saurabh Kumar Singh

AI/ML Engineer — Computer Vision · Generative AI · Vision-Language Model Evaluation

Portfolio · LinkedIn · Medium · Email


About

I build and evaluate vision and multimodal systems — detection and OCR pipelines that run on real video, and evaluation harnesses that tell you whether a model is actually any good.

  • 🔬 Currently at Snorkel AI, working on multimodal data and model evaluation
  • 🧠 Also building AI features at Claw LegalTech
  • 🎯 Focus: object detection, OCR, depth/3D from monocular video, VLM evaluation, agentic LLM pipelines
  • 🎓 M.Sc. Information Technology (AI & ML), IIIT Lucknow
  • 📍 India — open to remote and relocation
  • 📫 saurabhsingh802213@gmail.com

Featured projects

Project What it does Stack
YOLOv11 Object Detection Detects objects in images, uploaded video or a live webcam feed, with an adjustable confidence threshold Ultralytics YOLOv11, OpenCV, Streamlit
Multilingual OCR (Hindi + English) Extracts and keyword-searches Devanagari and Latin text in uploaded images EasyOCR, Flask
Industry Safety Detection PPE and safety-gear detection for industrial footage, containerised for deployment YOLOv7, Docker, Flask
2D → 3D Video Conversion Turns ordinary 2D video into stereoscopic 3D using monocular depth estimation Depth estimation, OpenCV, Flask
Presentation Gesture Control Drives presentation slides with hand gestures from a webcam — no clicker OpenCV, cvzone hand tracking
Video Summarizer Agent Agent that watches an uploaded video and answers questions about it phidata agents, Gemini, Streamlit

Tech

Languages Python · C++ · TypeScript / JavaScript · SQL

ML & Vision PyTorch · Hugging Face (transformers, datasets, accelerate) · PEFT / LoRA · timm · torchvision · OpenCV · albumentations · scikit-learn

Serving & Tooling FastAPI · Flask · Docker · Streamlit · pytest · ruff · MLflow / Weights & Biases

Pinned Loading

  1. YOLOv11-Object-Detection YOLOv11-Object-Detection Public

    Streamlit app that runs YOLOv11 detection on images, uploaded video, or a live webcam feed.

    Python 2

  2. 2D-To-3D-video-using-the-computer-vision- 2D-To-3D-video-using-the-computer-vision- Public

    Turns ordinary 2D video into stereoscopic 3D using monocular depth estimation.

    C 1

  3. Industry-Safety-Detection-Using-Computer-Vision Industry-Safety-Detection-Using-Computer-Vision Public

    YOLOv7-based PPE and industrial safety-gear detection, packaged with Docker for deployment.

    1

  4. Multilingual-OCR-Application-Hindi-and-English- Multilingual-OCR-Application-Hindi-and-English- Public

    OCR web app that extracts and searches Hindi (Devanagari) and English text in uploaded images.

    HTML 1

  5. presentation_gesture-using-computer-vision presentation_gesture-using-computer-vision Public

    Control presentation slides with hand gestures from a webcam feed — no clicker needed.

    Python 1

  6. Video-Summarizer-Agent Video-Summarizer-Agent Public

    Agentic app that watches an uploaded video and answers questions about its content.

    Python