Skip to content

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

🎥 MonoTrack3D

Monocular RGB Sparse Point Cloud SLAM using Computer Vision, Epipolar Geometry, and Modern Web Technologies

Python FastAPI OpenCV React TypeScript Three.js Docker

License: MIT GitHub Stars GitHub Forks


📖 About

MonoTrack3D is a full-stack monocular visual SLAM (Simultaneous Localization and Mapping) system. It reconstructs a sparse 3D point cloud and estimates the camera trajectory from a single ordinary RGB video — no depth sensor, no stereo rig, no LiDAR.

Instead of relying on an existing SLAM framework, the entire tracking and mapping pipeline is implemented from scratch on top of OpenCV and NumPy, covering:

  • 🎯 ORB feature detection & descriptor matching
  • 📐 Essential-matrix based relative pose estimation
  • 🧱 Two-view triangulation with outlier rejection
  • 🔁 PnP-based pose refinement against persistent landmarks
  • 🧭 Camera trajectory tracking and smoothing
  • ☁️ Sparse 3D point cloud generation

A user uploads a video through the web interface, the backend processes it frame-by-frame, and the frontend renders the reconstructed scene and camera path as an interactive 3D viewport.


📸 Application Preview

🏠 Landing / Hero

image

📤 Video Upload & 3D Viewport

image

📊 Stats & Trajectory

image

✨ Key Features

🧠 Visual SLAM Pipeline

  • Custom ORB feature detector and Hamming-distance feature matcher (Lowe's ratio test)
  • Essential matrix estimation via RANSAC + pose recovery
  • Two-view triangulation with depth, reprojection-error, and parallax filtering
  • PnP + RANSAC pose refinement against accumulated 3D landmarks
  • Motion validation to reject implausible translation/rotation jumps
  • Damped PnP correction so a single bad estimate can't destabilize the trajectory

🗺️ Mapping & Tracking

  • Persistent landmark manager (sparse 3D map)
  • Keyframe manager for selected camera poses
  • Trajectory smoothing for a cleaner camera path
  • Per-run statistics: frames processed, matches, inliers, processing time

🌐 Web Application

  • Drag-and-drop / file-picker video upload
  • Real-time processing status feedback
  • Interactive 3D point cloud viewer (orbit / zoom / pan)
  • Camera trajectory chart
  • Live tracking-quality statistics dashboard

⚡ Backend

  • FastAPI REST API
  • Stateless video processing (uploads are deleted after processing)
  • Modular, dependency-injected SLAM components
  • Dockerized with OpenCV runtime dependencies preinstalled

🎨 Frontend

  • React 19 + TypeScript + Vite
  • Three.js / React Three Fiber 3D point cloud rendering
  • Type-safe API client with defensive response parsing
  • Clean, dark "lab console" UI

🏗️ System Architecture

                           +-----------------------------+
                           |          End Users          |
                           +-------------+---------------+
                                         |
                                         |  Upload RGB Video
                                         ▼
                   +--------------------------------------+
                   |     React + TypeScript Frontend       |
                   |   (Three.js Point Cloud Viewer)       |
                   +----------------+---------------------+
                                    |
                          REST API (multipart/form-data)
                                    |
                                    ▼
                   +--------------------------------------+
                   |            FastAPI Backend            |
                   |         POST /process endpoint        |
                   +----------------+---------------------+
                                    |
                                    ▼
                   +--------------------------------------+
                   |           SLAM Pipeline               |
                   |----------------------------------------|
                   |  Feature Detection (ORB)                |
                   |  Feature Matching (BF + Ratio Test)     |
                   |  Pose Estimation (Essential Matrix)     |
                   |  Triangulation (Sparse 3D Points)       |
                   |  PnP Pose Refinement                    |
                   |  Keyframe / Landmark Management         |
                   |  Trajectory Smoothing                   |
                   +----------------+---------------------+
                                    |
                                    ▼
                    JSON Result (points + trajectory + stats)
                                    |
                                    ▼
                    Rendered in 3D Viewport (Three.js)

🔄 End-to-End Workflow

Upload RGB Video
      │
      ▼
Video Decoded Frame-by-Frame (OpenCV)
      │
      ▼
ORB Feature Detection
      │
      ▼
Feature Matching (Brute-Force + Ratio Test)
      │
      ▼
Essential Matrix Pose Estimation (RANSAC)
      │
      ▼
Motion Validation
      │
      ▼
Triangulation → Sparse 3D Landmarks
      │
      ▼
PnP Pose Refinement Against Landmarks
      │
      ▼
Keyframe & Trajectory Update
      │
      ▼
Final Point Cloud + Trajectory + Stats
      │
      ▼
Rendered in Browser (3D Viewer + Charts)

📊 Project Highlights

Metric Value
SLAM Approach Monocular, feature-based
Feature Detector ORB (up to 2000 features/frame)
Pose Estimation Essential Matrix + RANSAC
Pose Refinement PnP + RANSAC
Map Type Sparse 3D point cloud
Backend Framework FastAPI
Frontend Framework React (Vite)
3D Rendering Three.js / React Three Fiber
Containerization Docker Compose (2 services)
Supported Video Formats MP4, AVI, MOV, MKV

🛠️ Technology Stack

Frontend

Technology Purpose
React 19 UI development
TypeScript Type safety
Vite Build tool & dev server
Three.js 3D rendering engine
@react-three/fiber React renderer for Three.js
@react-three/drei Three.js helper components
Lucide React Icons

Backend

Technology Purpose
FastAPI High-performance REST API framework
Python 3.11 Backend development
Uvicorn ASGI server
python-multipart Video upload handling
OpenCV (headless) Feature detection, matching, geometry
NumPy Numerical computing
SciPy Spatial data structures (KD-Tree)

Computer Vision / SLAM

Technique Purpose
ORB Keypoint detection & description
Brute-Force Hamming Matcher Feature matching with Lowe's ratio test
Essential Matrix (RANSAC) Relative camera pose estimation
Triangulation 3D landmark reconstruction
PnP + RANSAC Pose refinement against known landmarks
KD-Tree (SciPy) Nearest-neighbor spatial queries

Infrastructure

Technology Purpose
Docker Containerized backend & frontend
Docker Compose Multi-service orchestration
Nginx Frontend static serving + API proxy

📂 Project Structure

MonoTrack3d
│
├── backend
│   ├── app
│   │   ├── slam
│   │   │   ├── features.py        # ORB feature detection
│   │   │   ├── matcher.py         # Feature matching (ratio test)
│   │   │   ├── pose.py            # Essential matrix + PnP pose estimation
│   │   │   ├── triangulation.py   # 3D point triangulation
│   │   │   ├── landmarks.py       # Sparse landmark map management
│   │   │   ├── keyframes.py       # Keyframe selection & storage
│   │   │   ├── trajectory.py      # Trajectory smoothing
│   │   │   └── pipeline.py        # Orchestrates the full SLAM pipeline
│   │   └── main.py                # FastAPI app & /process endpoint
│   │
│   ├── Dockerfile
│   └── requirements.txt
│
├── frontend
│   ├── src
│   │   ├── api
│   │   │   └── slam.ts            # Typed API client
│   │   ├── components
│   │   │   ├── VideoUploader.tsx
│   │   │   └── TrajectoryChart.tsx
│   │   ├── three
│   │   │   └── PointCloud.tsx     # 3D point cloud viewer
│   │   └── App.tsx
│   │
│   ├── Dockerfile
│   ├── nginx.conf
│   └── package.json
│
├── docker-compose.yaml
└── README.md

🧠 SLAM Pipeline

The core of the project is a hand-built monocular SLAM pipeline (backend/app/slam/pipeline.py) that processes an uploaded video frame-by-frame.

Step 1 — Feature Detection

Each frame is converted to grayscale and passed through an ORB detector, extracting up to 2,000 keypoints and their binary descriptors.

↓

Step 2 — Feature Matching

Descriptors from consecutive frames are matched using a brute-force Hamming matcher, filtered with Lowe's ratio test (threshold 0.75) to discard ambiguous matches.

↓

Step 3 — Pose Estimation

An Essential Matrix is estimated from matched point correspondences via RANSAC, and the relative camera rotation/translation is recovered from it.

↓

Step 4 — Motion Validation

Estimated motion is checked against sanity thresholds (maximum translation step, maximum rotation step) to reject implausible jumps caused by bad matches.

↓

Step 5 — Triangulation

Matched keypoints are triangulated into 3D landmarks, filtered by minimum/maximum depth, reprojection error, and parallax to remove unstable points.

↓

Step 6 — PnP Pose Refinement

The current pose is refined using solvePnPRansac against the accumulated set of persistent 3D landmarks. The correction is only partially applied (alpha = 0.35) so a single noisy PnP result cannot destabilize the whole trajectory.

↓

Step 7 — Keyframe & Landmark Management

Selected frames are stored as keyframes with their pose, keypoints, and descriptors. New landmarks are added to a persistent sparse map, and existing landmarks gain new observations.

↓

Step 8 — Trajectory Smoothing & Output

The camera trajectory is smoothed, and the final sparse point cloud, trajectory, and tracking statistics are returned as JSON.


🔌 API Overview

Health

GET /
GET /health

SLAM Processing

POST /process

Request: multipart/form-data with a file field (.mp4, .avi, .mov, .mkv)

Response:

{
  "success": true,
  "filename": "video.mp4",
  "result": {
    "frames_processed": 240,
    "average_keypoints": 812.4,
    "average_matches": 214.1,
    "successful_poses": 233,
    "average_inliers": 87.6,
    "processing_time": 42.7,
    "final_position": [1.2, 0.0, 3.4],
    "points": [{ "x": 0.1, "y": 0.2, "z": 1.5 }],
    "trajectory": [{ "x": 0.0, "y": 0.0, "z": 0.0 }]
  }
}

Uploaded videos are deleted from disk immediately after processing (success or failure).


⚙️ Installation Guide

1️⃣ Clone the Repository

git clone https://github.com/Kushagrahms/MonoTrack3d.git

cd MonoTrack3d

📦 Backend Setup

Navigate to Backend

cd backend

Create Virtual Environment

python -m venv venv

Activate Environment

Windows

venv\Scripts\activate

Linux / macOS

source venv/bin/activate

Install Dependencies

pip install -r requirements.txt

Start Backend Server

uvicorn app.main:app --reload

Backend runs on:

http://localhost:8000

💻 Frontend Setup

Navigate to frontend

cd frontend

Install packages

npm install

Set the backend URL (create frontend/.env.development)

VITE_API_URL=http://127.0.0.1:8000

Run development server

npm run dev

Frontend runs on

http://localhost:5173

🐳 Running with Docker

The entire stack can be started with a single command from the project root:

docker compose up --build
Service Port Description
backend 8000 FastAPI SLAM processing API
frontend 80 Nginx-served React app (proxies /process and /health to backend)

📁 Environment Variables

Variable Location Description
VITE_API_URL frontend/.env.development Base URL of the FastAPI backend

🚀 Deployment

Backend

  • Railway
  • Render
  • AWS EC2
  • DigitalOcean

Frontend

  • Vercel
  • Netlify
  • Any static host / Nginx container

Recommended

  • Deploy via the included docker-compose.yaml on a single VM for the simplest setup.

📈 Performance Highlights

✔ Stateless, per-request SLAM processing

✔ Automatic cleanup of uploaded video files

✔ Motion validation to reject unstable pose jumps

✔ Damped PnP correction for trajectory stability

✔ Outlier-filtered triangulation (depth, reprojection error, parallax)

✔ Dockerized OpenCV runtime for consistent deployment


🎯 Future Improvements

  • Loop closure detection
  • Bundle adjustment for global map optimization
  • Real camera calibration support (instead of approximated intrinsics)
  • Dense reconstruction option
  • GPU-accelerated feature extraction
  • Multi-video / session support
  • Export point cloud (PLY / PCD)
  • Real-time streaming input (webcam)
  • CI/CD pipeline
  • Automated backend test suite

📚 Learning Outcomes

This project strengthened my understanding of:

  • Monocular visual SLAM fundamentals
  • Epipolar geometry & the Essential Matrix
  • Feature detection & descriptor matching
  • Triangulation and 3D reconstruction
  • PnP-based pose refinement
  • FastAPI backend development
  • React + Three.js 3D visualization
  • Docker-based full-stack deployment

🤝 Contributing

Contributions are welcome!

If you'd like to improve the project:

  1. Fork the repository

  2. Create your feature branch

git checkout -b feature/NewFeature
  1. Commit your changes
git commit -m "Added New Feature"
  1. Push to the branch
git push origin feature/NewFeature
  1. Open a Pull Request

⭐ Support

If you found this project helpful,

please consider giving it a ⭐ on GitHub.

It helps others discover the project and motivates future development.


👨‍💻 Author

Kushagra Shrivastava

Backend Developer • Full Stack Developer • Computer Vision Enthusiast

GitHub


📄 License

This project is licensed under the MIT License.

See the LICENSE file for more information.


⭐ Thanks for visiting this repository!

If you enjoyed this project, don't forget to leave a ⭐.

Happy Coding! 🚀