Monocular RGB Sparse Point Cloud SLAM using Computer Vision, Epipolar Geometry, and Modern Web Technologies
MonoTrack3D is a full-stack monocular visual SLAM (Simultaneous Localization and Mapping) system. It reconstructs a sparse 3D point cloud and estimates the camera trajectory from a single ordinary RGB video — no depth sensor, no stereo rig, no LiDAR.
Instead of relying on an existing SLAM framework, the entire tracking and mapping pipeline is implemented from scratch on top of OpenCV and NumPy, covering:
- 🎯 ORB feature detection & descriptor matching
- 📐 Essential-matrix based relative pose estimation
- 🧱 Two-view triangulation with outlier rejection
- 🔁 PnP-based pose refinement against persistent landmarks
- 🧭 Camera trajectory tracking and smoothing
- ☁️ Sparse 3D point cloud generation
A user uploads a video through the web interface, the backend processes it frame-by-frame, and the frontend renders the reconstructed scene and camera path as an interactive 3D viewport.
- Custom ORB feature detector and Hamming-distance feature matcher (Lowe's ratio test)
- Essential matrix estimation via RANSAC + pose recovery
- Two-view triangulation with depth, reprojection-error, and parallax filtering
- PnP + RANSAC pose refinement against accumulated 3D landmarks
- Motion validation to reject implausible translation/rotation jumps
- Damped PnP correction so a single bad estimate can't destabilize the trajectory
- Persistent landmark manager (sparse 3D map)
- Keyframe manager for selected camera poses
- Trajectory smoothing for a cleaner camera path
- Per-run statistics: frames processed, matches, inliers, processing time
- Drag-and-drop / file-picker video upload
- Real-time processing status feedback
- Interactive 3D point cloud viewer (orbit / zoom / pan)
- Camera trajectory chart
- Live tracking-quality statistics dashboard
- FastAPI REST API
- Stateless video processing (uploads are deleted after processing)
- Modular, dependency-injected SLAM components
- Dockerized with OpenCV runtime dependencies preinstalled
- React 19 + TypeScript + Vite
- Three.js / React Three Fiber 3D point cloud rendering
- Type-safe API client with defensive response parsing
- Clean, dark "lab console" UI
+-----------------------------+
| End Users |
+-------------+---------------+
|
| Upload RGB Video
▼
+--------------------------------------+
| React + TypeScript Frontend |
| (Three.js Point Cloud Viewer) |
+----------------+---------------------+
|
REST API (multipart/form-data)
|
▼
+--------------------------------------+
| FastAPI Backend |
| POST /process endpoint |
+----------------+---------------------+
|
▼
+--------------------------------------+
| SLAM Pipeline |
|----------------------------------------|
| Feature Detection (ORB) |
| Feature Matching (BF + Ratio Test) |
| Pose Estimation (Essential Matrix) |
| Triangulation (Sparse 3D Points) |
| PnP Pose Refinement |
| Keyframe / Landmark Management |
| Trajectory Smoothing |
+----------------+---------------------+
|
▼
JSON Result (points + trajectory + stats)
|
▼
Rendered in 3D Viewport (Three.js)
Upload RGB Video
│
▼
Video Decoded Frame-by-Frame (OpenCV)
│
▼
ORB Feature Detection
│
▼
Feature Matching (Brute-Force + Ratio Test)
│
▼
Essential Matrix Pose Estimation (RANSAC)
│
▼
Motion Validation
│
▼
Triangulation → Sparse 3D Landmarks
│
▼
PnP Pose Refinement Against Landmarks
│
▼
Keyframe & Trajectory Update
│
▼
Final Point Cloud + Trajectory + Stats
│
▼
Rendered in Browser (3D Viewer + Charts)
| Metric | Value |
|---|---|
| SLAM Approach | Monocular, feature-based |
| Feature Detector | ORB (up to 2000 features/frame) |
| Pose Estimation | Essential Matrix + RANSAC |
| Pose Refinement | PnP + RANSAC |
| Map Type | Sparse 3D point cloud |
| Backend Framework | FastAPI |
| Frontend Framework | React (Vite) |
| 3D Rendering | Three.js / React Three Fiber |
| Containerization | Docker Compose (2 services) |
| Supported Video Formats | MP4, AVI, MOV, MKV |
| Technology | Purpose |
|---|---|
| React 19 | UI development |
| TypeScript | Type safety |
| Vite | Build tool & dev server |
| Three.js | 3D rendering engine |
| @react-three/fiber | React renderer for Three.js |
| @react-three/drei | Three.js helper components |
| Lucide React | Icons |
| Technology | Purpose |
|---|---|
| FastAPI | High-performance REST API framework |
| Python 3.11 | Backend development |
| Uvicorn | ASGI server |
| python-multipart | Video upload handling |
| OpenCV (headless) | Feature detection, matching, geometry |
| NumPy | Numerical computing |
| SciPy | Spatial data structures (KD-Tree) |
| Technique | Purpose |
|---|---|
| ORB | Keypoint detection & description |
| Brute-Force Hamming Matcher | Feature matching with Lowe's ratio test |
| Essential Matrix (RANSAC) | Relative camera pose estimation |
| Triangulation | 3D landmark reconstruction |
| PnP + RANSAC | Pose refinement against known landmarks |
| KD-Tree (SciPy) | Nearest-neighbor spatial queries |
| Technology | Purpose |
|---|---|
| Docker | Containerized backend & frontend |
| Docker Compose | Multi-service orchestration |
| Nginx | Frontend static serving + API proxy |
MonoTrack3d
│
├── backend
│ ├── app
│ │ ├── slam
│ │ │ ├── features.py # ORB feature detection
│ │ │ ├── matcher.py # Feature matching (ratio test)
│ │ │ ├── pose.py # Essential matrix + PnP pose estimation
│ │ │ ├── triangulation.py # 3D point triangulation
│ │ │ ├── landmarks.py # Sparse landmark map management
│ │ │ ├── keyframes.py # Keyframe selection & storage
│ │ │ ├── trajectory.py # Trajectory smoothing
│ │ │ └── pipeline.py # Orchestrates the full SLAM pipeline
│ │ └── main.py # FastAPI app & /process endpoint
│ │
│ ├── Dockerfile
│ └── requirements.txt
│
├── frontend
│ ├── src
│ │ ├── api
│ │ │ └── slam.ts # Typed API client
│ │ ├── components
│ │ │ ├── VideoUploader.tsx
│ │ │ └── TrajectoryChart.tsx
│ │ ├── three
│ │ │ └── PointCloud.tsx # 3D point cloud viewer
│ │ └── App.tsx
│ │
│ ├── Dockerfile
│ ├── nginx.conf
│ └── package.json
│
├── docker-compose.yaml
└── README.md
The core of the project is a hand-built monocular SLAM pipeline (backend/app/slam/pipeline.py) that processes an uploaded video frame-by-frame.
Each frame is converted to grayscale and passed through an ORB detector, extracting up to 2,000 keypoints and their binary descriptors.
↓
Descriptors from consecutive frames are matched using a brute-force Hamming matcher, filtered with Lowe's ratio test (threshold 0.75) to discard ambiguous matches.
↓
An Essential Matrix is estimated from matched point correspondences via RANSAC, and the relative camera rotation/translation is recovered from it.
↓
Estimated motion is checked against sanity thresholds (maximum translation step, maximum rotation step) to reject implausible jumps caused by bad matches.
↓
Matched keypoints are triangulated into 3D landmarks, filtered by minimum/maximum depth, reprojection error, and parallax to remove unstable points.
↓
The current pose is refined using solvePnPRansac against the accumulated set of persistent 3D landmarks. The correction is only partially applied (alpha = 0.35) so a single noisy PnP result cannot destabilize the whole trajectory.
↓
Selected frames are stored as keyframes with their pose, keypoints, and descriptors. New landmarks are added to a persistent sparse map, and existing landmarks gain new observations.
↓
The camera trajectory is smoothed, and the final sparse point cloud, trajectory, and tracking statistics are returned as JSON.
GET /
GET /health
POST /process
Request: multipart/form-data with a file field (.mp4, .avi, .mov, .mkv)
Response:
{
"success": true,
"filename": "video.mp4",
"result": {
"frames_processed": 240,
"average_keypoints": 812.4,
"average_matches": 214.1,
"successful_poses": 233,
"average_inliers": 87.6,
"processing_time": 42.7,
"final_position": [1.2, 0.0, 3.4],
"points": [{ "x": 0.1, "y": 0.2, "z": 1.5 }],
"trajectory": [{ "x": 0.0, "y": 0.0, "z": 0.0 }]
}
}Uploaded videos are deleted from disk immediately after processing (success or failure).
git clone https://github.com/Kushagrahms/MonoTrack3d.git
cd MonoTrack3d
cd backend
python -m venv venv
Windows
venv\Scripts\activate
Linux / macOS
source venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --reload
Backend runs on:
http://localhost:8000
Navigate to frontend
cd frontend
Install packages
npm install
Set the backend URL (create frontend/.env.development)
VITE_API_URL=http://127.0.0.1:8000
Run development server
npm run dev
Frontend runs on
http://localhost:5173
The entire stack can be started with a single command from the project root:
docker compose up --build
| Service | Port | Description |
|---|---|---|
| backend | 8000 | FastAPI SLAM processing API |
| frontend | 80 | Nginx-served React app (proxies /process and /health to backend) |
| Variable | Location | Description |
|---|---|---|
VITE_API_URL |
frontend/.env.development |
Base URL of the FastAPI backend |
- Railway
- Render
- AWS EC2
- DigitalOcean
- Vercel
- Netlify
- Any static host / Nginx container
- Deploy via the included
docker-compose.yamlon a single VM for the simplest setup.
✔ Stateless, per-request SLAM processing
✔ Automatic cleanup of uploaded video files
✔ Motion validation to reject unstable pose jumps
✔ Damped PnP correction for trajectory stability
✔ Outlier-filtered triangulation (depth, reprojection error, parallax)
✔ Dockerized OpenCV runtime for consistent deployment
- Loop closure detection
- Bundle adjustment for global map optimization
- Real camera calibration support (instead of approximated intrinsics)
- Dense reconstruction option
- GPU-accelerated feature extraction
- Multi-video / session support
- Export point cloud (PLY / PCD)
- Real-time streaming input (webcam)
- CI/CD pipeline
- Automated backend test suite
This project strengthened my understanding of:
- Monocular visual SLAM fundamentals
- Epipolar geometry & the Essential Matrix
- Feature detection & descriptor matching
- Triangulation and 3D reconstruction
- PnP-based pose refinement
- FastAPI backend development
- React + Three.js 3D visualization
- Docker-based full-stack deployment
Contributions are welcome!
If you'd like to improve the project:
-
Fork the repository
-
Create your feature branch
git checkout -b feature/NewFeature
- Commit your changes
git commit -m "Added New Feature"
- Push to the branch
git push origin feature/NewFeature
- Open a Pull Request
If you found this project helpful,
please consider giving it a ⭐ on GitHub.
It helps others discover the project and motivates future development.
Backend Developer • Full Stack Developer • Computer Vision Enthusiast
This project is licensed under the MIT License.
See the LICENSE file for more information.
If you enjoyed this project, don't forget to leave a ⭐.
Happy Coding! 🚀