Anything ever filmed, in 4D.
Hack the North 2026 finalist (top 12 of 350 teams) and 2nd place in the OpenAI track.
Orbis turns an ordinary video into a 4D scene: a full 3D world, plus time. The room is rebuilt in 3D, and the people and objects in it move through it as they did in the recording, in sync with the original sound. You are not locked to the camera. Put on a Quest, or use WASD in a browser, and walk anywhere while the moment plays: stand behind someone, watch a throw from the far side of the room, see it from angles the video never had.
In one scene you can also step in: walk up to a person and talk to them, or pick up the bottle they are holding, and it behaves like a physical object in your hand.
It works on phone clips and on movie shots. Built at Hack the North 2026.
flowchart LR
A[Video] --> B[Solve cameras and depth]
A --> C[Remove people, fill the holes]
C --> D[Generate the room]
B --> E[Rebuild and animate each person]
D --> F[Fit scale, floor and placement]
E --> F
F --> G[Walk around in it]
| Step | What happens | Built on |
|---|---|---|
| Cameras | Find where the camera was in every frame, plus depth | Pi3X |
| Clean | Mask the people out and fill the holes | Mask R-CNN, LaMa |
| Room | Generate a full Gaussian-splat room from the cleaned video | Marble, OpenAI vision |
| Correct | Train the room against the real frames (optional) | gsplat |
| People | Track each person and rebuild them as an animated avatar | LHM |
| Placement | Fit scale and floor so feet land where they stood | |
| Viewer | Walk around in a browser or a Quest, synced to the source video | Three.js, Spark, WebXR, OpenAI Realtime |
GPU stages run on Modal. Scenes are stored in S3.
| Idea | Tried | What happened |
|---|---|---|
| Rebuild the room from the footage | Depth Anything 3, VGGT, Brush | Great from the camera's spot, broken everywhere else |
| Generate only the missing views | Stable Virtual Camera, GeoNVS, GEN3C, Lyra | A courtyard that did not exist, grey mush, a shop sign turned to gibberish |
| Rebuild people from depth | V-DPM, PIFuHD | Melted limbs |
| Refine body poses against the video | our own optimiser | Worse. One camera cannot tell how far away a hand is |
A camera walking forward only records about 41° off its own path, so you cannot rebuild what was never filmed. What worked was generating the whole room, then correcting it with the footage.
Full list in the research log. Dead ends in known limits.
You need Bun 1.2.21+ and Chrome.
git clone --single-branch --branch main https://github.com/dtpu/OrbisEngine.git
cd OrbisEngine
bun install --frozen-lockfile
bun run demoOpen http://127.0.0.1:5399/demo.html and pick a scene.
Scenes are not in this repo. You have two ways to get one:
- Team credentials: put the read-only keys in
.env.local(see.env.example). - Your own clip: see below.
| Key | Does |
|---|---|
| W A S D | Walk |
| Drag or click | Look around |
| Shift | Run |
| Space | Play or pause |
| R | Reset |
| M | Overhead map |
- Enable Developer Mode and plug in the headset.
- Run
adb reverse tcp:5399 tcp:5399. - In the Meta Browser open
http://localhost:5399/demo.html?xr=1and press Enter VR.
B on the right controller opens the scene list. The desktop page shows what the headset sees.
To talk to a person and pick up the bottle, open fourd.html?demo=elevator&interact=1&xr=1 with
OPENAI_API_KEY set on the server. More options in developing.
You need uv, Python 3.11 or 3.12, FFmpeg, and keys for Modal, World
Labs and OpenAI in .env.author (.env.example lists them).
uv sync --locked --group inference
set -a; source .env.author; set +a
uv run --locked --group inference scripts/run_clip.py \
--clip /path/to/clip.mp4 --name myclip \
--marble video --all-people --fps 12 --skip-finetune --no-publish- The run stops once so you can check the cleaned frames. Run it again with
--gate-passto continue. - Running the same command after a crash resumes where it stopped.
- A 10-second clip takes about 25 minutes, roughly $0.15 of Modal GPU plus 1,600 Marble credits.
- Read paid recovery before retrying a failed paid stage.
Clips that work: one continuous shot, 5 to 15 seconds, people at least 150 px tall, some sideways camera movement, decent light.
- Devpost submission: the build story, what we tried first, and how OpenAI, Codex and Devin were used.
- Developing: code map, checks, Quest options, pipeline flags, repo history.
- Research log and known limits.
- Audio, objects, interactive scene, unattended runs.
Aayan Karmali, Austin Jian, Daniel Pu and James Li.
The Tears of Steel scene uses excerpts of (CC) Blender Foundation | mango.blender.org, CC BY 3.0. Phone clips are our own footage. Movie excerpts used in testing are not redistributed here.
Apache-2.0. Some models the pipeline uses are research-only; see licences.
