Pyongyang-VLM là MVP video person retrieval / pedestrian search: nhập mô tả bằng ngôn ngữ tự nhiên, tìm người trong video, render bounding box lên track phù hợp và xuất video kết quả.
- Natural-language person search
- Video person detection, tracking và cropping
- TBPS-CLIP image-text retrieval
- Track-level ranking
- Gradio demo UI
- Bounding box video rendering
- Optional conservative track stitching cho các track bị tách sau occlusion
Raw query
-> Query Understanding
-> normalized_text
Video
-> Vision Pipeline
-> tracklets / crops / bboxes
-> video embedding index
normalized_text + video index
-> Matching Engine
-> best_track_id + ranking
-> optional track stitching
-> renderer
-> output video
Luồng end-to-end hiện tại:
- User upload hoặc truyền video.
- Module 2 Vision Pipeline detect/track/crop persons.
- Demo pipeline build video embedding index bằng TBPS-CLIP image encoder.
- User nhập query tự nhiên.
- Module 1 Query Understanding dùng Vertex AI Gemini để normalize query.
- Module 3 Matching Engine encode text query bằng TBPS-CLIP, tính cosine similarity và rank track candidates.
- Có thể bật conservative track stitching để nối các
track_idcó khả năng là cùng người. - Renderer vẽ bbox lên người được chọn và xuất MP4.
- Python
>=3.12theopyproject.toml uvđược khuyến nghị để quản lý môi trườngffmpegđể render video H.264 MP4- CUDA là optional nhưng rất nên dùng khi export embeddings / chạy TBPS-CLIP
- Google Cloud project có Vertex AI enabled cho Module 1
- TBPS-CLIP checkpoint tại
weights/checkpoint_best.pth
git clone https://github.com/FWD-LeTung/Pyongyang-VLM.git
cd Pyongyang-VLM
pip install uv
uv syncNếu máy chưa có ffmpeg, cài bằng package manager của hệ điều hành. Ví dụ Ubuntu:
sudo apt-get update
sudo apt-get install -y ffmpegModule 1 cần Vertex AI Gemini. Tạo .env từ .env.example hoặc export trực tiếp các biến sau:
export GCP_PROJECT_ID="your-project-id"
export GCP_LOCATION="us-central1"
export GEMINI_MODEL="gemini-2.5-flash"Đăng nhập local bằng Application Default Credentials:
gcloud auth application-default loginTrên Colab, đăng nhập bằng gcloud auth login và gcloud auth application-default login.
Không commit credential, token, service account key hoặc nội dung .env thật vào Git.
TBPS-CLIP checkpoint cần nằm tại:
weights/checkpoint_best.pth
Tải checkpoint tại:
https://drive.google.com/file/d/1Y94znxB7J7UeczulGH9sAzZpScHOhaHT/view?usp=sharing
Checkpoint không được commit vào Git. Repo hiện có weights/yolov8n.pt cho detector; weights/checkpoint_best.pth nằm trong .gitignore.
Repo có test video tại:
data/test_videos/cctv_full_h264.mp4
uv run python demo/export_video_embeddings.py \
--video data/test_videos/cctv_full_h264.mp4 \
--output outputs/video_index/cctv_full_h264.pt \
--max-frames 0 \
--device cuda \
--precision fp16--max-frames 0 nghĩa là xử lý full video. Nếu máy không có CUDA, dùng:
uv run python demo/export_video_embeddings.py \
--video data/test_videos/cctv_full_h264.mp4 \
--output outputs/video_index/cctv_full_h264.pt \
--max-frames 0 \
--device cpu \
--precision fp32uv run python demo/query_video_embeddings.py \
--index outputs/video_index/cctv_full_h264.pt \
--query "viết mô tả bằng tiếng Anh hoặc tiếng Việt" \
--device cpu \
--precision fp32 \
--top-k 10 \
--save-debug-imagesDebug images mặc định được lưu vào:
outputs/debug_check/
uv run python demo/render_query_result.py \
--index outputs/video_index/cctv_full_h264.pt \
--video data/test_videos/cctv_full_h264.mp4 \
--query "viết mô tả bằng tiếng Anh hoặc tiếng Việt" \
--output outputs/rendered_result.mp4 \
--device cpu \
--precision fp32 \
--hold-frames 15 \
--force-renderAuto stitching là query-conditioned và conservative: chỉ nối track khi appearance, temporal gap, overlap và mutual-best checks đủ chắc.
uv run python demo/render_query_result.py \
--index outputs/video_index/cctv_full_h264.pt \
--video data/test_videos/cctv_full_h264.mp4 \
--query "viết mô tả bằng tiếng Anh hoặc tiếng Việt" \
--output outputs/rendered_result_stitched.mp4 \
--device cpu \
--precision fp32 \
--hold-frames 15 \
--force-render \
--auto-stitchDebug stitch candidates cho một target track:
uv run python demo/debug_track_stitching.py \
--index outputs/video_index/cctv_full_h264.pt \
--target-track-id 261Chạy local:
uv run python app_gradio.pyChạy với public link:
uv run python app_gradio.py --shareUI flow hiện tại:
- Upload video.
- Chọn
Precision. - Bấm
Process Video. - Nhập query.
- Bật/tắt
Auto stitch fragmented tracksnếu cần. - Bấm
Search & Render. - Xem output video trong UI.
Trong code hiện tại, Gradio tự chọn device bằng CUDA nếu có, nếu không thì CPU. Checkpoint mặc định là weights/checkpoint_best.pth.
Colab notebook:
https://colab.research.google.com/drive/1oGY7A0diKy6IbBVVXtns7o9jTdm53Hf8?usp=sharing
Các file config chính:
config/vision_pipeline.yamlconfig/matching_engine.yaml
Một số setting thường cần chỉnh:
reader.processing_fps: FPS xử lý video trong Vision Pipelinedetector.device: device cho detectordetector.confidence_threshold: ngưỡng detect persontracker.track_high_thresh,tracker.track_low_thresh,tracker.match_thresh: threshold trackerretrieval.checkpoint_path: path TBPS-CLIP checkpointretrieval.device: device mặc định của Matching Engineretrieval.precision:fp16hoặcfp32
CLI args như --device, --precision, --checkpoint, --vision-config, --matching-config có thể override config trong các demo script.
Các file output thường sinh ra:
outputs/video_index/*.pt: video embedding indexoutputs/debug_check/: debug frame/crop imagesoutputs/rendered_result.mp4: video render bboxoutputs/rendered_result_stitched.mp4: video render có auto stitchingoutputs/gradio_sessions/: session data của Gradio UI
outputs/, logs/, .env và weights/checkpoint_best.pth đang nằm trong .gitignore. Không commit checkpoint, video lớn, output .pt hoặc output .mp4.
Kiểm tra:
GCP_PROJECT_ID,GCP_LOCATION,GEMINI_MODEL- Vertex AI API đã enabled
gcloud auth application-default login- IAM permission của account đang dùng
- Model name còn hợp lệ với region đã chọn
Dùng CPU:
--device cpu --precision fp32Hoặc bật GPU runtime / cài CUDA đúng với môi trường.
Renderer cần ffmpeg để xuất H.264 MP4. Cài ffmpeg trước khi render.
Nếu OpenCV hoặc browser không đọc được video, convert sang H.264:
ffmpeg -i input.mp4 -c:v libx264 -pix_fmt yuv420p -movflags +faststart output_h264.mp4Render, upload/download và encode full video có thể chậm. Dùng video ngắn hơn hoặc đặt --max-frames nhỏ để test nhanh.
Detector/tracker có thể bị ID switch hoặc track fragmentation khi occlusion. Thử --auto-stitch hoặc inspect bằng demo/debug_track_stitching.py.
- Đây là MVP demo, chưa phải production multi-user system.
- Accuracy phụ thuộc detector/tracker, chất lượng video và TBPS-CLIP checkpoint.
- Track fragmentation vẫn có thể xảy ra khi người bị che khuất hoặc cảnh đông.
- Auto stitching cố tình conservative, nên có thể không nối nếu chưa đủ chắc.
- Full-video embedding export và rendering có thể chậm trên CPU.
uv run python -m compileall src demo app_gradio.py
uv run pytestgit status
uv run python -m compileall src demo app_gradio.py
uv run pytest