Skip to content

Repository files navigation

Wowza Streaming Engine · Video Intelligence Framework

Real-time computer vision for Wowza Streaming Engine — object detection, scene understanding, vision-language analysis, and synthetic video detection. Runs on your own servers, so your video never leaves your network.

Using VIF, incoming streams in Wowza Streaming Engine can be matched for real-time AI analysis across edge and cloud environments. The stream name determines which type of AI analysis is performed. This repository contains the Docker Compose setup files, working default configuration, and sample videos. VIF can also be installed directly on an existing Wowza Streaming Engine server using platform installers available from the Wowza portal.

Docker WSE License


How VIF Works

VIF has two parts:

  • Video Intelligence Service (VIS) — the analysis engine. Runs object detection and scene understanding models on an NVIDIA GPU, and coordinates VLM and synthetic video detection when enabled.
  • Video Intelligence Controller (VIC) — a plugin inside Wowza Streaming Engine (WSE). It uses the WSE transcoder to pull frames from your streams and send them to VIS over a WebSocket connection.

For a first test, one server with a supported GPU is simplest. For production, we recommend running VIS on its own GPU server, connected to the Controller over a dedicated WebSocket connection.

Tip

No-transcoder mode: VIC can skip transcoding entirely, but then it can only grab keyframes (your GOP sets the analysis rate) and overlays are not available. See Install without Docker for details.

Quick Summary

What it is Ready-to-run WSE configuration with prebuilt plugin JARs for video intelligence workflows
Primary workflows Object detection and scene understanding
Optional workflows VLM analysis and Synthetic Video Detection
How it runs docker compose up starts wse (Wowza Streaming Engine), manager (Engine Manager UI), and video-intelligence-service-gpu (Video Intelligence Service, or VIS, running on GPU). A one-shot vis-init helper runs first to prepare VIS mounts.
Alternative workflow (optional) docker compose --profile wse up starts only wse and manager when connecting to a remote VIS endpoint.
VI Service deployment Connect to a remote VI Service instance (wss://) or run VIF locally via Docker

Prerequisites

Tip

Installing without Docker? VIF also ships as platform installers (Windows, Linux x86_64, Linux Arm64) for existing Wowza Streaming Engine 4.11.1+ servers. Download them from the Wowza portal and follow the Install without Docker guide. Non-Docker installers provide full access to Object Detection and Scene Understanding. Vision Language Models and Synthetic Video Detection require additional configuration.

Table of Contents

Repository Layout

.
|- docker-compose.yaml                 # WSE + Manager + VIS containers
|- .env.example                        # Environment variables template
|- wse.standalone
|  |- conf/                            # WSE configuration
|  |- conf.modules/                    # VIF plugin configuration
|  |- lib/                             # Plugin JARs mounted into WSE
|  |- manager/                         # Manager UI extension assets
|  |- transcoder/                      # Transcoder templates used by VIF workflows
|- videos/                             # Sample VOD files for test publishing
|- docs/                               # Deployment docs

Local directories such as logs/, tmp/, vis/, and wse/ are gitignored. In the current docker-compose.yaml, WSE bind mounts and the VIS models mount are enabled. Docker Compose creates the local ./wse/* and ./vis/models directories on first startup if they do not already exist.

  • vis/ # Optional local VI service assets. ./vis/models/ stores model files, model weights/checkpoints (.pth), and TensorRT engines.
  • wse/ # Local bind-mounted WSE runtime folders for Docker (for example conf/, content/, transcoder/, logs/). This is separate from wse.standalone/ — see Persistent Volume Mounts below.
  • logs/ # Runtime log output

What you usually edit:

  • .env - license key, admin credentials, player/API keys
  • VIF stream/default configuration - use the VIF configuration page in WSE Manager, or the WSE REST API
  • wse/conf.modules/vif/ - if local WSE volume mounts are enabled, advanced users can directly edit VIF config files. Default.json defines global defaults and per-stream files (for example live_objectDotStar.json) override them.

Running Wowza VIF Locally (Quick Start)

Make sure you have a valid Wowza Streaming Engine key.

Before starting, confirm your machine meets the Compute Requirements (Self-Hosted VIF). The default local workflow runs all three containers (wse, manager, and video-intelligence-service-gpu).

Tip

Already running WSE 4.11.1+? You can add VIF to your existing server using platform installers instead of Docker. See wse.standalone/README.md for the standalone plugin setup.

1. Create .env from the example and edit values:

cp .env.example .env

Example:

WSE_LICENSE_KEY=REPLACE_WITH_YOUR_WSE_LICENSE_KEY
WSE_ADMIN_USER=admin
WSE_ADMIN_PASSWORD=CHANGE_THIS_PASSWORD
VIS_PROTOCOL=ws
VIS_HOST=video-intelligence-service.docker
VIS_PORT=5001
VIS_API_KEY=REPLACE_WITH_YOUR_VIS_API_KEY
VIS_LICENSE=REPLACE_WITH_YOUR_VIS_LICENSE

Important

VIS_HOST value depends on your deployment:

  • Local Docker: Set VIS_HOST=video-intelligence-service.docker. WSE and VIS run in the same Docker network; using localhost will not work because it resolves to the WSE container itself, not the VIS container.
  • Remote VIS: Set VIS_HOST to the hostname or IP of your remote VIS instance.

Caution

WSE_ADMIN_USER and WSE_ADMIN_PASSWORD bootstrap local Manager/REST access. Do not keep defaults — use a strong password.
VIS_API_KEY protects access to the Video Intelligence Service. Use a long, random, high-entropy key and rotate it regularly.

2. Verify your GPU and Docker can see it:

First, confirm the host GPU:

nvidia-smi

Then confirm Docker can reach it:

docker run --rm --gpus 'all' -e NVIDIA_DRIVER_CAPABILITIES=video,compute,utility nvidia/cuda:12.8.0-base-ubuntu22.04 nvidia-smi

You're ready when you see:

  • Driver version 570 or higher
  • CUDA version 12.8 or higher
  • Your GPU listed by name

Note

If the Docker command fails, you may need to install the NVIDIA Container Toolkit.

3. Start WSE + Manager + VIS:

docker compose up

Tip

If you change images or mounted config between runs, run docker compose down first.

4. Verify services are running:

Service URL
Manager UI http://localhost:8088
WSE REST API http://localhost:8087

Exposed ports:

Port Protocol Purpose
80 HTTP HLS
443 HTTPS TLS
554 RTSP RTSP
1935 TCP RTMP
8087 HTTP REST API
8089 HTTP REST API docs (if enabled)
10000 UDP SRT

Note

  • First startup can take 20+ minutes because VIS may download model weights and compile TensorRT engines for your GPU.
  • After restart, VIS usually needs about 1 to 3 minutes to become ready.

Persistent Volume Mounts

In the current docker-compose.yaml, WSE bind mounts and the VIS models mount are enabled by default. These map host folders into the containers:

  • ./wse/conf -> /usr/local/WowzaStreamingEngine/conf
  • ./wse/conf.modules -> /usr/local/WowzaStreamingEngine/conf.modules
  • ./wse/content -> /usr/local/WowzaStreamingEngine/content
  • ./wse/transcoder -> /usr/local/WowzaStreamingEngine/transcoder
  • ./wse/logs -> /usr/local/WowzaStreamingEngine/logs
  • ./wse/vif-vod-jobs -> /usr/local/WowzaStreamingEngine/vif-vod-jobs
  • ./vis/models -> /build/models

What this enables:

  • Persistent WSE config and runtime files across container recreation.
  • Editing WSE config directly in your repo and seeing those changes in the running container.
  • Keeping WSE logs on the host for troubleshooting and historical inspection.
  • Keeping transcoder templates/content under source control (or local backup) instead of only inside container storage.
  • Keeping VOD job records, results, and thumbnails across container recreation; without this mount they are lost when the container is replaced.
  • Persisting VIS model files, custom model weights/checkpoints, downloaded checkpoints, and generated TensorRT engines across restarts.

This persistence makes testing and iteration easier, but after major WSE, VIS, model, or plugin changes you may need to remove outdated persisted files before retesting:

  • ./wse/* can preserve older runtime or config state that masks image changes.
  • ./vis/models/engines can preserve stale TensorRT engines built for an older runtime, model set, or GPU environment.
  • Take care with ./vis/models/: it may also contain custom model weights (.pth) that you want to keep.

If you disable these mounts, WSE and VIS fall back to container filesystem defaults and local changes are not preserved when containers are recreated.

Testing Object Detection and VLM Analysis

In the default examples included in this repository, a stream with a name matching object.* triggers object detection using an RF-DETR Medium model capable of identifying up to 80 COCO object categories, including people, vehicles, animals, and more. VLM analysis covers open-vocabulary reasoning and scene-understanding style interpretation for overall activity.

Scene and VLM models are continuously evolving. Expect accuracy and behavior improvements over time as model versions are updated.

Out of the box, VIF renders bounding box overlays on the transcoded -vi stream and injects ID3 metadata into the HLS output. Webhook delivery to third-party platforms and local log file events are available. Developers can further customize event handling using Wowza Streaming Engine modules.

Note

  • Streams automatically appear in the VIF dashboard in Wowza Streaming Engine Manager.
  • Streams can be managed via REST or configured using the VIF configuration page in Wowza Streaming Engine Manager.
  1. Install FFmpeg.
  2. Start WSE using the quick-start steps above.
  3. Use the sample clips in ./videos/, or publish your own files. For best results, keep source videos under 720p resolution.
  4. For object detection, publish a stream to the WSE live application. The default configuration matches stream names using the object.* regex — any name starting with object will trigger object detection:
ffmpeg -stream_loop -1 -re -i "./videos/vi-object-detection-landscape.mp4" -r 25 -g 50 -c:v libx264 -preset veryfast -b:v 2000k -c:a aac -b:a 128k -f flv "rtmp://localhost/live/object_mystream1"
  1. For scene-understanding workflows, publish a stream matching the default scene.* rule. If the scene.* rule is enabled in the VIF configuration page or REST API, publish a stream to the WSE live application:
ffmpeg -stream_loop -1 -re -i "./videos/vi-scene-detection.mp4" -r 25 -g 50 -c:v libx264 -preset veryfast -b:v 2000k -c:a aac -b:a 128k -f flv "rtmp://localhost/live/scene_mystream1"
  1. For VLM analysis, start the full stack with the VLM sidecar and publish a stream matching the default vlm.* rule:
docker compose --profile default --profile vlm up -d
ffmpeg -stream_loop -1 -re -i "./videos/vi-object-detection-landscape.mp4" -r 25 -g 50 -c:v libx264 -preset veryfast -b:v 2000k -c:a aac -b:a 128k -f flv "rtmp://localhost/live/vlm_mystream1"
  1. Expected output for streams analyzed by VIF:

    • The -vi rendition includes overlays. For example, publish object_mystream1, then play http://localhost/live/object_mystream1-vi/playlist.m3u8 in an HLS player.
    • VIF live dashboard: View all incoming live VIF streams at http://localhost:8088/Home.htm#plugin/server/vif/shm.html.
    • VIF VOD dashboard: View all processed video on-demand VIF jobs at http://localhost:8088/Home.htm#plugin/server/vif/vod.html.
    • Live playback with overlays and ID3: Open the built-in player at http://localhost:8088/Home.htm#plugin/server/vif/playback.html to see overlays and ID3 metadata together. Both overlays and ID3 must be enabled in the stream configuration.
    • ID3 metadata is injected into HLS output for analyzed streams.
    • If the LogFiles listener is enabled, events are written to wowzastreamingengine_vi.log (under ./wse/logs/ when WSE log mounts are enabled).

See docs/VLM_GUIDE.md for VLM modes (Detect, Describe, and Custom), endpoint settings, and GPU tuning.

Analyzing Video Files (VOD)

VIF analyzes video files as well as live streams. A VOD job points any of the detectors — object, scene, VLM, or synthetic — at a file under the Engine content directory (./wse/content), runs it over the whole file as fast as the analysis backend allows, and leaves behind a queryable record: every detection stamped with its position on the video timeline, a provenance manifest describing what was run, and a thumbnail. Progress is reported in media time rather than wall-clock, because a file is analyzed as fast as the service answers — a 10-minute file can finish in well under a minute.

Jobs run in the background, survive Engine restarts, recover automatically from transient failures, and can push status updates to your own service by webhook, so a single POST is enough to fire and forget. They reuse the same configuration model as live streams: the same detector types, models, thresholds, class names, and listeners, either by naming a saved stream group config or by passing a config inline. Listeners that write into an output stream (overlay rendering, ID3 injection) do not apply to a file and are skipped; log and webhook listeners work exactly as they do for live streams.

Everything is driven over the WSE REST API under /v2/vif/vod, and the Manager submits and reports on jobs from its own VOD pages, including playback of the analyzed file seeked from the detection strip.

See docs/VOD_GUIDE.md for the walkthrough: accepted inputs and codecs, layering stream group configs with inline configs, reading and downloading results, lifecycle webhooks and named webhook secrets, retention and upload settings, the full API reference, and troubleshooting.

Default Model Coverage and Configuration

By default, VIF uses RF-DETR models for object detection (Nano, Small, Medium, Large). The default stream configuration uses RF-DETR Medium and supports up to 80 COCO classes.

The objects to detect and scene descriptions to analyze can be configured in the VIF configuration page or through the WSE REST API, allowing you to tailor detection behavior, routing, and analysis workflows to your deployment.

VIF also supports importing custom RF-DETR-based object detection models, enabling domain-specific detection beyond the default COCO classes.

To test a custom model, copy your model weights/checkpoint file (.pth) into ./vis/models/, set checkpoint_path for the matching stream in the VIF configuration page or REST API, and restart VIS so the service loads the new model.

Updating VIF Configuration

You can update VIF defaults and per-stream settings in two ways:

  • WSE Manager UI: Open the VIF configuration page to add or update stream rules (for example object.*, scene.*, vlm.*, and synthetic.*), including event listeners, class names, thresholds, and service connection settings.
  • REST API: Use the WSE REST API for scripted updates to defaults and per-stream overrides.

Note

You can use any stream naming terms or regex patterns that fit your workflow; these rule names are user-defined and simply control how traffic is matched to each detector type.

VIF configuration files are stored under conf.modules/vif/:

  • Default.json contains global defaults.
  • Per-stream files (for example live_objectDotStar.json) override those defaults.

For the full field reference and additional API examples, see README.wse-plugin.md.

Plugin Configuration Reference

Use README.wse-plugin.md as the detailed configuration guide for the Wowza Streaming Engine (WSE) Video Intelligence plugin.

  • Includes field-level options and behavior for VIF settings exposed through the Manager UI, REST API, and the underlying conf.modules/vif configuration files.
  • Useful when tuning object detection settings (for example model_name, checkpoint_path, thresholds, and per-stream overrides).

Synthetic Video Detection (Optional)

VIF can flag synthetic / AI-generated video on a live stream via the optional detector_type: "synthetic" analyzer, backed by the NVIDIA Synthetic Video Detector (SVD) NIM. It is opt-in and bring-your-own endpoint: bring up the bundled NIM sidecar with docker compose --profile default --profile svd up -d, or point a stream at a hosted/self-hosted SVD endpoint. The NIM image is access-gated through NVIDIA NGC (VI-550 partnership) and needs an NVENC/NVDEC GPU (T4/A10/A16/A40/L4/L40/L40S/RTX 4090/5090/RTX PRO 6000 Blackwell — not A100/H100/B100).

See docs/SYNTHETIC_VIDEO_DETECTOR.md for the full deployment, GPU support matrix, air-gapped pre-seed, stream-config, and EU AI Act context.

Deploying Your Own VIF Service (Optional)

If you want your own Video Intelligence Service deployment (for example on NVIDIA Brev), see docs/VIS_DEPLOYMENT.md.

Compute Requirements (Self-Hosted VIF)

Recommended baseline for self-hosted VIF: approximately 8 concurrent 720p streams analyzed at up to 10 FPS under standard production conditions using an RF-DETR medium model, assuming GPU-backed inference and typical overlay/event generation settings.

  • 1x NVIDIA T4 (16 GB) minimum
  • 1x NVIDIA L4 (24 GB) recommended
  • 8 vCPU minimum
  • 16 vCPU recommended
  • 32 GB RAM
  • 32 GB disk storage minimum
  • NVMe storage recommended
  • 10 GbE networking preferred for multi-stream deployments

Actual capacity depends on the model variant, frame preprocessing, overlay configuration, event listeners, and downstream integrations. Higher density workloads, heavier models, or full-frame-rate analysis may require additional resources.

Running VIS Only (Optional)

docker compose --profile vi-service up

Use this profile only when you intentionally want the VIS container by itself (for example, debugging or integration with an external WSE instance).

Tip

docker compose down may be required between version changes.

Running on NVIDIA Jetson (Optional)

NVIDIA Jetson Orin devices (Orin Nano and AGX Orin) are supported with Jetpack 7.2. The repository ships with a small overlay, docker-compose.jetson.yaml, that merges on top of the base compose file and swaps the image tags that differ.

Note

For the Jetson Orin Nano, you will need the 8GB model (Orin Nano ships in 4GB/8GB). Run the VIS container alone first to build the initial models (~15–20 min).

docker compose -f docker-compose.yaml -f docker-compose.jetson.yaml --profile vi-service up

There are two ways to run with Jetson support:

docker compose -f docker-compose.yaml -f docker-compose.jetson.yaml up

Or set COMPOSE_FILE once in your .env so a bare docker compose up (and down, logs, etc.) picks up the overlay automatically:

# in .env
COMPOSE_FILE=docker-compose.yaml:docker-compose.jetson.yaml

About

Wowza Streaming Engine's Video Intelligence Framework

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages