TL-DW turns a local meeting recording into:
- a normalized, timestamped speech transcript;
- a small set of meaningful screen frames with optional local OCR;
- ordered multimodal sections in
analysis.json; - a readable transcript document;
- an optional evidence-based meeting report generated by any image-capable model configured in Pi.
Speech, frame extraction, and OCR run locally. Only the bounded transcript sections and selected images are sent to a model when you explicitly run the Pi analysis command.
The Python CLI:
- probes the recording with FFprobe;
- extracts mono 16 kHz audio;
- applies the merged speech-focused normalization pipeline by default;
- transcribes with
faster-whisper(default:large-v3, beam size 5, VAD enabled); - scans frames cheaply at a low rate, scores visual changes, keeps one anchor plus meaningful changes per time section, and removes the rest;
- optionally enriches retained frames with RapidOCR in an isolated worker process;
- writes a portable
analysis.jsonthat aligns speech and images chronologically.
The default audio filter is:
highpass=f=80,dynaudnorm=f=250:g=15:p=0.95,loudnorm=I=-16:TP=-1.5:LRA=11
Audio/video speed-up is deliberately not applied. It can alter ASR accuracy and would require timestamp rescaling; local ASR plus visual deduplication saves model cost without that quality risk.
The project extension at .pi/extensions/tl-dw.ts registers /tl-dw-analyze. It:
- lets you choose an authenticated image-capable model;
- analyzes sections sequentially, carrying a compact evidence state forward;
- sends at most the selected frames for the current section;
- checkpoints each section so interrupted work can resume without repaying completed calls;
- performs a final text-only synthesis;
- writes the report as
<summary>.md, named from its subject, plus usage metadata.
- Python 3.13+
- uv
ffmpegandffprobeonPATH- Pi 0.99.2 or newer for the included extension
- enough disk/RAM for the chosen Whisper model
On macOS:
brew install ffmpeg uvuv syncOptional NVIDIA runtime libraries:
uv sync --extra cuda
export LD_LIBRARY_PATH="$(uv run python -c 'import os; import nvidia.cublas.lib; import nvidia.cudnn.lib; print(os.path.dirname(nvidia.cublas.lib.__file__) + ":" + os.path.dirname(nvidia.cudnn.lib.__file__))')"Optional wtpsplit SaT paragraph segmentation (the lightweight built-in segmenter is the default):
uv sync --extra satCUDA still requires a compatible NVIDIA driver. Current faster-whisper/CTranslate2 releases expect CUDA 12 cuBLAS and cuDNN 9. Use --hardware-profile cuda-pascal for GTX 10-series cards and cuda-modern for newer GPUs.
High-quality Italian example:
uv run tl-dw \
--video "ScreenRec-2026-09-30-16.04.32.mp4" \
--whisper-language it \
--hardware-profile macbook \
--segment-unchaptered \
--timestamp-paragraphs \
--add-table-of-contentsThe equivalent portable entrypoint is:
uv run python main.py --video "recording.mp4"Process every supported video in videos/:
uv run tl-dwForce offline model loading after Whisper is cached:
uv run tl-dw --video "recording.mp4" --local-files-onlySpeech:
--whisper-model: model name/path; defaultlarge-v3.--whisper-language: language code such asitoren; empty enables detection.--hardware-profile:auto,cpu,macbook,cuda-pascal, orcuda-modern.--whisper-device,--whisper-compute-type: low-level overrides.--beam-size: default5.--no-vad: disable voice-activity filtering.--normalize-audio/--no-normalize-audio: enabled by default.--normalization-filter: custom FFmpeg filter.--initial-prompt: vocabulary/style hint for Whisper. Embedded titles and chapter names are used when this is omitted. A title that only repeats an opaque recording filename is ignored.--local-files-only: prohibit model downloads.
Visual analysis:
--frame-sample-sec: visual scan interval; default5seconds.--section-seconds: analysis section duration; default120seconds.--max-frames-per-section: default4.--frame-min-change: minimum novelty score; default0.015.--frame-max-dimension: retained image width limit; default1600.--no-ocr: keep images but skip local OCR.--no-frames: create a transcript-only bundle.
Document formatting:
--segment-unchaptered--timestamp-paragraphs--add-table-of-contents--sat-model simple(default) or a SaT model installed through thesatextra
Run uv run tl-dw --help for the full list.
Start Pi in this repository so it discovers the project extension (review it and grant project trust when prompted):
piThen run:
/tl-dw-analyze output/recording
You can also pass analysis.json directly. Pi opens a picker containing configured image-capable models.
Choose a model explicitly:
/tl-dw-analyze output/recording --model openai-codex/gpt-5.6-sol
Re-running the same bundle/model reuses a completed report or section checkpoints. To intentionally pay for a fresh analysis:
/tl-dw-analyze output/recording --model provider/model --force
For development or use outside the repository's automatic extension discovery:
pi --extension .pi/extensions/tl-dw.tsBy default, artifacts are stored in output/<video-slug>/:
| File | Purpose |
|---|---|
metadata.json |
FFprobe metadata |
audio.wav |
mono 16 kHz ASR input, normalized by default |
transcript.txt |
complete readable timestamped ASR |
transcript.jsonl |
structured ASR segments |
transcript_info.json |
model/runtime/preprocessing settings |
frames/ |
selected images only |
frames.json |
frame timestamps, novelty scores, reasons, and OCR |
ocr.jsonl |
raw and filtered OCR observations |
visual_notes.json |
deduplicated OCR notes for the transcript document |
analysis.json |
ordered Pi-ready speech/image sections |
chapters.json |
rendered paragraph/chapter structure |
<summary>.md |
readable full transcript plus OCR notes and the selected screenshots they describe; the filename and heading are a short summary of the transcript |
pi-analysis/section-*.json |
resumable Pi section checkpoints |
pi-analysis/run.json |
selected model and aggregate token/cost usage |
<summary>.md |
final multimodal meeting report; the filename is a short subject summary, and useful screenshots are embedded with relative links |
large-v3is the quality-first ASR default. Usesmallormediumwhen local runtime matters more.- A two-minute section with at most four selected frames is the default compromise between visual recall and model cost.
- Every section retains an anchor even when the screen is static; extra images must pass a visual-change threshold.
- OCR supplements images but does not replace them. The Pi model receives both because desktop OCR can be noisy.
- Final synthesis is text-only; images are paid for only during their relevant section.
pi-analysis/run.jsontotals successful calls. Provider dashboards remain authoritative for failed calls whose usage may not be returned.- Section checkpoints are model-, extension-, and bundle-specific, preventing stale results from being silently reused.
- The transcript does not perform speaker diarization. The Pi prompt explicitly avoids inventing speaker identities.
- OCR quality depends on font size and resolution; the vision model should treat OCR as fallible.
- Screen changes shorter than the frame scan interval can be missed. Lower
--frame-sample-secfor rapidly changing demos. - Meeting analysis sends selected images and transcript text to the chosen model provider; review privacy requirements first.
uv run --extra dev pytest