Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

TL-DW - Too Long; Didn't Watch

TL-DW turns a local meeting recording into:

  • a normalized, timestamped speech transcript;
  • a small set of meaningful screen frames with optional local OCR;
  • ordered multimodal sections in analysis.json;
  • a readable transcript document;
  • an optional evidence-based meeting report generated by any image-capable model configured in Pi.

Speech, frame extraction, and OCR run locally. Only the bounded transcript sections and selected images are sent to a model when you explicitly run the Pi analysis command.

How it works

1. Local preparation

The Python CLI:

  1. probes the recording with FFprobe;
  2. extracts mono 16 kHz audio;
  3. applies the merged speech-focused normalization pipeline by default;
  4. transcribes with faster-whisper (default: large-v3, beam size 5, VAD enabled);
  5. scans frames cheaply at a low rate, scores visual changes, keeps one anchor plus meaningful changes per time section, and removes the rest;
  6. optionally enriches retained frames with RapidOCR in an isolated worker process;
  7. writes a portable analysis.json that aligns speech and images chronologically.

The default audio filter is:

highpass=f=80,dynaudnorm=f=250:g=15:p=0.95,loudnorm=I=-16:TP=-1.5:LRA=11

Audio/video speed-up is deliberately not applied. It can alter ASR accuracy and would require timestamp rescaling; local ASR plus visual deduplication saves model cost without that quality risk.

2. Pi analysis

The project extension at .pi/extensions/tl-dw.ts registers /tl-dw-analyze. It:

  • lets you choose an authenticated image-capable model;
  • analyzes sections sequentially, carrying a compact evidence state forward;
  • sends at most the selected frames for the current section;
  • checkpoints each section so interrupted work can resume without repaying completed calls;
  • performs a final text-only synthesis;
  • writes the report as <summary>.md, named from its subject, plus usage metadata.

Requirements

  • Python 3.13+
  • uv
  • ffmpeg and ffprobe on PATH
  • Pi 0.99.2 or newer for the included extension
  • enough disk/RAM for the chosen Whisper model

On macOS:

brew install ffmpeg uv

Install

uv sync

Optional NVIDIA runtime libraries:

uv sync --extra cuda
export LD_LIBRARY_PATH="$(uv run python -c 'import os; import nvidia.cublas.lib; import nvidia.cudnn.lib; print(os.path.dirname(nvidia.cublas.lib.__file__) + ":" + os.path.dirname(nvidia.cudnn.lib.__file__))')"

Optional wtpsplit SaT paragraph segmentation (the lightweight built-in segmenter is the default):

uv sync --extra sat

CUDA still requires a compatible NVIDIA driver. Current faster-whisper/CTranslate2 releases expect CUDA 12 cuBLAS and cuDNN 9. Use --hardware-profile cuda-pascal for GTX 10-series cards and cuda-modern for newer GPUs.

Prepare a meeting

High-quality Italian example:

uv run tl-dw \
  --video "ScreenRec-2026-09-30-16.04.32.mp4" \
  --whisper-language it \
  --hardware-profile macbook \
  --segment-unchaptered \
  --timestamp-paragraphs \
  --add-table-of-contents

The equivalent portable entrypoint is:

uv run python main.py --video "recording.mp4"

Process every supported video in videos/:

uv run tl-dw

Force offline model loading after Whisper is cached:

uv run tl-dw --video "recording.mp4" --local-files-only

Important preparation options

Speech:

  • --whisper-model: model name/path; default large-v3.
  • --whisper-language: language code such as it or en; empty enables detection.
  • --hardware-profile: auto, cpu, macbook, cuda-pascal, or cuda-modern.
  • --whisper-device, --whisper-compute-type: low-level overrides.
  • --beam-size: default 5.
  • --no-vad: disable voice-activity filtering.
  • --normalize-audio / --no-normalize-audio: enabled by default.
  • --normalization-filter: custom FFmpeg filter.
  • --initial-prompt: vocabulary/style hint for Whisper. Embedded titles and chapter names are used when this is omitted. A title that only repeats an opaque recording filename is ignored.
  • --local-files-only: prohibit model downloads.

Visual analysis:

  • --frame-sample-sec: visual scan interval; default 5 seconds.
  • --section-seconds: analysis section duration; default 120 seconds.
  • --max-frames-per-section: default 4.
  • --frame-min-change: minimum novelty score; default 0.015.
  • --frame-max-dimension: retained image width limit; default 1600.
  • --no-ocr: keep images but skip local OCR.
  • --no-frames: create a transcript-only bundle.

Document formatting:

  • --segment-unchaptered
  • --timestamp-paragraphs
  • --add-table-of-contents
  • --sat-model simple (default) or a SaT model installed through the sat extra

Run uv run tl-dw --help for the full list.

Analyze with Pi

Start Pi in this repository so it discovers the project extension (review it and grant project trust when prompted):

pi

Then run:

/tl-dw-analyze output/recording

You can also pass analysis.json directly. Pi opens a picker containing configured image-capable models.

Choose a model explicitly:

/tl-dw-analyze output/recording --model openai-codex/gpt-5.6-sol

Re-running the same bundle/model reuses a completed report or section checkpoints. To intentionally pay for a fresh analysis:

/tl-dw-analyze output/recording --model provider/model --force

For development or use outside the repository's automatic extension discovery:

pi --extension .pi/extensions/tl-dw.ts

Output

By default, artifacts are stored in output/<video-slug>/:

File Purpose
metadata.json FFprobe metadata
audio.wav mono 16 kHz ASR input, normalized by default
transcript.txt complete readable timestamped ASR
transcript.jsonl structured ASR segments
transcript_info.json model/runtime/preprocessing settings
frames/ selected images only
frames.json frame timestamps, novelty scores, reasons, and OCR
ocr.jsonl raw and filtered OCR observations
visual_notes.json deduplicated OCR notes for the transcript document
analysis.json ordered Pi-ready speech/image sections
chapters.json rendered paragraph/chapter structure
<summary>.md readable full transcript plus OCR notes and the selected screenshots they describe; the filename and heading are a short summary of the transcript
pi-analysis/section-*.json resumable Pi section checkpoints
pi-analysis/run.json selected model and aggregate token/cost usage
<summary>.md final multimodal meeting report; the filename is a short subject summary, and useful screenshots are embedded with relative links

Quality and cost choices

  • large-v3 is the quality-first ASR default. Use small or medium when local runtime matters more.
  • A two-minute section with at most four selected frames is the default compromise between visual recall and model cost.
  • Every section retains an anchor even when the screen is static; extra images must pass a visual-change threshold.
  • OCR supplements images but does not replace them. The Pi model receives both because desktop OCR can be noisy.
  • Final synthesis is text-only; images are paid for only during their relevant section.
  • pi-analysis/run.json totals successful calls. Provider dashboards remain authoritative for failed calls whose usage may not be returned.
  • Section checkpoints are model-, extension-, and bundle-specific, preventing stale results from being silently reused.

Current limitations

  • The transcript does not perform speaker diarization. The Pi prompt explicitly avoids inventing speaker identities.
  • OCR quality depends on font size and resolution; the vision model should treat OCR as fallible.
  • Screen changes shorter than the frame scan interval can be missed. Lower --frame-sample-sec for rapidly changing demos.
  • Meeting analysis sends selected images and transcript text to the chosen model provider; review privacy requirements first.

Tests

uv run --extra dev pytest

About

too long; didn't watch

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages