Skip to content

Fix UVQ 1.5 frame selection and speed up CPU inference - #31

Merged
andreas-pastor merged 2 commits into
google:mainfrom
slhck:optimize-cpu-inference
Sep 29, 2026
Merged

andreas-pastor merged 2 commits into
google:mainfrom
slhck:optimize-cpu-inference

Conversation

@slhck

@slhck slhck commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Original issues

Frame selection

With the default --fps 1, UVQ should score one frame per second. On a 60 fps test clip, it reported source frame indices 0, 60, 120, …, 540, but it actually scored frames 0, 1, 53, 113, …, 473. The score for a reported frame could therefore belong to a different frame.

The reason was that on main, the reader resized frames and then used FFmpeg's output -r 1, while the model calculated reported indices from the requested rate.

Those two paths did not select and report the same frames.

Note: Correcting the selection can change UVQ scores because different frames are evaluated!

Speed

On CPU, inference was quite slow, and memory use was extremely high. We propose optimizations to the tensors and video pipeline to improve speed and memory pressure.

What this branch changes

Area main This branch
Frame selection Resize, then use FFmpeg output -r to reduce the frame rate. Select frames with an fps filter before resizing, using start_time=0:round=up; pass those frames through without a second frame-rate conversion. On the test clip, UVQ now scores and reports 0, 60, 120, …, 540.
Video memory Write the decoded video to a temporary raw file, then load and normalize the full clip in memory. Read FFmpeg output in bounded batches and normalize directly to NumPy float32. Keep the existing padding behavior for short decodes. Massively improves memory usage.
CPU inference Build a tensor for the whole clip, then infer in fixed batches of 24 frames. Infer as batches arrive. Default to one frame per CPU batch and up to four PyTorch CPU threads; expose --batch_size and --threads for tuning. Use channels-last model weights on CPU. CUDA keeps its previous layout and batch default. Leads to speedup.
FFmpeg invocation Run a shell command assembled from paths and options. Pass an argument list directly to FFmpeg, so paths with spaces or shell characters work instead of breaking or causing security issues.

UVQ 1.0 behavior is not changed.

Measured results

We tested on locally generated 10-second H.264 clips. Tests used six physical cores of a Ryzen 9 9955HX.

Comparison 1080p60 2160p60
main → reader, batching, and CPU defaults 7.88 s → 6.77 s 10.39 s → 7.84 s
Previous branch revision → channels-last CPU weights 5.53 s → 3.70 s 6.16 s → 4.23 s

We used two different test clips for the two comparisons, so the speedups cannot be combined.

For a 60-second 1080p input, peak memory use fell from about 18 GiB to 1.5 GiB.

Results on smaller CPUs have not yet been measured.

Validation

  • Three video-reader unit tests pass, covering filter order, paths with shell characters, float32 conversion, padding, and incomplete frames.
  • A full CPU CLI run on a 60 fps clip reported the intended indices 0, 60, …, 540.
  • Python compilation and git diff --check pass.

Acknowledgements: Implementation and testing was done with the help of GPT-5.6 Sol, the summary of performance notes on this PR and its comparison against main were created by GPT-6 Sol.

@slhck

slhck commented Sep 29, 2026

Copy link
Copy Markdown
Contributor Author

@andreas-pastor Would appreciate your feedback here!

@andreas-pastor

Copy link
Copy Markdown
Collaborator

Hi @slhck, thanks for the PR!

The core improvements look great!

A few points to address before merging:

  1. Docstrings: The Args: and Returns: sections were stripped from load_video_1p5, and iter_video_batches_1p5 lacks Args: and Yields: blocks. Please restore/add standard Google-style docstrings.
  2. transpose filter order: In _video_1p5_command, place transpose=2 after fps= (fps={video_fps}:start_time=0:round=up,{transpose_filter}scale=...) to avoid rotating frames that will immediately be dropped and save even more inference time.
  3. batch_size default: In uvq_inference.py, batch_size defaults to 1 on CPU, but in UVQ1p5.infer() it still defaults to 24. Please default batch_size to 1 on CPU in infer() as well for direct library callers.
  4. Diff hygiene: Please revert the unrelated import reordering in uvq_inference.py/uvq1p5.py and whitespace edits in README.md to minimize churn.

Happy to take over if not enough time on your side.

Thanks!

@slhck
slhck force-pushed the optimize-cpu-inference branch from f7eea44 to c80e147 Compare September 29, 2026 16:49
@slhck

slhck commented Sep 29, 2026

Copy link
Copy Markdown
Contributor Author

Thanks for the quick review! I updated the branch, please check once more. Sorry for the import sorting and white space, it's auto-formatting and OCD at work.

@andreas-pastor

Copy link
Copy Markdown
Collaborator

Hi @slhck, thanks for addressing the previous feedback so quickly!

Three small final points before we merge:

  1. README.md: Could you remove the ### Optimized CPU inference section? The CLI flag docs under Options are sufficient, and the architectural details are already well captured in the PR description.
  2. shlex import: In utils/video_reader.py, shlex is only used for shlex.join(command) in one log line. Could you replace that with ' '.join(command) and drop import shlex?
  3. Empty decode error: In utils/video_reader.py line 302, could we switch raise AssertionError("...: 0 < {single_frame_size}") to a RuntimeError? Since 0 is frames and single_frame_size is bytes, a cleaner message would be:
    if num_real_frames < 1:
      raise RuntimeError(f"Decoding failed to output a single frame for '{filepath}'")

(and update AssertionError to RuntimeError in the docstring).

I cannot do it myself as it seems you did not allow "Allow edits by maintainers" on the pull request or something similar?

Thanks again!

@slhck
slhck force-pushed the optimize-cpu-inference branch from c80e147 to f0bdd3d Compare September 29, 2026 17:41
@slhck

slhck commented Sep 29, 2026

Copy link
Copy Markdown
Contributor Author

shlex import: In utils/video_reader.py, shlex is only used for shlex.join(command) in one log line. Could you replace that with ' '.join(command) and drop import shlex?

That is not a safe replacement. shlex ensures the command works fine when individual components contain spaces or special characters. I would prefer to keep it; it's just the idiomatic way of running shell commands from python. Please see: https://docs.python.org/3/library/shlex.html#shlex.join — that's also what #30 tried to address.

I updated the README and assertion in the meantime.

@andreas-pastor

Copy link
Copy Markdown
Collaborator

In that case, maybe use subprocess.list2cmdline(command) since already in used here to avoid shlex import for just a join?

or a one liner like cmd_str = " ".join(f'"{arg}"' if " " in arg else arg for arg in command) ?

@slhck

slhck commented Sep 29, 2026

Copy link
Copy Markdown
Contributor Author

Sure, allowing edits by maintainers is on so you should be able to merge this as you see fit!

Replace shlex.join with subprocess.list2cmdline to avoid the extra
shlex import while preserving quoted formatting for args with spaces.
@andreas-pastor
andreas-pastor merged commit c204ecc into google:main Sep 29, 2026
7 checks passed
@andreas-pastor

Copy link
Copy Markdown
Collaborator

Thanks a lot for the PR and the quick iterations, Werner! The speedups and frame-selection fixes are great improvements for UVQ 1.5!

@slhck

slhck commented Sep 29, 2026

Copy link
Copy Markdown
Contributor Author

Absolutely, thanks for the quick merge!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants