Repository navigation
Fix UVQ 1.5 frame selection and speed up CPU inference - #31
Conversation
|
@andreas-pastor Would appreciate your feedback here! |
|
Hi @slhck, thanks for the PR! The core improvements look great! A few points to address before merging:
Happy to take over if not enough time on your side. Thanks! |
f7eea44 to
c80e147
Compare
|
Thanks for the quick review! I updated the branch, please check once more. Sorry for the import sorting and white space, it's auto-formatting and OCD at work. |
|
Hi @slhck, thanks for addressing the previous feedback so quickly! Three small final points before we merge:
(and update AssertionError to RuntimeError in the docstring). I cannot do it myself as it seems you did not allow "Allow edits by maintainers" on the pull request or something similar? Thanks again! |
c80e147 to
f0bdd3d
Compare
That is not a safe replacement. I updated the README and assertion in the meantime. |
|
In that case, maybe use or a one liner like |
|
Sure, allowing edits by maintainers is on so you should be able to merge this as you see fit! |
Replace shlex.join with subprocess.list2cmdline to avoid the extra shlex import while preserving quoted formatting for args with spaces.
|
Thanks a lot for the PR and the quick iterations, Werner! The speedups and frame-selection fixes are great improvements for UVQ 1.5! |
|
Absolutely, thanks for the quick merge! |
Original issues
Frame selection
With the default
--fps 1, UVQ should score one frame per second. On a 60 fps test clip, it reported source frame indices0, 60, 120, …, 540, but it actually scored frames0, 1, 53, 113, …, 473. The score for a reported frame could therefore belong to a different frame.The reason was that on
main, the reader resized frames and then used FFmpeg's output-r 1, while the model calculated reported indices from the requested rate.Those two paths did not select and report the same frames.
Note: Correcting the selection can change UVQ scores because different frames are evaluated!
Speed
On CPU, inference was quite slow, and memory use was extremely high. We propose optimizations to the tensors and video pipeline to improve speed and memory pressure.
What this branch changes
main-rto reduce the frame rate.fpsfilter before resizing, usingstart_time=0:round=up; pass those frames through without a second frame-rate conversion. On the test clip, UVQ now scores and reports0, 60, 120, …, 540.float32. Keep the existing padding behavior for short decodes. Massively improves memory usage.--batch_sizeand--threadsfor tuning. Use channels-last model weights on CPU. CUDA keeps its previous layout and batch default. Leads to speedup.UVQ 1.0 behavior is not changed.
Measured results
We tested on locally generated 10-second H.264 clips. Tests used six physical cores of a Ryzen 9 9955HX.
main→ reader, batching, and CPU defaultsWe used two different test clips for the two comparisons, so the speedups cannot be combined.
For a 60-second 1080p input, peak memory use fell from about 18 GiB to 1.5 GiB.
Results on smaller CPUs have not yet been measured.
Validation
0, 60, …, 540.git diff --checkpass.Acknowledgements: Implementation and testing was done with the help of GPT-5.6 Sol, the summary of performance notes on this PR and its comparison against
mainwere created by GPT-6 Sol.