You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Native objects from the cvcuda + PyNvVideoCodec chain (cvcuda.Tensor, cvcuda.ExternalBuffer, PyNvVideoCodec.DecodedFrame) crash the Python interpreter when released in a specific pattern, even after just a single release event. Different surface conditions than the prior reports I link below, but the same error class (Python GC racing with native dealloc).
I want to flag this as a possibly-regressed version of the same bug class the maintainers have fixed before (see #72, #188, #208's "potential race condition with Python garbage collection fixed with Pybind upgrade" note), rather than a fresh discovery -- the prior fixes appear to have only partially stuck. The repro below triggers the crash with a different code path than the prior reports (single-threaded in-scope list reassignment + cvcuda.cvtcolor via DLPack handoff from PyNvVideoCodec, vs. those reports' multithreading + cvcuda.resize directly).
Steps/Code to reproduce bug
# Self-contained repro: open a SimpleDecoder, build a cvcuda chain for# one frame, reassign the list, then try to decode another frame. Crashes# on the second batch's cvcuda call (segfault on the freed native pointer).# Run with any H.264 mp4 as ./clip.mp4; can be generated with# `ffmpeg -f lavfi -i testsrc=duration=30:size=1920x1080:rate=30 -c:v libx264 clip.mp4`importosos.environ["HF_HUB_DISABLE_PROGRESS_BARS"] ="1"importtorchimportPyNvVideoCodecasnvcimportcvcudadecoder=nvc.SimpleDecoder("clip.mp4", use_device_memory=True)
keepalive= []
forbatch_idxinrange(2):
frame=decoder.get_batch_frames(1)[0]
raw=torch.from_dlpack(frame)
h15, w=raw.shapereshaped=raw.view(1, h15, w, 1).contiguous()
cv_in=cvcuda.as_tensor(reshaped, "NHWC")
cv_out=cvcuda.cvtcolor(cv_in, cvcuda.ColorConversion.YUV2RGB_NV12)
ext=cv_out.cuda()
result=torch.from_dlpack(ext)
keepalive.append((raw, reshaped, cv_in, cv_out, ext, result))
# The trigger: a plain in-scope list reassignment, no function boundary.keepalive= []
Crash signature on the second batch's cvcuda call:
Linux + driver 595.84 + Python 3.12: bare SIGSEGV (rc=-11) -- the raw signal is delivered before Python's signal handler runs.
Other platforms / driver versions: Fatal Python error: PyThreadState_Get: the function must be called with the GIL held, but the GIL is released followed by SIGABRT.
I have a self-contained script that handles the subprocess boundary, captures exit codes, and classifies the crash: https://github.com/gavmor/cvcuda-bug-repros/blob/main/multi_batch_repro.py (it's also preserved alongside the other repros in that repo, in case future investigations need additional test cases).
The KEEPALIVE pattern in the repro is the same one I had to use in my own production code to avoid the crash (process-lifetime, never-cleared list, os._exit(0) at the end) -- the chain is real, it works, and the workaround is documented in my own project's feasibility writeup.
Expected behavior
keepalive = [] is a plain Python list reassignment. It should free the previous batch's objects normally; the next batch's decoder.get_batch_frames(1) and subsequent cvcuda calls should run on fresh objects without referencing anything from the prior batch.
Environment overview
Environment location: Bare-metal
Method of cuDF install: not using cuDF; cvcuda installed with pip
Driver: NVIDIA 595.84 (also reproduced on prior drivers per the prior linked issues)
Same bug present in every version combination I tested: cvcuda 0.15.0, 0.16.0, 0.17.0 × PyNvVideoCodec 2.0.0, 2.0.5, 2.1.0, 2.2.2. The bug predates all publicly available versions on PyPI.
PyNvVideoCodec 2.2.2 / SimpleDecoder reproduces; ThreadedDecoder also reproduces (same crash, same trigger). cvcuda.cvtcolor is the specific operator that crashes on the second batch; the same chain without cvtcolor does not crash.
[BUG] Runtime Error when do Multithreading #72 -- same error class (pybind11_object_dealloc: unregistered instance), different trigger (multithreading + cvcuda.resize). Fixed in v0.3.1.
This new report's trigger is different from those (single-threaded, in-scope reassignment, cvcuda.cvtcolor via DLPack from PyNvVideoCodec, no cvcuda.resize involved), but the underlying race and the crash class are clearly the same family. My current best guess: the prior Pybind11 upgrade fixed one specific path but did not address all release paths through the cvcuda Tensor / ExternalBuffer / DLPack chain. Happy to provide a C++-level repro or a debugger trace if helpful.
Additional context
Worth noting: this bug is the blocking issue for a video keyframe-extraction pipeline I'm developing (zero-copy NVDEC -> CV-CUDA -> CLIP). The bug caps the per-process frame count to ~1300 (24 GB ÷ ~18 MB per frame in KEEPALIVE), since every frame's pixel buffers must be retained process-lifetime to avoid the crash. So the workaround is real and works, but the bug genuinely blocks the "process every frame of a large corpus" use case.
Describe the bug
Native objects from the cvcuda + PyNvVideoCodec chain (
cvcuda.Tensor,cvcuda.ExternalBuffer,PyNvVideoCodec.DecodedFrame) crash the Python interpreter when released in a specific pattern, even after just a single release event. Different surface conditions than the prior reports I link below, but the same error class (Python GC racing with native dealloc).I want to flag this as a possibly-regressed version of the same bug class the maintainers have fixed before (see #72, #188, #208's "potential race condition with Python garbage collection fixed with Pybind upgrade" note), rather than a fresh discovery -- the prior fixes appear to have only partially stuck. The repro below triggers the crash with a different code path than the prior reports (single-threaded in-scope list reassignment +
cvcuda.cvtcolorvia DLPack handoff from PyNvVideoCodec, vs. those reports' multithreading +cvcuda.resizedirectly).Steps/Code to reproduce bug
Crash signature on the second batch's cvcuda call:
SIGSEGV(rc=-11) -- the raw signal is delivered before Python's signal handler runs.Fatal Python error: PyThreadState_Get: the function must be called with the GIL held, but the GIL is releasedfollowed bySIGABRT.RuntimeError: pybind11_object_dealloc(): Tried to deallocate unregistered instance!raised as a Python exception. All three are the same underlying race; the different surface forms depend on which code path catches the dealloc first.I have a self-contained script that handles the subprocess boundary, captures exit codes, and classifies the crash: https://github.com/gavmor/cvcuda-bug-repros/blob/main/multi_batch_repro.py (it's also preserved alongside the other repros in that repo, in case future investigations need additional test cases).
The KEEPALIVE pattern in the repro is the same one I had to use in my own production code to avoid the crash (process-lifetime, never-cleared list,
os._exit(0)at the end) -- the chain is real, it works, and the workaround is documented in my own project's feasibility writeup.Expected behavior
keepalive = []is a plain Python list reassignment. It should free the previous batch's objects normally; the next batch'sdecoder.get_batch_frames(1)and subsequent cvcuda calls should run on fresh objects without referencing anything from the prior batch.Environment overview
SimpleDecoderreproduces;ThreadedDecoderalso reproduces (same crash, same trigger).cvcuda.cvtcoloris the specific operator that crashes on the second batch; the same chain withoutcvtcolordoes not crash.Environment details
Related issues
This new report's trigger is different from those (single-threaded, in-scope reassignment,
cvcuda.cvtcolorvia DLPack from PyNvVideoCodec, nocvcuda.resizeinvolved), but the underlying race and the crash class are clearly the same family. My current best guess: the prior Pybind11 upgrade fixed one specific path but did not address all release paths through the cvcuda Tensor / ExternalBuffer / DLPack chain. Happy to provide a C++-level repro or a debugger trace if helpful.Additional context
Worth noting: this bug is the blocking issue for a video keyframe-extraction pipeline I'm developing (zero-copy NVDEC -> CV-CUDA -> CLIP). The bug caps the per-process frame count to ~1300 (24 GB ÷ ~18 MB per frame in KEEPALIVE), since every frame's pixel buffers must be retained process-lifetime to avoid the crash. So the workaround is real and works, but the bug genuinely blocks the "process every frame of a large corpus" use case.