Repository navigation
Conversation
play_sound() set its playbin to PLAYING and never watched the bus, so after playback finished the last playbin idled in PLAYING with the file loaded indefinitely. Its pipeline threads were never released (pollen-robotics#840), and on hardware the idle pipeline can spontaneously replay the loaded file — observed on a Reachy Mini Lite (daemon 1.9.0) as a byte-identical replay of wake_up.wav every 12,198 s with no log trace. Add a bus signal watch per playbin: on EOS or ERROR the playbin is set to NULL, its watch removed, and the stored reference cleared. The superseded-playbin path in play_sound() and stop_playing() now share the same teardown helper, so their bus watches are dropped as well. Fixes pollen-robotics#840 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The daemon-side play_sound (GstMediaServer, used for wake-up/sleep and /api/media/play_sound) has the same missing-EOS-teardown bug as the GStreamerAudio twin fixed in the previous commit — and on Reachy Mini Lite it is the implementation that actually runs. A playbin left in PLAYING after EOS was observed replaying its loaded file every ~12198s with no log trace (11 recorded occurrences over two days; replay content tracks the most recently played file). Applying the same bus-watch teardown here resolved it: the periodic replay window passed silent after the fix. Routes stop_sound and the stale-playbin swap in play_sound through the shared _teardown_playbin helper, and parametrizes the teardown unit tests over both implementations. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Flagging an overlap: #1332 prerolls Context that might help either way, since the failure mode is easy to miss: a leaked playbin doesn't just leak a thread. It stays in I caught eleven of these over three nights with a capture tap on the sink (nothing appears in the journal or the stream graph — only the wire hears it), then predicted five in advance to the second before patching. Correlation between replays is 1.000000; max sample difference is zero. Worth noting Related: #840 described the same leak and was closed for inactivity. Tests are in |
Summary
play_sound()creates aplaybin, sets it toPLAYING, and never watches its bus. After playback finishes, the last playbin idles inPLAYINGwith the file loaded indefinitely — until the nextplay_sound()call, which may never come.This bug exists in both
play_soundimplementations, and this PR fixes both:GStreamerAudio.play_sound()(media/audio_gstreamer.py) — first commitGstMediaServer.play_sound()(media/media_server.py) — second commit; this is the implementation the daemon actually uses for wake-up/sleep sounds andPOST /api/media/play_sound(at least on Reachy Mini Lite, daemon 1.9.0)This is the remaining half of #840 (closed for inactivity, but still present on
main). #915 fixed the accumulation between plays by storing the playbin and tearing down the previous one, but nothing tears down the last one, so its pipeline threads are never released and the pipeline stays live.The surprising symptom: spontaneous replay
Beyond the thread leak documented in #840, an idle
PLAYINGplaybin spontaneously replays its loaded file every 12,198 s (3 h 23 m 18 s). Observed on a Reachy Mini Lite, daemon 1.9.0 (Linux host, PipeWire), over a two-day tap of the speaker sink:wake_up.wavplays, the replay is the wake sound; we then playeddance1.wavand predicted (in advance) that the next phantom would be the dance jingle instead — it was, to the second, matchingdance1.wavat r = 0.999. Eachplay_sound()tears down the previous playbin (the 892 stop playing has no effect on audio started by play sound gstreamer backend #915 behavior), which also cancels its pending replay — single-slot semantics that matchself._playbinexactly.object.serial), no API clients, no SSH sessions, no other audio processes.Audio captures of the sink, waveform correlation analysis, and the event timeline are available on request.
Why both files matter
We initially shipped only the
GStreamerAudiofix to the affected robot — and the replays continued unchanged, because the daemon-side path goes throughGstMediaServer.play_sound(), which has the identical bug. Only after patchingmedia_server.pydid the replays stop. A fix to one twin without the other does not cure a deployed robot.Fix
Add a bus signal watch to each playbin created by
play_sound(). OnEOSorERRORthe playbin is set toNULL, the watch removed, and the stored reference cleared (guarded by identity, so a late message from a superseded playbin cannot tear down a newer one). The superseded-playbin path inplay_sound(),stop_playing()(GStreamerAudio) andstop_sound()(GstMediaServer) share the same teardown helper, so their bus watches are removed as well.After the fix (verified on the same robot, both patches applied):
play_sound()behavior unchanged audibly; sounds play normally (wake/sleep chimes, API plays).NULL(logged) and its PipeWire stream closes; the robot idles with zero open playback streams (pre-fix: the zombie stream persisted indefinitely).Tests
tests/unit_tests/test_playbin_teardown.py, in the same fake-selfstyle astest_audio_gstreamer.py, now parametrized over both implementations (GStreamerAudio,GstMediaServer):_playbin_playbinreference intactEOSandERRORbus messages trigger teardown (ERRORalso logs)ruff check/ruff formatclean on the touched files (ruff 0.12.0 as pinned);mypy --stricterror delta zero vsmain.