Version: gallery-server v5.7.0 (Docker), gallery-ml v5.7.0-cuda on a remote ML host (RTX 4090), Postgres 14-vectorchord0.4.3-pgvectors0.2.0, Valkey 9.
Setup: ~84k assets (images + external read-only libraries), 5 users. Pet detection rfdetr-nano (minScore 0.75), pet recognition pet-recognition-base (maxDistance 0.55, minFaces 1).
What happened: I started a full reset (PUT /api/jobs/petRecognition {"command":"start","force":true}). While detection jobs ran, the RSS of immich_server rose roughly linearly, about 0.4 GiB per 1,000 completed jobs. It reached 17.9 GiB of a 19.5 GiB limit and was not released on its own. Restarting the server container freed the memory immediately. The resumed queue then finished normally, with no failed jobs.
Expected: memory stays flat during a long pet detection/recognition run.
Workaround: pause the queue, restart immich_server and resume it whenever memory passes a threshold.
Notes: ML ran in a separate remote container, so the growth is in the server process itself. I haven't tested whether face detection/recognition resets show the same pattern at the same scale. Happy to provide a heap snapshot if you tell me which worker/flag to use.
Version: gallery-server v5.7.0 (Docker), gallery-ml v5.7.0-cuda on a remote ML host (RTX 4090), Postgres
14-vectorchord0.4.3-pgvectors0.2.0, Valkey 9.Setup: ~84k assets (images + external read-only libraries), 5 users. Pet detection
rfdetr-nano(minScore 0.75), pet recognitionpet-recognition-base(maxDistance 0.55, minFaces 1).What happened: I started a full reset (
PUT /api/jobs/petRecognition {"command":"start","force":true}). While detection jobs ran, the RSS ofimmich_serverrose roughly linearly, about 0.4 GiB per 1,000 completed jobs. It reached 17.9 GiB of a 19.5 GiB limit and was not released on its own. Restarting the server container freed the memory immediately. The resumed queue then finished normally, with no failed jobs.Expected: memory stays flat during a long pet detection/recognition run.
Workaround: pause the queue, restart
immich_serverand resume it whenever memory passes a threshold.Notes: ML ran in a separate remote container, so the growth is in the server process itself. I haven't tested whether face detection/recognition resets show the same pattern at the same scale. Happy to provide a heap snapshot if you tell me which worker/flag to use.