Skip to content

Managed store sizing: make the derivation platform-aware (Windows 487 cap only on Windows, Linux cgroup limits, commit headroom) and write it to one service-owned config file instead of appended blocks #4215

Description

@erikdarlingdata

Problem

The derivation is one formula for every platform. DeriveMemorySettings caps shared_buffers at min(25% RAM, 1 GB). The cap has two stated reasons: co-location/double caching, and the Windows error-487 shared-memory reattach failures (#1559/#1581; upstream BUG #14050/#18954). The 487 mechanism is Windows-only. PostgreSQL on Linux maps shared memory with mmap/SysV and has no reattach step, so a Linux store is capped for a Windows reason. In the other direction, a Linux store in a container must size from the cgroup limit, not host RAM, or shared_buffers + work_mem headroom gets OOM-killed. On Windows the shared segment counts against the commit limit (RAM + pagefile), which nothing checks today.

The config is appended, not owned. Every sizing change is a new marker block appended to postgresql.conf (v1–v14; last occurrence wins). On the three production stores:

Reading the effective value means reading the whole file bottom-up, and an operator can't tell "the service set this" from "someone edited this".

Proposal

1. Platform-aware derivation (pure, testable, same inputs as the host profile in the companion issue):

2. One service-owned file. Put include_dir 'darling.conf.d' (or include_if_exists 'darling-managed.conf') in postgresql.conf once, and have the service rewrite that file wholesale on every start from the current derivation. It's atomic (write to temp, then rename) and has a single header recording the derivation inputs: platform, RAM/limit, CPUs, hypertable count, formula version.

  • The v1–v14 append blocks are migrated once into the file and removed. The migration is a single, idempotent step with a backup of the original conf.
  • Operator overrides keep working: ALTER SYSTEM (postgresql.auto.conf) is read after postgresql.conf and wins, and the host-profile check reports it as operator-override.
  • Restart-only settings (shared_buffers, worker counts) set pending_restart. The check reports it, and the service logs it once instead of re-appending.

Pins:

  • derivation unit tests per platform (Windows cap, Linux cgroup, the commit clamp);
  • a conf-ownership test: two starts with a changed fingerprint leave exactly one managed file and zero appended blocks;
  • a migration test from a conf carrying today's v1–v14 blocks.

Context: measured 2026-09-25 on the three production managed stores (m7i/r7i, 8 vCPU, 31.5 → 63 GiB). A test on one store raised shared_buffers to 8 GB and work_mem to 64 MB. 487 retries stayed at the pre-change rate (~1/h, all LOG-level retries, 0 FATAL) in the first minutes. The 24 h verdict will be added here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    client-siteOwned by the client-site agents (other laptop). Local sessions never pick these up.enhancementNew feature or requestin-progressActively being worked by a local session or its agents (PR open or in flight)

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions