Established by experiment¶
Each of these was measured, not reasoned about. A specification cannot derive them, getting them wrong produces failures that are hard to diagnose, and several constrain the architecture rather than the implementation.
- libvips cannot survive
forkonce it has evaluated an image. AfterrequireandVips.concurrency_setthe process has three threads and forks children that work; the first image evaluation takes it to five, and from then on every forked child deadlocks infutex_do_wait, permanently. Reproducible. This is the whole reason for thebefore_forkandbefore_worker_bootsplit, and the reasonbefore_forkmay require and configure but must never evaluate. - Two descriptors over
SCM_RIGHTSwork end to end, withVips::Source.new_from_descriptorandVips::Target.new_to_descriptor, and the kernel enforces the access modes: writing an input or reading an output raisesErrno::EBADF. - Reopening
/proc/self/fd/Ndefeats a read-only descriptor, because it is a freshopenrechecked against the inode and does not inherit the original flags. Never use it to turn a descriptor into a filename. Copy instead. - An empty Docker named volume takes its ownership from the image of whichever container mounts it, even if an earlier container already mounted it, provided it is still empty. So accessory and app boot order does not matter. A bind mount instead takes the host directory's ownership, which is why local development needs the directory created first.
- Kamal 2.11 hard-codes
--network kamalfor app roles. Only accessories acceptnetwork, and only as an accessory key. Underoptions:it is additive rather than overriding — Kamal emits its own--networkfirst — and Docker then refuses the container. Accessories can targetroles: [web, jobs], and are not updated by a deploy. - A worker killed by a resource limit produces a bare end of stream, which is why the supervisor must
hold the connection and report
killed. Amemorybreach does this too, roughly a third of the time: libvips 8.18 dereferences null on its own out-of-memory path and takesSIGSEGVatvips_image_decodeafter printing the correct diagnostic, and GLib's non-nullableg_mallocaborts. So the reap-and-report path carriesmemoryas well asfsize. - A sibling process's
/proc/<pid>/memisEACCESatptrace_scope = 1, but its/proc/<pid>/environis readable. Verified with two same-UID siblings forked from one parent: the environ read returned the victim's canary. Yama restrictsPTRACE_MODE_ATTACH, whichmemneeds, and does not restrictPTRACE_MODE_READ, whichenvironneeds. - A forked process cannot change what its own
/proc/self/environshows. That view is the exec-time environment, soENV.deletein a worker is invisible to a reader. Only anexeced child gets a fresh one, which is whyunsetenv_otherson the tool spawn is the control. - Docker cannot mount
/procwithhidepid.--security-opt proc-opts=hidepid=2is a Podman feature; Docker rejects it outright, and remounting inside the container needsCAP_SYS_ADMIN. - LibreOffice corrupts itself when two instances share a
$HOMEprofile. That is the origin of slots. A hardened conversion measures at roughly 613ms, which sizes a soffice cell's deadline concretely. RLIMIT_ASis unusable andRLIMIT_DATAis expensive. For a real variant whose peak RSS is 45MB,RLIMIT_ASmust be at least 1536MB to succeed reliably and fails nondeterministically for a 400MB band below that;RLIMIT_DATAworks at 704MB.RLIMIT_DATAcharges private writable anonymous mappings and ignoresPROT_NONEreservations, read-only private file mappings, andMAP_SHAREDentirely. It does charge thread stacks, so shrinkingRLIMIT_STACKat container entry is worth real headroom — andRLIMIT_STACKcannot be changed afterexec, because glibc snapshots it at init.- Ruby reserves about 450MB of
RLIMIT_DATAat boot and never touches it. A single ~404MB writable anonymous region, introduced in 3.3 and unchanged in 3.4 and 4.0, unaffected by the GC and malloc environment knobs. It is why thememoryfloor is what it is, and whymemorycannot be read as how much a bomb may consume. - A cgroup memory kill is a prompt, silent
SIGKILLwith no diagnostic, and the kernel chose the allocating worker rather than the supervisor in every trial, because badness is RSS-proportional. The argument for a per-worker limit is the diagnostic, not the choice of victim. - Plain
require "vips"leaves the libheif plugins un-dlopened. They then load lazily inside the worker, after its limits are on, wheredlopenfails with only a warning and the process continues with HEIC and AVIF missing — turning a limit breach intounreadablefor a whole format family.require "image_processing/vips"orVips.block_untrusted truemaps them in the supervisor instead. forkcosts about 2.8ms, measured asfork+exit!+waitfrom a 58MB three-thread parent with libvips required and never evaluated. It is a small part of the cell's fixed overhead, not most of it.- The cell's fixed overhead is copy-on-write settling, and it is proportional to the supervisor's
resident heap. A worker's first pipeline run takes about 7,900 minor faults against 1,564 for a warm
process. Forking the same work from a 265MB parent instead of a 58MB one took faults to about 25,900
and added roughly 52ms per request. So preloading generously in the supervisor makes every request
slower for the life of the deployment.
RLIMIT_ASand libvips thread-pool size were both tested and neither moves it. - macOS has no finite
RLIMIT_DATA.Process.setrlimitrejects one withEINVAL, and the inherited hard limit is already infinity, so it is not a privilege problem. A cell there runs with its memory clamp unenforced and warns once. Every other limit is strict on both platforms. - Re-opening
/dev/fd/Nis checked against the opening process's credentials and the file's mode. Not against the caller's. So a cell handed a descriptor for a mode0600file the application owns cannot open it by name at all, and every operation that gives a tool a filename dies asEACCES. Operations that read the descriptor directly are unaffected, which is why the failure looks selective. The mode is also what enforces invariant 4 once a tool holds a filename: at0400even the owner is refusedO_RDWR, and at0200even the owner is refused a read. That holds only while the cell does not own the file. Changing a mode needs ownership, andcap-drop ALLleaves no capability that overrides it — so a shared group enforces the invariant and a shared uid cannot.