Container¶
This page describes the cell's container: how to build its image, what each container flag does, how to bound OpenMP, and how to check an accessory before and after you deploy it.
A cell is a second container beside the application, on the same host, that shares one volume holding its
UNIX sockets. The examples use Kamal. Any orchestrator that can set the same docker run flags works. The
complete Kamal configuration for both containers is in the README, under
Build and deploy the cell.
For the settings inside the cell, see Cell settings. For how the numbers constrain each other, see Tuning. For where scratch lives, see Scratch. To test an image that you built yourself, see Conformance.
Build an image¶
Hot Cell publishes no base image. bin/rails hotcell:install writes a hotcell/ directory into the
application that holds a complete Dockerfile, the cell's Gemfile, its config.rb, and an
operations/ directory. That Dockerfile is the whole recipe, and you can customize it. Build it from
its own directory:
docker build -t your-image:latest hotcell/
Keep the following in mind:
- Every gem is inside the blast radius. Keep the cell's
Gemfileshort. A larger supervisor heap also costs every request. See Tuning. - Load only the operations that your image has tools for. A cell reports what it loaded in
describe, and a client checks that inventory at boot to catch a cell that doesn't carry an operation that it wants. An operation whose tool is missing passes that check and fails at the first request instead. So requiring an operation without installing its tool is worse than not requiring it. - The cell's image must create its own socket mount point, owned by the cell's user:
/run/hotcell/cell, not only/run/hotcell. Docker creates a missing last level as root, and the cell then can't create a socket in it. The installedDockerfiledoes this. The application's image needs nothing at its own mount point. - Keep the application's and the cell's lockfiles in step. They resolve Hot Cell separately.
HotCell.describe_cellswarns at boot when the cell'shotcell-serverversion differs from the application'shotcell-clientversion.
Container flags¶
These are docker run flags. Under Kamal, they go in an accessory's options:, and Kamal supplies no
value of its own for any of them.
network is the exception, and it's the flag that the design depends on most. It's an accessory key of
its own, a sibling of image: and roles:. Kamal always emits its own --network, kamal by default.
An entry under options: adds a second --network rather than replacing the first, and Docker refuses
the container:
docker: conflicting options: cannot attach both user-defined and non-user-defined network-modes
Give each cell its own socket volume. Two accessories that share one would write work.sock over each
other.
Without Kamal, the cell container needs --volume hotcell-sockets:/run/hotcell/cell and the security flags
below. The application container needs --volume hotcell-sockets:/run/hotcell/active_storage,
--group-add 10001, and HOTCELL_ROOT=/run/hotcell.
Performance flags¶
These flags have no defaults: Docker applies no limit for a flag that you omit, so a cell without them can take the whole host. Tune each one to your workload and your hardware. The examples are from the README's accessory, and they aren't recommendations.
| Flag | Example | Description |
|---|---|---|
cpus |
2 |
The share of the host that this cell can use. Start the cell's concurrency at twice this number, and match the image's OMP_NUM_THREADS to it. See Bound the OpenMP thread pools. |
memory |
2g |
The cgroup limit, which counts every worker and the tmpfs. Size it from concurrency × peak RSS plus the tmpfs. Keep it above the cell's memory. On a disk-backed scratch, there's no tmpfs term. See Scratch. |
memory-swap |
2g |
Set it equal to memory. If you omit it, Docker allows twice memory in swap, and the memory limit no longer holds. |
tmpfs size |
size=512m |
Scratch for all concurrent workers together. Size it from concurrency times the most scratch that one request holds. See Constraints between the numbers. Moving scratch onto disk separates it from memory. See Scratch. |
ulimit: stack |
leave it unset | See Don't lower the stack limit. |
Security flags¶
Use these values. They're what a cell is for. If you omit one, the protection is gone, and the cell still serves requests exactly as before.
| Flag | Recommended | Description |
|---|---|---|
network |
none |
Removes every network interface. A tool that's persuaded to fetch a URL can't reach anything. Set it as an accessory key, not under options:. |
read-only |
true |
Makes the root filesystem read-only. A cell writes only to /tmp and to the socket volume. |
tmpfs flags |
nosuid,nodev,noexec |
noexec stops a dropped binary from running from /tmp, where a request's files live. It doesn't cover the socket volume, and it doesn't stop ruby payload.rb. On a disk-backed scratch, Docker sets none of these flags. See Scratch. |
cap-drop |
ALL |
Removes every Linux capability. |
security-opt |
no-new-privileges:true |
Prevents a setuid binary from regaining what cap-drop removed. |
user |
10001:10001 |
Runs the cell with no home directory and no shell. Without user namespace remapping, this is a host uid in the ordinary range, so pick one that your hosts don't give to a person. |
pids-limit |
512 |
Bounds the number of processes in the cell. It must clear concurrency plus the threads and subprocesses that one toolchain starts, so check it when you raise concurrency. |
Don't lower the stack limit¶
Leave ulimit: stack unset. A container inherits the Docker daemon's value, normally 8MB.
Lowering it to 2MB buys a worker about 24MB more room, and nothing for an operation that shells out. It
costs far more than that: a thread that overflows the smaller stack dies on SIGSEGV, and the cell
reports that as killed with cause crashed, which is transient. The caller's job retries the request
against a limit that fails it again, for as long as the job keeps trying, and nothing in the verdict
points at the setting. Raise the cell's memory instead.
The stack limit is a container flag rather than a config.rb setting because glibc reads it at exec,
before Process.setrlimit could run. bin/conformance and bin/load set no stack limit, and
examples/gate checks that they don't.
Bound the OpenMP thread pools¶
Set OMP_NUM_THREADS and OMP_THREAD_LIMIT in the image. The installed Dockerfile sets both. On a large
host, neither is optional: without them, a cell dies. hotcell:install doesn't change an existing
Dockerfile, so for a cell installed before these variables existed, add both by hand and rebuild the
image.
OpenMP sizes its thread pool from the host's core count. A cpus: limit is a CFS quota, not an affinity
mask, so it doesn't lower that count: on a 98-core host, libvips and ImageMagick ask for 98 threads. Each
thread stack is 8MB of private anonymous memory, which the cell's memory limit charges as
RLIMIT_DATA. Under a 1280MB limit, 96 of those stacks fit and 104 don't. Past that line,
pthread_create returns EAGAIN, which libgomp treats as unrecoverable. It writes the following line to
standard error and calls exit(1):
libgomp: Thread creation failed: Resource temporarily unavailable
The worker dies before it answers, so the caller gets a transient failure, and the job retries against a host that fails the same way. This killed 285 Basecamp workers in production on 2026-08-31 and forced a rollback.
That line now reaches a log. The supervisor captures what a worker writes to fd 2 and attaches its tail to
the worker.killed event, as hotcell.stderr, and to the failure that the caller receives. See
What a worker wrote to fd 2.
Size OMP_NUM_THREADS from the container's cpus. The number follows the allocation, not the example. It's
a thread count, so round a fractional quota down, and round a value below 1 up to 1. OMP_THREAD_LIMIT is
the backstop against a library that raises the count itself by calling omp_set_num_threads, which is
what ImageMagick does for MAGICK_THREAD_LIMIT.
The cell forwards both variables to the tools that it runs. A tool sees the environment that its
operation wrote for it, not the worker's own, which is invariant 9. So the image's
variables alone would bound in-process libvips and nothing else. Operation#run_tool and mini_magick both
pass the pair from the cell's environment.
A deploy to staging or beta can't catch a regression here, because the failure exists only at
production's core count. That's how it reached production. So the guard is a test:
hotcell-client/test/install_test.rb checks the installed Dockerfile's variables, and an image that you
customize needs its own test.
To verify a built image, write a GIF inside it and count /proc/self/task during the write. GIF is the
cheapest probe: it's the only common output format that quantizes, and quantization is where the threads
appear. Unbounded, the count tracks the visible cores. Bounded, it stays at the limit.
Verify an accessory¶
bin/conformance supplies its own Docker flags, so it proves that the image works when the flags are
right. It doesn't read your Kamal configuration, and it passes against an image that you're about to
deploy with cap-drop missing.
Before you deploy¶
Print the command that Kamal runs. kamal config doesn't show it: it prints the merged configuration, not
the command, so a --network emitted twice doesn't appear there at all.
bundle exec ruby -rkamal -e '
config = Kamal::Configuration.create_from(
config_file: Pathname.new("config/deploy.yml"), destination: "production", version: "check")
puts Kamal::Commands::Accessory.new(config, name: :images).run.flatten.join(" ")
'
--network must appear once on that line, as --network none. Two of them is the options: mistake
described under Container flags, and Docker rejects the container at boot with exit
status 125.
After you deploy¶
Read the flags on the running container:
docker inspect <container> --format '{{json .HostConfig}}' | jq '{
NetworkMode, ReadonlyRootfs, CapDrop, SecurityOpt, PidsLimit, Memory, MemorySwap, Tmpfs, Binds
}'