Service Readiness and Boot Order: A Started Process Is Not a Usable Environment

The moment an environment supervises a second service, "the process started" and "the environment works" stop being the same fact. A database accepts connections seconds after its process appears in the process table. A dev server compiles for longer than that. A queue worker attaches only after the broker it depends on is up. An agent released at process-start runs its first command into one of those gaps, and the result is a verification that means nothing: tests fail because the database was not listening yet, or pass because the component they exercise never came up at all.

Multi-service is the normal case. Ramp's Inspect bundles its full product stack, Vite, Postgres, Redis, Temporal, and RabbitMQ, into every sandbox (Modal case study, Feb 19, 2026). Versioned images decide what is installed in an environment like that; three runtime mechanisms decide when it is safe to use: a health check that defines ready per service, a dependency graph that orders startup on health rather than on launch, and a small set of idempotent commands that re-run on every boot to absorb drift the snapshot cannot.

Ready is a passing probe, not a running process

A health check gives each service a machine-checkable definition of ready. Two probe shapes cover nearly everything: a command that must exit 0 (pg_isready, a curl against a status endpoint) or a URL probe, where an http check expects a 2xx and a tcp check expects the connection to be accepted. Four knobs make it fit real services: the interval between probes, a total timeout, a retry budget, and a grace period before the first probe so a slow starter is not declared dead while it warms up.

Without a probe, the only observable fact about a service is "has not exited yet," which is equally true of a misconfigured service in the window before it crashes.

The distinction is settled infrastructure knowledge. Kubernetes separates liveness (restart it) from readiness (send it traffic) and withholds traffic until the readiness probe passes (Kubernetes docs). Docker Compose's plain depends_on waits only for the dependency's container to start, and condition: service_healthy is the opt-in that waits for its health check (Compose docs). An agent platform inherits the problem with a different consumer of the signal: an orchestrator deciding whether to release an agent onto the stack rather than a load balancer deciding whether to route a request.

Dependents start on their dependencies' health, in graph order

Each service declares what it depends on, and the supervisor resolves the declarations into a dependency graph via topological sort. Services launch concurrently where the graph allows, but a service with dependencies starts only after every one of them has passed its health check.

Starting the web app after the database process exists reintroduces the exact race the health check closed. The graph should be validated when the environment definition is parsed, not discovered at boot: a dependency cycle or a reference to an undeclared service is a definition error that should fail the build loudly, not a boot that hangs forever waiting on itself.

Launching services in list order approximates the graph well enough to pass local testing. It holds until the day the database restore is slow or the broker recovers a large journal, and then every dependent service starts against a dependency that is not there.

Every-boot hooks catch what the snapshot cannot

Snapshot economics split the work into two piles. Heavy installation runs once, at image build time, and gets baked into the snapshot. But a snapshot is a photograph of a moment, and the code keeps moving after the shutter: new migrations land, dependency lockfiles change, caches go stale. Ramp bounds this drift at the image layer by rebuilding snapshots every 30 minutes, so a fresh sandbox is at most 30 minutes out of date (Modal case study).

A boot hook is a named, lightweight command that runs after all services are healthy, on every boot, fresh build and snapshot restore alike: apply pending database migrations, clear a cache, refresh dependencies against the lockfile. A migration cannot run before the database is ready, which is why the hooks sit behind the health-check gate rather than beside it.

Because hooks run every boot, they must be idempotent. A migration runner that applies only what is pending qualifies; a seed script that inserts rows unconditionally does not, and it will corrupt state on the second boot. The same discipline that makes retries and reconciliation safe at the orchestration layer applies here at the environment layer.

One aggregate readiness signal gates the agent

All of this machinery exists to produce a single fact the orchestrator can wait on: every declared service is healthy and the environment is usable. The agent's first action, or its first turn in a design where the reasoning process boots separately per turn, waits behind that signal.

The stakes are the verifiable-outcomes premise. A test run is evidence only if the stack under test was actually up. Released early, the agent hits one of two failure modes: the loud one, where tests fail against a database that is not listening and the agent burns the session debugging a phantom bug in correct code, and the silent one, where an assertion passes vacuously because the worker it exercises never attached.

Per-task environments raise the frequency. When every task boots its own environment, a readiness race that bit a laptop once a month runs hundreds of times a day. A service that exhausts its health-check retries should fail the boot and name the service, the probe, and the error, feeding the same diagnosis loop as a failed environment build.

The primitives are published; the agent-facing gate mostly is not

Health checks, ordered startup, and post-start hooks are documented exhaustively by Compose and Kubernetes, and any team can assemble them. What the published agent-platform record does not yet document is the composition: gating an agent's release on aggregate stack readiness. Ramp is the closest public datapoint, and its write-up names the stack contents and the snapshot cadence, not the readiness ordering inside a sandbox. No published team has put figures on what early release costs in wasted sessions, so weigh the prescription as reasoning from mechanism rather than as a measured result.