Retries, Idempotency, and Reconciliation: The Reliability Layer Under Provisioning
Renting a managed sandbox buys compute and isolation. Reliable delivery between your own services is a property of the pipeline you wire around it, and that pipeline is typically four independently deployed parties (agent topology maps the shape): an orchestrator that owns tasks and sessions, a fleet service that places VMs on hosts, a host agent that boots them, and an in-VM agent that runs setup and reports readiness by calling home over HTTP. None of these hops is transactional; each is an async HTTP call that can succeed, fail, or succeed invisibly.
Every callback arrives twice, late, or never
- Twice. The in-VM agent POSTs environment-ready, the orchestrator commits the state change, and the 200 is lost on the way back. The agent cannot distinguish "processed, response lost" from "never processed," so it must retry, and duplicates are a certainty.
- Late. Two callbacks about the same resource race through different paths and land in either order, so a reader can act on one fact before the other has been written anywhere.
- Never. A transient outage outlives the sender's retry budget, or the sender crashes mid-retry.
- Never, and no sender. The VM dies before the reporting agent starts: a host crash, a kernel that never boots, a config parse error ahead of process launch. No retry policy fixes this class; only polling finds it.
Dedupe at ingestion: one delivery id, one transaction
The sender mints a unique delivery id per logical event and reuses it on every retry. The receiver records it in a deliveries table with a unique constraint, in the same database transaction as the state change the event causes. A retried delivery that already landed hits the constraint and the handler no-ops: no state written twice, no listeners re-fired.
The same-transaction detail is load-bearing. Check-then-write as two steps leaves a crash window: the process dies after writing state but before recording the delivery, and the retry replays the side effects.
Everything the event triggers sits behind the dedupe: if environment-ready dispatches a repo-preparation job and advances a session, a duplicate must skip all of it, so the dispatch lives inside the guarded handler rather than in a second listener with its own idea of freshness.
The same pattern recurs one level up, where external webhook providers redeliver workflow events on their own schedule: event routing and dedupe.
Bounded retry, sorted by whether retrying can help
Sender policy is short: retry transport errors and 5xx responses, never 4xx. A transport error or a 5xx is plausibly transient, a deploy or a restart or a load spike. A 4xx means the request itself is wrong and will fail identically on every attempt; retrying it just delays the real signal.
Keep the budget small. Three attempts with a one-then-five-second backoff rides out a deploy blip. A durable outbox is the wrong tool inside a machine that is minutes from destruction: the queued state dies with the disk it lives on.
Sort channels by the consequence of a drop. Telemetry (metrics, log shipping) can stay fire-and-forget; losing a data point costs a gap on a dashboard. State-bearing events (ready, failed, turn complete) get the retry budget, because a dropped one strands a state machine: a session waits forever on a readiness signal that already fired.
Reconcilers catch the failure that never reports itself
A reconciler is a sweep job on a timer: query for resources stuck in a transitional state past a deadline, poll the authoritative system for what actually happened to the underlying resource, and force-resolve. Mark the record failed, release anything waiting on it, clean up what the far side left behind. Run it every minute or so, and make the job itself idempotent, since two overlapping sweeps are just delivery-twice again.
This is the only mechanism that covers the no-sender class. When an environment build's VM dies before the reporting agent starts, the reconciler notices the record has sat in "building" too long, asks the fleet service whether the VM still exists, and flips it to failed when the answer is gone or dead. That flip turns a silent hang into a failed state a user can see and diagnose (when environment builds fail).
Kubernetes controllers are the canonical form of the pattern: control loops that move observed state toward desired state instead of trusting that every edge-triggered event arrived (Kubernetes docs). The minimal version is one sweep per resource type that can get stuck.
Per resource type is the trap. The sweep you wrote for environment builds does not cover the session VMs you add next quarter, and the missing sweep is invisible until a resource strands with no error, usually as a user-facing timeout. The change that adds a resource type should add its reconciler.
For racing callbacks, designate an authority and read-repair
When two services learn about the same resource by different paths, ordering is a coin flip: the in-VM agent reports readiness straight to the orchestrator while the host agent reports the VM's port allocations to the fleet service, so the orchestrator can act on readiness before the fleet service's row has the ports.
The fix is a designated authority plus read-repair. Decide which system sits closest to each fact (the host allocated the ports, so the host layer is the authority) and have intermediate services backfill from it on a miss: when the orchestrator asks for the VM and the ports are absent, the fleet service queries the authority synchronously, stores the answer, and returns it. The common case, where the callback already landed, skips the round trip.
Writes racing writes get the database version of the same idea: upserts with conflict resolution (ON CONFLICT DO UPDATE in Postgres) so two overlapping writers converge on one row instead of deadlocking or duplicating.
Price this layer into the assemble path
Build, assemble, or buy prices the assemble path as integration plus governance: you rent the sandbox layer, then own the seams and the compliance story. Delivery reliability is a third cost, orthogonal to both. Every seam you own needs the twice-late-never treatment, and every resource type or callback you add extends the matrix.
It gets under-budgeted because none of it shows up in a working demo. A demo never loses a packet; production loses one in week two, and the symptoms (a duplicated job, a session stuck at boot, a half-populated read) each look like an application bug until the pattern emerges. The published in-house builds are silent here: Shopify documents its append-only event log and Stripe its pre-warmed devboxes, but nobody in the public record documents webhook dedupe, retry policy, or reconciler cadence, so the layer appears on no checklist you can copy.
Two questions audit the layer. For each hop: what happens when this message arrives twice, late, or never? For each resource with an async reporter: who notices when the reporter dies before first contact?