Per-Turn Brains: Dispose the Harness Every Turn and Resume Becomes the Boot Sequence
One environment per task sets disposability at task granularity: one sandbox per task, torn down after. The brain can run on a finer unit than that. In the per-turn model, a fresh harness boots for every message, runs one exchange, and is destroyed when the turn's terminal event lands. One published implementation runs exactly this way (Wallfacer case study; disclosure: Wallfacer funds this guide).
Nothing in the brain/hands/session split forbids it; Anthropic already states the premise, "nothing in the harness needs to survive a crash," because the session log sits outside it (Scaling Managed Agents).
Per-turn disposal removes idle compute, accumulated state, and long-lived credentials
No idle brain. A conversation is mostly human think-time. A task-lived brain either sits allocated between messages or needs eviction machinery to reclaim it, which is why Shopify built process-exit-and-rehydrate for idle sessions (Shopify case study). Per-turn disposal makes eviction unconditional: the brain exists only while the model is working.
No accumulated state. Config drift, leaked temp files, a wedged CLI process: each is bounded to one exchange, because the machine that developed the problem is gone before the next message arrives.
Upgrades become image promotion. Every turn boots from the current harness image, so shipping a new harness build or a new agent CLI version means promoting a snapshot (versioned images). There is no long-lived process to drain and no fleet of half-upgraded brains.
Short-lived credentials by construction. The machine holding the model API key and the session's bearer token dies minutes after it boots, so a credential's maximum exposure on the brain is the length of one exchange.
Resume stops being a recovery branch and becomes the boot sequence
The session teaches resume as replay from the external log, and its motivating cases are exceptional: a crashed harness, an idle task whose host got reclaimed. Shopify's eviction-and-rehydration made replay the common path at tens-of-thousands-of-sessions scale (Shopify case study). Per-turn disposal finishes the move: replay-from-log runs on message two, message three, message forty-seven, in every ordinary conversation.
The resume path cannot rot. A recovery branch exercised only during crashes is untested code that fails exactly when you need it; a resume path that is also the boot sequence gets tested by every conversation.
Boot cost also lands on the critical path of every message, so the brain has to boot from a restored snapshot in seconds, and the image has to stay small. The scheduling levers in placement, warm pools and snapshot-cache locality, apply to the brain as much as the hands.
Rebuilding conversation state: re-render canonical events, or hand the vendor its own bytes back
The durable log holds canonical events. The fresh brain needs a conversation. Which path gets you from one to the other depends on whether you own the agent loop.
Re-render from the canonical log. Anthropic's published design: a new harness calls wake(sessionId), fetches the log with getSession(id), and assembles the context window itself (Anthropic). One representation, vendor-neutral, and the harness controls exactly what enters the context. It works because that harness makes raw model calls and owns the context-assembly logic end to end.
Replay the vendor's own transcript. If the brain wraps a vendor CLI instead of raw model calls, the loop's internal state is the CLI's own session artifact, in a format you do not control. The CLIs already resume across invocations from state they keep on disk; Claude Code persists sessions locally and exposes --resume (CLI reference). The per-turn design piggybacks on that machinery: the turn's last act is to upload the CLI's on-disk session artifact to the control plane as an opaque blob, and the next boot's first act is to put the same bytes back where the CLI expects them and pass the resume flag. The blob is never parsed, so the vendor's format can change without breaking you.
The blob path costs storage twice, a canonical stream for rendering and audit plus an opaque per-vendor artifact for resume, and the artifact's path and format differ per CLI, which is work for the adapter layer (multi-vendor harness adapters). The canonical log stays the source of truth; the blob is a cache of the vendor's private state.
Rebooting the writer every turn adds a contract: a fresh brain retries event uploads after transient failures, so every event needs a unique id and a session-scoped sequence, and ingest has to deduplicate (retries, idempotency, and reconciliation).
The fresh brain cannot see the code, so orientation is pushed to it before the first tool call
Placement establishes that the harness makes no assumption about where the sandbox lives. Per-turn disposal surfaces the consequence: the brain has no filesystem view of the hands. The import-and-symlink bridges in instruction files assume the agent opens AGENTS.md or CLAUDE.md in the repo checkout, and this brain has no checkout. It cannot list the environment's services, read its manifest, or learn which port the dev server holds, except by burning its opening turns on tool-call archaeology.
The fix is out-of-band orientation. The control plane renders the environment's shape, the repos, the services and their ports, the named commands, the connected tool servers, conventions like the commit contract, into instruction-file content and ships it in the boot payload. The harness writes the full content to every instruction filename the CLI might read, then spawns the CLI, so orientation costs zero opening tool calls.
Duplicating full content per filename trades cleanliness for reliability; the one-source-of-truth argument in instruction files still holds, because the source is the renderer, and the files are per-boot output nobody edits.
Rendering per boot also makes orientation self-updating: add a service or connect an MCP server and the next turn's brain sees it, no rebuild. The same channel can carry a table of contents for the team's written knowledge, so the agent knows what it can look up (knowledge base for agents).
The costs: boot on the critical path, doubled storage, and stale inherited config
Boot lands on every message: snapshot restore makes it a few seconds (Wallfacer case study reports roughly 2 seconds for VM restore), and no source publishes a measured end-to-end per-turn overhead, so budget for it. Storage doubles where the vendor-transcript path is used. Anything the CLI wrote locally mid-turn that reached neither the event log nor the transcript blob is gone by design.
The trap: a machine restored from a snapshot inherits the environment of the machine the snapshot was taken from, including a stale session id pointing at some other conversation's log. The boot contract must treat the freshly delivered config as authoritative and overwrite inherited values, not merge around them.