The Swappable Brain Across Vendors: One Canonical Event Schema, One Adapter Per CLI

The published case for a swappable brain is same-vendor. Anthropic's context-reset workaround for Claude Sonnet 4.5 became "dead weight" on Claude Opus 4.5, and because the harness was decoupled, the model upgrade was a swap rather than a re-architecture (Anthropic). Swapping across vendors is a stronger claim, and teams already want it: Stripe runs three coding agents against one synced rule set (Stripe case study).

A brain that is replaceable across vendors, the property the reference architecture treats as a payoff of the brain/hands split, takes two things the public record is nearly silent on: a canonical event schema that everything downstream consumes, and a per-vendor adapter that translates into it before anything durable sees a vendor-specific shape.

A vendor CLI differs in three places, and none of them is the argv

Event vocabulary. Every agent CLI with a non-interactive mode streams structured events to stdout, and no two vendors stream the same shapes. One emits messages with typed content blocks, tool calls inline as tool_use paired with tool_result in the following message. Another emits item-completed events, command executions, MCP tool calls, reasoning items.

Transcript on disk. Each CLI persists its own resumable session artifact, in its own format, at its own location. One writes a JSONL file at a deterministic path derived from the working directory; another keeps a database file under its home directory. The vendor's resume flag reads that artifact and nothing else.

Tool transport. MCP support differs in capability, not just configuration. One CLI dials an HTTP MCP endpoint directly and surfaces its tools as function calls. Another does not surface HTTP-transport tools as first-class function-call handles, so its adapter bakes a small stdio relay per server wrapping the same endpoint.

The instruction file is the same divergence one layer up: each vendor reads a different filename, so a multi-vendor platform writes the per-session context to every filename its adapters read (one source of truth, bridged to each harness).

Translate at the edge, before the durable log sees a byte

The adapter's translation runs in the streaming path: parse each stdout line, rewrite it into the canonical schema, append it to the session log. Downstream of that point nothing is vendor-specific. The UI renders one shape, the audit trail stores one shape, and the workflow engine that settles steps from terminal events reads one shape.

Skipping the canonical layer does not save the translation work; it distributes it. The chat renderer, the audit query, and the step-settling logic each grow a branch per vendor, and every new vendor multiplies across every consumer instead of adding one adapter.

Two schema decisions do most of the work. Model the canonical shape on the richest vendor format you already render, so at least one adapter is the identity function and translation cost lands only on vendors that diverge.

Synthesize stable tool identifiers for vendors whose events do not carry tool names in the shape you filter on: "shell command ran" gets one fixed name, "patch applied" another, so a query like "which turns touched files" filters on canonical names rather than switching per vendor. Terminal events, turn completed cleanly and turn aborted, are canonical too, emitted by the wrapper from the CLI's exit status rather than trusted to any vendor's vocabulary.

Store the raw bytes next to the canonical event

Each durable event row should carry three things: the canonical payload, the verbatim bytes the vendor emitted, and a provider tag naming which adapter produced it.

  • Translation is code, and it will have bugs. When a turn renders wrong, the raw bytes distinguish "the vendor emitted something new" from "the adapter rewrote it badly," and a fixed adapter can re-derive the canonical rows.
  • The canonical schema will change. Raw bytes are the migration path: re-translate history instead of declaring old sessions unreadable.
  • An audit trail that holds the unmodified original with provenance lets the auditor read the same bytes the vendor emitted, with the rendering derived rather than authoritative.

Where the translation is the identity, store both columns anyway. The redundancy costs a text column and keeps the next vendor's adapter a drop-in.

Resume round-trips the vendor's transcript verbatim

The canonical log answers rendering, audit, and replay questions. It cannot resume the vendor CLI, because each CLI reconstructs its context from its own on-disk artifact, full of vendor-internal fields the canonical schema deliberately does not model.

So the transcript rides alongside the event stream as a second durable artifact. At end of turn, read the vendor's on-disk file and persist it as an opaque blob keyed by the vendor's session id. At the next turn's start, write the same bytes back to the path the CLI expects and invoke its resume flag. Never inspect, migrate, or normalize the blob: an edited blob means the platform now owns compatibility with a format the vendor never published.

The round-trip has to go through durable storage because of the per-turn brain lifecycle: when the machine that ran the CLI is destroyed after every turn, the transcript's only home between turns is off the machine, exactly like the session's own state. One managed implementation runs this shape today: two vendor adapters, a fresh harness machine per turn, and the session record persisted outside both machines.

The adapter interface is four methods; the dispatch sites are the real checklist

The interface a new vendor implements is small: a preflight that renders vendor-specific config files (MCP config, relay scripts, home-directory settings), an argv layout for the spawn (prompt, streaming-output flag, resume flag, model override), the on-disk transcript path, and the event translation. The wrapper's main loop stays vendor-agnostic.

Outside the wrapper, the conditional sites accrete: which API key variable gets stamped into the harness environment, which model ids the backend whitelists, what assistant name the UI shows, the picker entry, the settings field where a user pastes the vendor key. None of these is architecturally interesting, and every one is missable, with failures that surface only on the first live turn: the CLI exits on a missing key, or events land tagged with the wrong provider.

Keep the dispatch sites written down and greppable, because the length of that list, not the elegance of the adapter interface, measures how pluggable the brain actually is. And make the adapter degrade per feature rather than per turn: a malformed MCP server definition should remove that one server for the turn with a warning, never fail the turn, because the turn's cost is a dead session in front of a user and the feature's cost is one missing tool.