The Workflow Is a State Machine: Typed Steps, Bounded Loop-Backs, Early Exit

A review gate answers one question: may this work advance. It does not answer what happens when the answer is no, or how a task with five steps moves between them. The systems with the biggest published wins answer with a state machine. Airbnb's 97% automated test migration ran as a state-machine pipeline with retry loops. Stripe orchestrates its minions with Blueprints, a state machine interleaving agent loops with deterministic nodes so required steps always run. Both name the pattern; neither publishes the mechanics.

A workflow is an ordered list of typed steps. Each step declares a failure route and a success route. Four routing values cover everything in the published record: halt, continue, loop back on a budget, and exit early.

Steps are typed: an agent acts, a person decides, a wait watches

  • An AI step. Runs one agent turn in an isolated environment and produces work product plus an outcome. Each attempt is a fresh turn against durable task state, not a resumed process (see the per-turn brain lifecycle), which is what makes an attempt safely repeatable.
  • A human step. Parks the machine. Nothing advances until a named person acts, and the act is the record. Review as a gate covers the enforcement; the gate is a step with routes like any other, and the author step and approver step are different actors by construction (author, never approver).
  • A wait step. Re-checks a condition on an interval: is CI green, did the canary hold, has the fix deployed. Pass advances, fail routes like any failure, and the re-checks count against a budget.

All three kinds produce the same artifact, one recorded attempt with a kind, an owner, and a typed outcome (every step has an owner). That symmetry is what lets the machine route on outcomes without caring who produced them.

Every step declares its failure route: halt, continue, or loop back on a budget

Halt. The default. The task stops, marked as an error, and surfaces in a queue where a person decides what happens next. Halt is also the floor under every other route: any budget that exhausts lands here.

Continue. The failure is advisory. A nice-to-have enrichment step that fails should not kill a task whose real work succeeded.

Loop back. The route names an earlier step. A review step that rejects sends the task back to the implement step, bounded by a per-step attempt count, and exhausting the count halts. This is the outer loop, across steps, and it is distinct from the inner loop inside one step where an agent iterates against CI (closing the feedback loop). Stripe bounds its inner loop at two CI rounds, then the branch goes to a human.

A loop-back that replays the original prompt just buys another sample from the same distribution. The retry has to carry memory: each new attempt receives an account of prior attempts and why each was rejected, so attempt three argues with the review feedback instead of repeating attempt one. Airbnb's pipeline is the nearest published analog, regenerating retry prompts from the failure state rather than replaying the original, and handing files that exhausted the loop to humans.

Success routes too: a happy branch can end the task early

The success route has two values: continue, the default, and complete, which ends the whole task as a success at the current step and skips the rest. Consider bug triage: a decision step checks whether the report reproduces. The valid branch enriches the issue and is finished; the invalid branch should flow to a human confirmation and a close step. With early exit, the decision step's happy path completes the task and its failure route carries the other branch onward, so one step routes two ways and both paths land in the record. Without it, the branch gets modeled inside a prompt, as an instruction to skip the remaining steps, which the machine cannot enforce and the record cannot show.

This model is deliberately weaker than a general DAG: one ordered list, loops only backward, exits only forward. That is enough for every published shape, implement then review then approve then merge, triage with a happy exit, and the restriction keeps every execution readable as a straight line with annotated detours.

The verdict that routes is a schema-bound tool call

Routing needs an unambiguous outcome, the per-step version of verifiable outcomes. A review step declares an output schema, approved plus a list of must-fix items, and the reviewing agent submits its verdict through a tool call, so the verdict lands as a typed event rather than a sentence to grep for. An approved: false is a step failure, routed like any other, which is what drives the review-and-fix loop without a human relaying the rejection.

A wait budget counts checks instead of watching the clock

The wait step's budget is a count: re-check every five minutes, up to six checks. A count beats a wall-clock deadline here because each check has its own cost and duration; when the check is itself an agent turn, a forty-minute timeout is a fuzzy guess at the number of checks the author actually meant. An exhausted budget is a failure and takes the step's failure route, so "wait for CI, stop if it never greens" needs no extra machinery.

The workflow assumes every attempt runs once and settles once; a VM that dies mid-check or a completion webhook that delivers twice is not a workflow failure to route, it is an infrastructure problem for retries, idempotency, and reconciliation to absorb below the step layer.

The published record names the pattern and stops there

Tasks enter the machine from routed events and leave it done, completed early, or halted in a human queue, with one recorded action per step attempt in between. That per-attempt trail is what makes a five-step, three-loop task auditable after the fact, where a hand-rolled fix-loop leaves only a chat transcript.

The published record supports the pattern and its results, Airbnb's 97% and Stripe's 1,300+ weekly PRs, and goes no further: no company has published its routing vocabulary, its loop budgets, or its verdict schemas. The vocabulary above is a synthesis, consistent with what the published systems describe.