Event Routing: Compile the Trigger Once, Dedupe Every Spawn
Work intake establishes where tasks come from: listeners on chat, issue trackers, GitHub, and cron. Between the listener and the spawned task sits the router, the component that decides whether this event starts that workflow. Three properties make it worth designing deliberately: its input is attacker-controlled, every new source can demand new integration code, and webhooks redeliver.
The spawn decision is a consequential action fed untrusted content
A webhook payload is attacker-controlled text. Anyone who can open an issue, comment on a PR, or file a support ticket writes into it. And spawning a workflow is a consequential action: the spawned task checks out code, runs tools, and opens pull requests. The design rule the lethal trifecta quotes applies to the router without modification: "once an LLM agent has ingested untrusted input, it must be constrained so that it is impossible for that input to trigger any consequential actions."
The intuitive router violates it. Having a model read each inbound payload and judge whether it matches a workflow's intent puts an LLM between attacker text and task creation on every event, so a crafted issue body that talks the classifier into a match spawns an agent. Even with no attacker, that router is non-deterministic: the same event, replayed, can route differently, so you cannot test a routing change against history or reconstruct why an incident's task spawned.
Compile the plain-English condition at authoring time
The model's job moves to authoring time. The workflow author writes the condition once, as a sentence: run this when a bug-labeled issue is opened in the payments repo. A compiler turns it into a deterministic filter, a (source, type) pair plus a small match expression over payload fields: equality, membership in a list, boolean combinators. Ground the compiler on sampled events from the account's own history, so it names payload paths that actually exist, and have it preview which recent events would have matched before the author saves.
At edit time the model reads the author's own sentence and stored payloads the team already trusts. At runtime there is no model at all: routing is an indexed lookup on (source, type) followed by filter evaluation. That path is deterministic, replayable against history, testable in CI, and cheap enough to run on every event of a busy account.
The trade is fuzzy semantics. "Only spawn on serious-looking bug reports" does not compile. Move that judgment one step downstream: spawn deterministically on label and repo, and make the first step of the workflow a triage decision that can end the task early when the report is not a real bug. The model still judges, but it judges inside a state machine whose gates bound what the judgment can trigger, and every verdict leaves a step record.
Schedules ride the same path. The compiler classifies "every weekday at 9am" as a schedule and emits a cron expression plus timezone. The tick job fires it by emitting a synthetic event from a reserved internal source, routed exactly like an external event, so a scheduled spawn is a recorded, queryable fact the task links back to. Reserve that internal source at the ingest boundary; otherwise an outside caller can forge a scheduler event and fire your cron workflow on demand.
One generic ingest URL replaces per-provider listeners
The channel table in work intake is a list of bespoke integrations, and each new external system adds another listener. The alternative is to normalize at the edge: pick one event shape (CloudEvents hands you source, type, and data) and mint a generic collector URL per external source. The team pastes that URL into anything that can POST a webhook: the error tracker, the CRM, the ticket queue.
The collector does three cheap things. The URL itself is the credential, an identifier plus a secret; store the secret hashed and return 404 on any mismatch so the URL never confirms a collector exists. It stamps its configured source and derives type from the delivery, by fixed default, named header, or a JSON path into the body, falling back to the default rather than dropping the event. Then it writes through the same ingestion path as every first-party event, so routing, dedupe, and the event log apply identically.
Signature verification is the graduated-trust upgrade. Most senders sign deliveries with HMAC-SHA256 over the raw body, so a collector can hold an optional signing secret and reject unsigned or mismatched deliveries. Start with the URL credential, add the signature check when the workflows a source can spawn get consequential, and rotate the URL secret like any other credential.
Redelivery is normal, so spawning is idempotent twice over
Webhook senders retry on timeouts and 5xx by design: a delivery that succeeded, but whose response got lost, arrives again. Your own internal emitters retry too. A router that spawns one task per received event will eventually spawn two tasks for one issue, silently, and the duplicate does real work: an agent boots, writes code, and may open a competing PR. Idempotency lives at two layers.
- Delivery-level. Every delivery gets an id: the provider's own delivery id when present, else a deterministic hash over the payload's identifying fields. Insert it against a unique index in the same transaction that stores the event; a duplicate trips the index and the whole call no-ops. This absorbs bit-for-bit redelivery.
- Entity-level. The trigger declares a dedupe key rendered from payload fields, repo plus issue number for issue-driven work. When an open task on the same workflow already carries that key, the new event attaches to the existing task as an update instead of spawning a second one, and the dedupe itself is recorded as an event. This absorbs genuinely distinct deliveries about the same entity: the issue edited, relabeled, or synced in by a second integration.
Schedule ticks need their own guard: claim each due trigger with a conditional first-writer-wins update so overlapping sweeps cannot double-fire, and re-base the next run on now so an outage replays as one catch-up fire, not a burst.
The record is silent on routing mechanics, so hold the bar yourself
Brex publishes the closest thing to a routing statement, one sentence: "The orchestrator routes them to the agent pool" (Brex). How matching works, whether spawns are idempotent, and what a redelivered webhook does are unpublished in every system the intake corpus covers. The design above is drawn from one platform build (case study), and its properties are checkable on any implementation:
- No model call anywhere in the per-event path.
- The same event, replayed, routes identically, so routing changes are testable against history.
- A new external source is configuration: a URL and a type rule, no new listener code.
- A redelivered webhook or an overlapping cron tick cannot create a second task.
- Every spawn, dedupe, and scheduled fire is a queryable record linking the event to the task it produced.