Agent-friendly documentation: a Markdown index for AI agents is at /llms.txt. Every page is also available as clean Markdown by appending .md to its URL (e.g. /pricing.md), or by requesting it with an Accept: text/markdown header.

BuildSEOAI
All posts
Engineering· 10 min read

Eight Agents, One Graph: Keeping Failure Local

A slow submission stage should never stall discovery. How a state machine per campaign lets every stage retry and resume without corrupting the run.

AP
Anurag Pattnaik
Platform engineering, Andolasoft · June 20, 2026

The failure mode we designed against is specific and familiar: one publisher with a slow form holds a lock, the submission stage backs up behind it, and a campaign that is ninety per cent finished displays as running for eleven hours. Nothing is broken. Nothing is progressing either, and no screen in the product can tell you which.

The fix is not better error handling. It is making sure a failure has somewhere small to be.

The graph

A campaign is a directed graph of eight stages, each owning one transition and each independently re-runnable. Nothing in the pipeline calls the next stage directly; a stage writes a state and the scheduler decides what becomes eligible.

STAGE
RETRY
Intake
Free
Discovery
Free
Qualification
Free
Content
Free
Form mapping
Free
Submission
Once
Verification
Free
Reporting
Free

The column that matters is the retry policy, and it is a property of the stage rather than of the queue. Seven of the eight are free to re-run because re-running them recomputes a value. One of them files a form on somebody else's website.

The unit of retry is one publisher, not one campaign

Every stage owns a state transition on a single unit of work — one publisher at one stage — and never on the campaign. The campaign's own status is derived from its children on read. It is a view, not a source of truth, and it is never written by a worker.

If the unit of retry is the campaign, you have not built a pipeline. You have built a batch job with a progress bar.

This is what makes failure local. A publisher whose form mapping fails is one row in a failed state with a reason attached; discovery keeps running, content generation keeps running, and the other two hundred rows proceed. The operator sees a campaign that is running with three items needing attention, which is the truth, instead of a campaign that is stuck, which is not.

Submission is the stage that is not idempotent

Recomputing a score is free. Regenerating a description is nearly free. Filing the same directory form twice creates a duplicate listing, an annoyed moderator, and occasionally a ban on the domain you were building links for.

So submission gets an at-most-once guard rather than an at-least-once one. A claim token is written for the publisher-attempt before the browser is opened, and it is written in the same transaction that moves the row into the submitting state. A worker that cannot claim the token does not submit — it exits. Crash recovery on that stage is deliberately conservative: an attempt that was claimed but never resolved is surfaced to a human rather than retried, because the system genuinely cannot tell whether the form went through.

Timeouts have to be shorter than the queue's patience

The most expensive bug we have shipped in this area was not a crash. Discovery hung forever, and the campaign displayed as discovering publishers with no error anywhere in the logs. The cause was a broad exception handler around the stage body. The queue enforces its own soft time limit by raising inside the task, and a handler that catches everything catches that too — so the stage swallowed its own deadline, logged a warning, and waited.

Two rules came out of it. A stage never catches the base exception class without re-raising the queue's control exceptions, and every stage has an internal timeout strictly below the worker's soft limit so the deadline is hit by our code first, where it can be recorded as a stage failure with a reason.

The third change was a reaper. A row that has been in a running state for longer than the stage's maximum plausible duration is transitioned to failed by a periodic sweep, whether or not any worker ever reports back. Anything that depends on a worker surviving to report its own death will eventually stall silently.

Commit as you go, and re-bind the tenant when you do

All-or-nothing commits are the other way a long stage loses its work. Discovery now commits publishers incrementally, so a stage that dies at minute nine keeps the eight minutes of results it earned. There is a sharp edge in that change worth flagging: our row-level security uses a session variable to scope every query to a tenant, and committing a transaction resets it. Incremental commits have to re-bind that variable on the new transaction, or the next write lands with no tenant context at all.

Restart without corruption

Operators restart campaigns, usually after fixing something upstream, and the naive implementation races: the old run's in-flight workers finish after the new run has started and write their results over it. Every campaign therefore carries a run epoch. Workers capture it when they claim work and compare it on write. Results produced under a superseded epoch are discarded rather than applied, so restart is a clean boundary instead of a merge of two runs.

It is worth saying plainly that we found this with an end-to-end test that restarted a campaign twice in quick succession, not by reasoning about it. Concurrency bugs in a stage graph are cheap to reason about and expensive to believe.

TAKEAWAYS
Make the unit of state one work item at one stage; derive campaign status on read so no worker ever writes it.
Give the one non-idempotent stage an at-most-once claim token and escalate unresolved attempts to a human instead of retrying.
Never catch the base exception class in a task body — it swallows the queue's own time limit and the stage hangs with no error.
Add a reaper for rows stuck in a running state; anything that relies on a worker reporting its own death will stall silently.
Commit long stages incrementally, and re-bind the tenant session variable after every commit.
Share this post

See the Score Run on Your Category

We will qualify live publisher inventory in your vertical and show you every signal behind it.

Schedule a Demo