# Eight Agents, One Graph: Keeping Failure Local

> Agent-friendly Markdown of https://www.buildseo.org/blog/eight-agents-one-graph — full index for AI agents: https://www.buildseo.org/llms.txt

_Engineering · 10 min read — Anurag Pattnaik, Platform engineering, Andolasoft · June 20, 2026_

A slow submission stage should never stall discovery. How a state machine per campaign lets every stage retry and resume without corrupting the run.

The failure mode we designed against is specific and familiar: one publisher with a slow form holds a lock, the submission stage backs up behind it, and a campaign that is ninety per cent finished displays as running for eleven hours. Nothing is broken. Nothing is progressing either, and no screen in the product can tell you which.

The fix is not better error handling. It is making sure a failure has somewhere small to be.

## The graph

A campaign is a directed graph of eight stages, each owning one transition and each independently re-runnable. Nothing in the pipeline calls the next stage directly; a stage writes a state and the scheduler decides what becomes eligible.

| STAGE | RETRY |
| --- | --- |
| Intake | Free |
| Discovery | Free |
| Qualification | Free |
| Content | Free |
| Form mapping | Free |
| Submission | Once |
| Verification | Free |
| Reporting | Free |

The column that matters is the retry policy, and it is a property of the stage rather than of the queue. Seven of the eight are free to re-run because re-running them recomputes a value. One of them files a form on somebody else's website.

## The unit of retry is one publisher, not one campaign

Every stage owns a state transition on a single unit of work â€” one publisher at one stage â€” and never on the campaign. The campaign's own status is derived from its children on read. It is a view, not a source of truth, and it is never written by a worker.

> If the unit of retry is the campaign, you have not built a pipeline. You have built a batch job with a progress bar.

This is what makes failure local. A publisher whose form mapping fails is one row in a failed state with a reason attached; discovery keeps running, content generation keeps running, and the other two hundred rows proceed. The operator sees a campaign that is running with three items needing attention, which is the truth, instead of a campaign that is stuck, which is not.

## Submission is the stage that is not idempotent

Recomputing a score is free. Regenerating a description is nearly free. Filing the same directory form twice creates a duplicate listing, an annoyed moderator, and occasionally a ban on the domain you were building links for.

So submission gets an at-most-once guard rather than an at-least-once one. A claim token is written for the publisher-attempt before the browser is opened, and it is written in the same transaction that moves the row into the submitting state. A worker that cannot claim the token does not submit â€” it exits. Crash recovery on that stage is deliberately conservative: an attempt that was claimed but never resolved is surfaced to a human rather than retried, because the system genuinely cannot tell whether the form went through.

## Timeouts have to be shorter than the queue's patience

The most expensive bug we have shipped in this area was not a crash. Discovery hung forever, and the campaign displayed as discovering publishers with no error anywhere in the logs. The cause was a broad exception handler around the stage body. The queue enforces its own soft time limit by raising inside the task, and a handler that catches everything catches that too â€” so the stage swallowed its own deadline, logged a warning, and waited.

Two rules came out of it. A stage never catches the base exception class without re-raising the queue's control exceptions, and every stage has an internal timeout strictly below the worker's soft limit so the deadline is hit by our code first, where it can be recorded as a stage failure with a reason.

The third change was a reaper. A row that has been in a running state for longer than the stage's maximum plausible duration is transitioned to failed by a periodic sweep, whether or not any worker ever reports back. Anything that depends on a worker surviving to report its own death will eventually stall silently.

## Commit as you go, and re-bind the tenant when you do

All-or-nothing commits are the other way a long stage loses its work. Discovery now commits publishers incrementally, so a stage that dies at minute nine keeps the eight minutes of results it earned. There is a sharp edge in that change worth flagging: our row-level security uses a session variable to scope every query to a tenant, and committing a transaction resets it. Incremental commits have to re-bind that variable on the new transaction, or the next write lands with no tenant context at all.

## Restart without corruption

Operators restart campaigns, usually after fixing something upstream, and the naive implementation races: the old run's in-flight workers finish after the new run has started and write their results over it. Every campaign therefore carries a run epoch. Workers capture it when they claim work and compare it on write. Results produced under a superseded epoch are discarded rather than applied, so restart is a clean boundary instead of a merge of two runs.

It is worth saying plainly that we found this with an end-to-end test that restarted a campaign twice in quick succession, not by reasoning about it. Concurrency bugs in a stage graph are cheap to reason about and expensive to believe.

**Takeaways**

- Make the unit of state one work item at one stage; derive campaign status on read so no worker ever writes it.
- Give the one non-idempotent stage an at-most-once claim token and escalate unresolved attempts to a human instead of retrying.
- Never catch the base exception class in a task body â€” it swallows the queue's own time limit and the stage hangs with no error.
- Add a reaper for rows stuck in a running state; anything that relies on a worker reporting its own death will stall silently.
- Commit long stages incrementally, and re-bind the tenant session variable after every commit.

## Keep Reading

- [The Hidden Cost of Manual Link Submission: A 5-Year ROI Case Study](https://www.buildseo.org/blog/the-hidden-cost-of-manual-link-submission.md): One agency spent $840k on manual link building over five years. We mapped where every dollar went, and why their cost per verified link was 12Ã— higher than they reported.
- [Cost Per Link: How To Calculate It, Why Vendors Hide It, and What It Actually Means](https://www.buildseo.org/blog/cost-per-link-what-it-actually-means.md): Every link-building platform quotes a different number. Here's how to calculate the one that matters: cost per link still live at 90 days. And why that number is the only one your CFO should care about.
- [The Link Database Advantage: Why Campaign Two Should Cost 20% Less Than Campaign One](https://www.buildseo.org/blog/link-building-database-compounding.md): Campaign one teaches you everything about your vertical. Campaign two should cost less, take less time, and deliver more verified links â€” because you kept what you learned.

## See the Score Run on Your Category

We will qualify live publisher inventory in your vertical and show you every signal behind it. Schedule a demo: https://calendly.com/buildseo-sales/30min

---

[All posts](https://www.buildseo.org/blog.md) · [BuildSEO AI](https://www.buildseo.org/index.md)

