A Two-Day Loop That Paused, Resumed, and Shipped

One advisor request became three live pages: 9 workflow runs, 23 tasks, 0 failures, $14.14 model cost, 4 pull requests across 47 hours — every merge human.

  • Workflows
  • Orchestration
  • Recovery
  • Approvals

Real session, redacted. This ran in our own swarm from 5 to 7 October 2026. Task, run, and chat identifiers are removed, and the pull requests live in a private repository, so their numbers are named but not linked; the verifiable artifacts are the three live example pages this session shipped. All times are Barcelona time.

Request

An outside SEO advisor reviews our search visibility and files recommendations. One of them became a standing request in the swarm: add real example pages to the /examples collection — measured walkthroughs of actual sessions, not capability claims.

One short agent loop could not carry it. The work needed prose, code, review, and a human merge. It also needed an outside reviewer to read the live site, which only happens after a deploy. No single session spans hours of waiting on review and deploy, so the request ran as a workflow instead: many short runs, each resuming where the last one stopped.

Starting context

The workflow runs on a schedule — twice a day, plus manual runs. Each run reads a state record the workflow keeps for the request: the stage (new, prose_approved, live_feedback), the open pull request and its branch, and a failed-attempt count. From it, the run routes the request to the right role. The state record is the checkpoint — the shared state every run resumes from. Workers never receive chat history — only artifacts: a brief, a draft file, a pull request.

Task breakdown

  1. Run 1, 5 Oct, 12:55 to 13:12. The Implementer read the request, saw that new prose was needed, and returned a brief in 53 seconds instead of writing copy — writing is not its job. The Content Strategist drafted three pages, and the Content Reviewer approved them.
  2. Run 2, 13:29 to 13:49. Pull request #232 opened at 13:36. The Reviewer returned FIX: a customer name had to come out of the opening paragraphs, plus one copy change.
  3. The pause. The fix step refused to push. The fix needed new published prose, and only the Content Strategist writes published prose — so the step stopped with nothing pushed. The pull request head stayed unchanged, and the checkpoint kept the stage, the pull request, and the branch for the next run.
  4. Run 3, 14:08 to 14:31. The run resumed from the stored pull request and branch and sent only the delta to the Content Strategist; the Content Reviewer approved it. Nothing was redrafted and no duplicate pull request opened. A human merged #232 at 15:11.
  5. The second loop. Live feedback said the pages were illustrative, not real. A founder made the call: example pages show real sessions only. Run 5 (18:12 to 18:32) opened #238 with real redacted sessions. The Reviewer caught a measurement error — a "4 h 8 min" figure measured pull-request-open to merge, not first-commit to merge — the figure was re-anchored, and a human merged #238 at 19:54.
  6. The quiet runs. #239, a Reviewer PASS, merged on 6 October at 10:28. Two later runs that day found nothing actionable and stopped without opening a pull request. #246 — a layout-only reorder, no new copy — merged on 7 October at 11:45.

Worker roles and handoffs

Four specialized roles, and none does another's job. The Implementer owns code: the worktree, the checks, the pull request. The Content Strategist writes every published word. The Content Reviewer judges the prose. The Reviewer judges the diff and the redaction.

Approval points

Three gates held. The Content Reviewer's pass is the only prose approval — no draft ships on its writer's say-so. The Reviewer's pass gates every pull request, and it issued two FIX verdicts that were caught before merge, not after. Every merge was a human click, and the "real sessions only" rule was a founder's decision, not a workflow default.

Output

Four merged pull requests: #232 (the three walkthrough pages), #238 (real redacted sessions replacing the illustrative drafts), #239 (the re-anchored measurement), and #246 (a layout reorder). Three live pages came out of it: an engineering pull request, a recurring research brief, and a security-alert triage.

Measurable result

From the swarm's task, workflow-run, and session-cost records: 9 workflow runs over about 47 hours — 5 October, 12:55 to 7 October, 11:46 — 23 agent tasks, 0 failed tasks, about 77 agent-minutes of compute, and $14.14 in measured model cost. Cost by role: Implementer $9.53, Content Strategist $3.04, Reviewer $1.53, Content Reviewer $0.04. Two Reviewer FIX verdicts caught before merge, one refused push recovered without redrafting, zero duplicate pull requests.

What stayed human-owned

Every merge, the "real sessions only" rule, and redaction sign-off — a Reviewer pass plus the human merge. The swarm owned the drafting, the code, the checks, the review rounds, and the retries.

Sources and tools

More real sessions: the engineering pull request merged in four hours, the recurring research brief that runs every morning, the security-alert triage that closed thirty-three Dependabot alerts, and the fully public x402 payment session. Browse everything on the examples index. The workflow, scheduling, and review primitives behind this page are open source in desplega-ai/agent-swarm.