Back to writing
October 7, 2026·10 min read

We deleted --resume. Continuity got better.

Why Agent Swarm dropped claude --resume and codex.resumeThread for a 2,000-token context preamble that survives container restarts and harness re-routing.

session continuitycontext preambleharness routingdeep diveagent-swarm
Distracted boyfriend meme. The woman in red: me chasing native --resume for session continuity. The boyfriend: --resume, silently forgets everything when a container restarts. The girlfriend: the session_logs table I already own.
The state was in our database the whole time.

In late May 2026 we removed the feature most people assume a coding-agent orchestrator cannot live without: native session resume. No more claude --resume <UUID>. No more codex.resumeThread(id). The module that used to pick a session to resume still exists, but it now returns nothing to resume and logs what it skipped. Continuity got better, not worse.

This post explains why. It also defends a belief we now hold firmly enough to delete code over: a conversation should belong to the orchestrator's database, not to a harness transcript file. Anything that moves work between runners has to rebuild context, not resume it.

Why does native session resume look like the right primitive?

Every agent harness ships some form of resume. Claude Code has claude --resume. Codex can resume a thread. The pitch is simple: the harness keeps a transcript, you hand back an ID, and the agent continues where it stopped, with nothing summarized away.

For one developer running one agent on a laptop, that is the right primitive. The transcript is on your disk, the disk persists, and the harness owns the format.

An orchestrator works differently. It runs agents in disposable worker containers, redeploys them, routes tasks between harnesses, and spawns follow-up tasks from parents. In that setting native resume has three failure modes, and we hit all three.

Failure mode 1: the transcript is a file, and files die

A resume handle points at state on the worker's filesystem. When the worker container restarts (a deploy, an OOM kill, an autoscaler reschedule), that filesystem is gone. The session ID stored in your database now points at nothing.

The deprecation PR describes what happened next: --resume <UUID> either errored with “session not found” or silently launched a fresh session with no context. The second case is the bad one. The agent starts work with no memory of the thread, and the user concludes it forgot the conversation. A crash pages someone. Silent context loss ships nonsense.

Persistent volumes would patch this, but then every task depends on landing on the worker that holds its transcript. We had kept workers stateless on purpose (see why our workers cannot touch the database). Resume would have pulled that coupling back in.

Failure mode 2: unbounded context saturates

Resume only appends. Every turn stays in the transcript, and every resumed session carries all of it forward. Long-running tasks are what an orchestrator is for, so the context keeps growing.

In our workers this showed up as sessions dying with SIGTERM, exit code 143. From the outside it looked like flaky infrastructure. The header of context-preamble.ts names the real cause: “the SIGTERM-143 context-saturation failure mode seen with unbounded session resumes.” Resume has no eviction policy and no budget.

Past a certain length, full fidelity stops helping. An agent carrying a long tail of stale tool output is not better informed than one holding a tight brief. It is slower, costs more, and is more likely to anchor on old history. We covered the general version of this in our context compaction deep dive.

Failure mode 3: resume cannot cross a harness boundary

This one matters most. cezar, an MIT-licensed tool that runs Claude Code, Codex, OpenCode and pi on the subscriptions you already pay for, says you can “switch a task’s runner, model or account mid-thread without losing the conversation.” That is the right goal. The question is which primitive gets you there.

It cannot be native resume. A Claude Code session ID means nothing to Codex, and a Codex thread means nothing to OpenCode. Transcript formats are per-harness. The moment a task moves to another runner, because a model is rate-limited or another agent is free, every stored resume handle is useless.

We hit a sharper version of this in our own schema. In src/be/db.ts, derived tasks (resume follow-ups, review follow-ups, re-dispatches) deliberately do not inherit the parent's model. A parent's model is a concrete, provider-specific string such as claude-opus-4-8. When a resume of that task was claimed by a Codex worker, it died at session start with 400 model is not supported when using Codex. The model now resolves from the assignee agent's own configuration. If continuity is bound to the harness, routing is bound to it too.

What replaced resume?

The replacement is boring, which is why it works. Everything that matters already lives in the orchestrator's database: task briefs, parent-child links, outputs, progress notes, and the session_logs rows that record what agents did. So instead of pointing a harness at its own transcript, the worker rebuilds a short brief over the API and puts it at the top of a fresh session.

First, the old resolver became an observability shim. This is the current resume-session.ts, trimmed:

export const RESUME_DEPRECATED_REASON = "native resume deprecated — using context preamble";

export function resolveResumeSession(
  _currentProvider: ProviderName,
  candidates: ResumeSessionCandidate[],
): ResumeSessionResolution {
  const skipped: ResumeSessionSkip[] = [];

  for (const candidate of candidates) {
    const sessionId = candidate.sessionId?.trim();
    if (!sessionId) continue;
    skipped.push({
      source: candidate.source,
      sessionId,
      provider: candidate.provider,
      reason: RESUME_DEPRECATED_REASON,
    });
  }

  return { skipped }; // resumeSessionId is always undefined
}

The adapters lost their resume branches in the same change: no --resume argument for Claude, no resumeThread for Codex. Continuity now comes from two builders in context-preamble.ts:

  • buildContextPreamble, capped at 2,000 tokens. It walks up to 5 ancestors. The immediate parent gets its brief, outcome and attachments inline. Older ancestors get a one-line pointer the agent can expand with get-task-details.
  • buildResumeContextPreamble, capped at 4,000 tokens. Used when an interrupted task is resumed. It keeps the original brief verbatim and adds one-line summaries of the last 50 session_logs rows, so the agent does not redo finished work. Log lines and steering messages pass through scrubSecrets first.
  • A hard character cap, not a target. The budget is roughly 4 characters per token. Anything over it is cut and marked as truncated. In the resume variant, the oldest log lines go first and the task brief is never cut.
export const CONTEXT_PREAMBLE_MAX_TOKENS = Number(
  process.env.CONTEXT_PREAMBLE_MAX_TOKENS || "2000",
);
// ~4 chars per token (conservative approximation for mixed code/prose)
export const CONTEXT_PREAMBLE_MAX_CHARS = CONTEXT_PREAMBLE_MAX_TOKENS * 4;
export const CONTEXT_PREAMBLE_MAX_ANCESTORS = 5;

export const CONTEXT_PREAMBLE_RESUME_MAX_TOKENS = Number(
  process.env.CONTEXT_PREAMBLE_RESUME_MAX_TOKENS || "4000",
);
/** How many of the most recent session_logs rows to inspect for tool-call summary. */
export const CONTEXT_PREAMBLE_RESUME_SESSION_LOG_LIMIT = 50;

Where the preamble goes matters more than you would expect. Our task prompts often start with a slash skill such as /work-on-task, and the skill resolvers only expand it when it is on line one. So the preamble goes right after that line, never before it:

/** Keep slash skills on line one so native and adapter resolvers can expand them. */
export function prependContextPreamble(prompt: string, preamble: string): string {
  if (!preamble) return prompt;
  const newline = prompt.indexOf("\n");
  const firstLine = newline === -1 ? prompt : prompt.slice(0, newline);
  if (!/^\/[a-z0-9:_-]+(?:[ \t].*)?$/.test(firstLine.trim())) {
    return preamble + prompt;
  }
  const rest = newline === -1 ? "" : prompt.slice(newline + 1);
  return `${firstLine}\n${preamble}${rest}`;
}

Does a context preamble beat native session resume?

For orchestrated work, yes. A bounded preamble rebuilt from the orchestrator's database survives container restarts, caps context growth, and works across harnesses: the three places native resume fails. Native resume still wins for a single interactive session on a machine whose disk persists.

The test that settled it was plain. We restarted the worker container in the middle of a Slack thread, the exact case that used to produce “session not found” or a silent fresh start. The fresh container picked up 8 more follow-up tasks: 8 preamble injections, 0 --resume calls, each against a distinct parent. There was no transcript to lose.

The comparison that matters
  • Container restart. Native resume: “session not found” or a silent fresh start. Preamble: the same brief, rebuilt from the database.
  • Context growth. Native resume: unbounded, ending in SIGTERM-143. Preamble: capped at 2,000 or 4,000 tokens.
  • Cross-harness routing. Native resume: impossible, IDs are per-harness. Preamble: plain text any harness can read.
  • Fidelity. Native resume: the full transcript, noise included. Preamble: a lossy summary, but the loss is bounded and chosen.

Where should session state live?

In the orchestrator's database, as structured rows, not in a harness transcript file. Transcript files die with containers, grow without bound, and only one harness can read them. Database rows are durable, queryable, easy to summarize, and readable by any runner you add later.

That is what cezar's claim gets right, whatever its implementation: switching runners mid-thread only works if the layer above the harness owns the conversation. If the conversation lives in Claude Code's transcript, Claude Code owns your task. If it lives in your database, you own the task, and harnesses become interchangeable executors.

It changes cost too. A resumed session sends its whole history back to the model on every turn, which prompt caching discounts but does not remove. A preamble sends a few thousand tokens. Over a long thread with dozens of follow-ups, that gap adds up.

What does not work

Persistent volumes for transcripts. They fix failure mode 1 and nothing else. You keep unbounded growth, you keep harness lock-in, and you add volume lifecycle and scheduling affinity to your operations work.

Replaying logs verbatim. Pasting raw session_logs rows into the prompt is resume with extra steps: the same growth, the same noise, the same saturation. The 50-row window and the one-line summaries are the feature. Losing detail is the point.

Inheriting the parent's model on derived tasks. Model identifiers are local to a harness. Copy them onto follow-ups and the first cross-harness route fails at session start. Let the assignee pick its model.

Assuming the preamble covers every case. It does not, and the PR says so. The standard preamble only fires for tasks with a parent. A standalone task paused during a graceful shutdown resumes with just its brief and progress notes. Native resume could do better there, but only when the disk survived, which is not the failure we were fixing. Scrubbing has a gap too: today it runs on log lines and steering messages in the resume variant, while the standard preamble copies the parent's brief and output as stored.

The pattern, if you want to copy it

You do not need Agent Swarm for this. It comes down to four decisions:

  • Record what agents do in your own database, as structured rows, at the orchestrator layer. If the only record of a session is the harness transcript, you do not own the session.
  • Rebuild context for every new session from those rows: summarized, aware of ancestors, and hard-capped. Pick a budget and enforce it by cutting the oldest material first.
  • Scrub secrets when you build the preamble. It will cross harness boundaries and land in other vendors' logs. Plan for that.
  • Never let routing inherit harness-local identifiers. Models, session IDs and thread IDs are all per-harness. The router assigns them fresh.

Do those four things and the harness becomes what it should have been all along: a replaceable executor. Restart the container, swap the runner, re-route the task. The conversation survives, because it was never in the harness.

For more on keeping context small on purpose, see context window management for agents. The source for everything above is linked below.

/ references

Sources and further reading

FAQ

Why did Agent Swarm deprecate native session resume?

Native resume points at a transcript file on the worker's disk. When the worker container restarts, that file is gone, and resume either fails with 'session not found' or silently starts a fresh session. Unbounded resumed sessions also saturated context, and a harness-specific session ID cannot follow a task to a different harness.

What is a context preamble for coding agents?

A context preamble is a bounded summary of prior work, injected at the start of a fresh agent session. In Agent Swarm it is capped at 2,000 tokens, covers up to 5 ancestor tasks, and is inserted right after the prompt's leading slash-skill line.

How does the preamble survive container restarts?

The worker rebuilds it on every follow-up task from the parent-task chain stored in the orchestrator's API database. Nothing lives on the worker's disk, so a restarted container builds the same preamble the old one would have.

Can a task move from Claude Code to Codex mid-thread?

Yes. Continuity is plain text built from the orchestrator database, so any harness can read it. Derived tasks deliberately do not inherit the parent's concrete model, because a Claude model string sent to a Codex worker fails at session start.

What token budget should a session preamble use?

Agent Swarm uses 2,000 tokens for follow-up tasks and 4,000 tokens for resuming an interrupted task, both configurable by environment variable. The cap is the point: unbounded resumed sessions are what led to SIGTERM-143 context saturation.

/ keep reading
/ get started

Build your swarm tonight.

Talk with us about Cloud, or fork it on GitHub. Either way, your agents start compounding today.