Skip to content
Cameron Mills

All projects

Steward, the managed agent runner

Keeping long AI agent runs alive through usage limits and full contexts

The runner behind every unattended Claude Code job. It winds a run down before the usage limit, resumes the same conversation after the reset, rotates a run whose context is too long into a fresh session, and feeds it all from a timed work queue.

Problem

An unattended Claude Code run can hit the account's usage limit mid-task, be cut off by a sleeping Mac, or run so long that every turn re-sends a huge conversation and cost grows with the square of its length.

Approach

Keep the conversation and a written checkpoint. Every run is told it may be paused at any moment and must keep a five-part checkpoint written for a future self with no memory.

Architecture

Every unattended run goes through one manager, agentctl.py. Each run gets a directory with the event stream, a readable log and the checkpoint, and any abnormal end writes a crash report with the exact resume command. Each outcome (done, wound down, limited, timed out, interrupted) has its own exit code.

Work arrives through a timed queue of new briefs, resumes and job re-kicks. Resumes go first, so a paused build finishes before new work starts. An always-on scheduler runs what is due every minute and writes a heartbeat the health report checks. The /build pipeline launches each of its stages through this same queue.

What I built

  • Wind-down. At 90% of the five-hour usage window the run is stopped and a resume queued for just after the reset, reusing the session id so the whole conversation comes back. The resume message points the agent at its checkpoint, not a re-read of the tree.
  • Context rotation. Past 130k tokens, the session ends at the next tool result and a fresh one starts from the brief and checkpoint alone, at most four times per launch.
  • Liveness and caps. A run counts as alive if the manager, the session or a recent log write says so. A run that writes no checkpoint for 20 minutes is nudged once. The default cap is 90 minutes, and a flag file can lift time caps while keeping the wind-down. The manager is 1,842 lines of Python.

Results

One build was wound down at 99% usage, resumed hours later and finished. Rotation was simulated to save about 12% of spend. It was first tested only on synthetic events. A live planning stage has since rotated at about 137k tokens and finished in its second session.

What didn't work

Five parallel agents once took the usage window from 53% to 100% in ten minutes, so the line now drops four points for each other running run. Runs were marked orphaned while still working, because the sandbox hid their process ids or the session had outlived its manager. Now only a positive dead reading closes a run. Rotation has cut off work in flight, such as subagent results and a test run. The checkpoint now has to carry decisions, file paths and verify commands.

Stack

  • Python
  • Bash
  • launchd
  • Claude Code