Skip to the lesson
CivOps AI Academy · F15AI Fleet: Splitting the Build, Managing Context, Counting the Cost
0%

Chapter 1 · Session 15

Splitting the build for up to eight agents

Eight agents do not make you eight times faster. They make you faster only when the work is split into tasks that do not touch the same files, in an order that respects what depends on what, and when you can check what comes back.

35 minTasks in wavesOne owner per fileAmdahl's lawYou are the bottleneck

By the end of this chapter you can

  • Break one workflow into tasks an agent can finish and prove on its own.
  • Arrange tasks into waves by their dependencies, and give every file exactly one owner.
  • Use Amdahl's law to estimate what a fleet will really save, and size the fleet to your review time.
  • Set up each worker's isolated workspace: its own folder, branch, database and ports.

Where Session 15 starts

By now one agent has built a workflow with you (Session 13) and a dashboard on top of it (Session 14). You have seen how to brief an agent and read what it hands back (Session 12). Session 15 multiplies that: up to eight agents working at the same time, each on its own task. Element F01, chapter 11, explained how a fleet is put together and what a night of it costs. This element is the hands-on version: how to split the work, how to keep each agent's context useful, and how to keep the bill in proportion to what you get.

A good task for a fleet

A fleet task is a brief one agent can finish and prove on its own, without waiting for another agent and without touching files another agent is changing. A good task has five properties:

PropertyTestExample
SmallOne pull request a person can read in under 30 minutesThe review queue's reject button with its reason field
IndependentNeeds nothing from another task in the same waveThe operator form does not wait for the dashboard
Owned filesWrites only to the folders the brief namesapp/review/ and test/review/
ProvableHas acceptance tests written before the work startstest/review/reject.test.mjs, which the agent may not edit
BoundedHas a turn limit or a budgetStop after 80 turns or $8, whichever comes first

Waves: respect the dependencies

Some tasks need others first. Nothing can query a table before the migration that creates it has merged. Draw the tasks as boxes with an arrow from each task to the ones that need it. Tasks with no arrows into them form the first wave; when they merge, the next wave can start [1].

Tasks in wavesEight tasks in three waves: the schema first, then four tasks in parallel, then three, each arrow a dependency. Wave 1 · foundationsT1schema migrationWave 2 · in parallelT2operator formT3review queueT4widget viewsT5access testsWave 3 · in parallelT6manager dashboardT7phone audit fixesT8runbook pageA task starts when the tasks it depends on have merged; inside a wave, no two tasks touch the same files
A workflow split into eight tasks and three waves. The schema goes first and alone. Four tasks then run in parallel, then three more.

Notice that eight tasks did not become eight agents at once. The widest wave here has four tasks. Starting more agents than a wave has tasks only produces agents waiting, or worse, agents inventing work that collides.

One owner per file

The most common cause of a broken night is two agents writing the same file. Even if each change is right, the second merge conflicts with the first, or quietly undoes it. Give every folder in a wave exactly one owner, and keep the shared files (the package list, shared configuration, the database schema) with the lead: the person or orchestrator agent who merges.

File ownership across agentsA grid of four agents by six folders: each folder written by exactly one agent, package.json by the lead only. File ownership for wave 2: each folder has one owner; shared files belong to the leadagent-1agent-2agent-3agent-4db/migrations/writesapp/operator/writesapp/review/writeslib/widgets/writestest/access/writespackage.jsonlead only: dependencies and shared configurationTwo agents writing one file is the most common cause of a broken night
File ownership for wave 2. Each agent writes only its own folders; a dependency change goes through the lead.

Isolation, set up once per worker

Each worker gets its own copy of the repository on its own branch (a git worktree is the cheapest way: a second folder sharing one repository), its own test database and its own port range, so its tests cannot reset another agent's data or stop another agent's server [2][3]. Claude Code can create worktrees for its sessions and subagents; any agent tool works with a worktree you make by hand [4].

Terminal: four isolated workspaces (Postgres on your laptop; test data only)
# Wave 2: four workers, each with its own folder, branch, database and ports. Run from the repository folder.
for N in 2 3 4 5; do
  git worktree add ../agent-$N -b wave2/agent-$N
  createdb "platform_agent_$N"
  printf 'DATABASE_URL=postgres://localhost:5432/platform_agent_%s\nPORT=%s\n' "$N" "$((3000 + N * 100))" > ../agent-$N/.env.test
done
git worktree list   # four folders, four branches

How much faster, really?

In 1967 Gene Amdahl pointed out that the speed-up from running work in parallel is capped by the part that stays serial [5]. In a fleet, the serial part is mostly you: reading each pull request, merging them one at a time, and sorting out anything that conflicts. If one fifth of the total effort is serial, eight agents give at best about 3.3 times the speed of one.

Amdahl's law for an agent fleetSpeed-up at eight agents is 5.9 times with 5% serial work, 3.3 times with 20%, and 2.1 times with 40%. Speed-up with n agents when a share of the work is serial (review, merge, fixing conflicts)1×2×4×6×8×12345678agents working in parallel5% serial: 5.9×20% serial: 3.3×40% serial: 2.1×dashed: perfect 8× (never happens)
Amdahl's law for a fleet, drawn to scale. The serial share, mostly your review and merging, decides the speed-up far more than the number of agents.

The formula: speed-up = 1 ÷ (s + (1 − s) ÷ n), where s is the serial share and n the number of agents. Two practical lessons follow. First, shrink s: small pull requests, tests that decide for you, and a merge order fixed in advance. Second, size the fleet to your review time. If you can carefully review four pull requests in the morning, run four agents, not eight.

Exercise · Split one workflow into a wave plan35 minutes

You need: Your Session 13 workflow; paper or a Markdown file, docs/waves.md, in your repository; Git

You will turn the next piece of your platform into tasks, waves and file owners, and set up the workspaces.

Outcome: A written wave plan with owners, a fleet size you chose from your own review time, and isolated workspaces ready for Session 16.

Knowledge check

Two tasks in the same wave both need to add a package to package.json. What is the fix?

Knowledge check

Your review and merging are 20% of the total effort. Roughly what speed-up can eight agents give?

Knowledge check

Which task is ready to hand to a fleet worker?

References

  1. Anthropic: Building effective agents (the orchestrator-workers pattern). https://www.anthropic.com/engineering/building-effective-agents
  2. Git documentation: git-worktree. https://git-scm.com/docs/git-worktree
  3. PostgreSQL documentation: createdb. https://www.postgresql.org/docs/current/app-createdb.html
  4. Claude Code docs: Worktrees. https://code.claude.com/docs/en/worktrees
  5. Amdahl, G. M. (1967). Validity of the single processor approach to achieving large scale computing capabilities. AFIPS Spring Joint Computer Conference. https://doi.org/10.1145/1465482.1465560
  6. Anthropic: How we built our multi-agent research system. https://www.anthropic.com/engineering/multi-agent-research-system

Chapter 2 · Context management

Context: what each agent knows

An agent knows only what is in its context window: its instructions, your project's memory file, the brief, and everything it has read so far. Managing that context well makes agents more accurate and cheaper at the same time.

35 minProject memory under 200 linesA stable start gets cachedClear between tasksSubagents keep it clean

By the end of this chapter you can

  • Describe what fills an agent's context window and which part the prompt cache can reuse.
  • Write a short project memory file that every agent reads, and keep it stable.
  • Use clearing, compaction and subagents to keep each agent's context relevant.
  • Read an agent's cache statistics and spot what is breaking the cache.

The context window

A model has no memory between requests. Every request an agent makes sends the whole conversation so far: the system instructions and tool definitions, your project's memory file, the brief, every file it has read, every command's output, and every earlier turn. That bundle is the context, and the most a model can take in one request is its context window. Current models offer windows from about 200,000 tokens up to a million [1].

What fills the context windowStacked from top: system prompt and tools, project memory, the brief, accumulated files and turns, and new tool results. One request from an agent 40 turns into a task (about 90,000 tokens), drawn to scaleSystem prompt and tool definitionsstable: cachedProject memory (CLAUDE.md or AGENTS.md)stable: cachedThe briefstable: cachedFiles read, command output, earlier turnsgrows: cached once sentThis turn's new tool resultsnew: written to the cacheKeep the top stable (same instructions, same order) and most of every request is billed at the cache-read price
What fills one request, drawn to scale. The top three layers are the same on every request, so the prompt cache serves them at a fraction of the price.

More context is not better

A bigger window does not mean you should fill it. Anthropic's engineering team describes context rot: as the number of tokens in the window grows, a model's ability to recall what is in it accurately goes down [2]. Earlier research found models use information at the start and end of a long input better than information in the middle [3]. For a fleet, the lesson is practical: give each agent what its task needs, and no more.

Project memory: one short file every agent reads

Agent tools read a project memory file at the start of every session: CLAUDE.md for Claude Code, AGENTS.md for many other tools [4][5]. It is where the rules that never change live: how to run the tests, the folder layout, the rules from your matrix and your clamps, what never to do. Claude Code's documentation recommends keeping each such file under about 200 lines, because longer files cost context on every request and are followed less reliably [4].

CLAUDE.md (or AGENTS.md): short, stable, the same for every agent
# Plant platform: rules for every agent
## Run
- npm ci, then npm test (unit and end to end against a fresh local database). Both must pass.
- Verify from a fresh clone before any pull request.
## Never
- Never edit a test under test/acceptance/ (the lead owns them).
- Never put a key in a file, a prompt or a commit. Keys live in Vercel and Supabase settings.
- Never write outside the folders your brief names.
## Rules
- Every table has row-level security generated from docs/matrix.md. A new route is closed until the matrix opens it.
- Every screen passes the phone check at 360 and 390 px: 44 px tap targets, 16 px fields, 15 px text.
- Money is integer cents. Times are stored in UTC and shown in plant time.
## Layout
- app/ screens · lib/ shared rules · db/migrations/ schema · test/ tests · docs/ intent, matrix, measures

Keeping each agent's context relevant

TechniqueWhat it doesWhen to use it
Clear between tasksStarts a fresh session (/clear in Claude Code); the memory file and brief load again, the old task's files do notEvery time an agent moves to an unrelated task
CompactionReplaces the history with a summary (/compact, with instructions about what to keep)A long task that must continue, close to the window's limit
SubagentsA helper with its own fresh context does a side job (search the code, read a long log) and returns only a short answerExploration that would otherwise flood the main context
Point, don't pasteGive file paths and let the agent read what it needs, instead of pasting whole files into the briefEvery brief
Short tool outputRun tests in a mode that prints failures, not every passing lineLong test suites and build logs

All five come from the same idea: the agent's working context should hold the task, not the history of everything it has done [7][8][9].

Context growth and clearingSixty requests: one long session grows to about 170 thousand tokens and costs $2.63; clearing between three tasks keeps it under 70 thousand and costs $2.03. Context size per request over three tasks (thousands of tokens)0k60k120k180k1214160request numberone long session: $2.63cleared per task: $2.03
Context per request over three tasks, drawn to scale. Clearing between tasks keeps requests small, which is cheaper and, because of context rot, more accurate.

Reading the cache statistics

Claude Code's /usage command shows, for the current session, the share of input tokens served from the cache, the number of cache misses and, where it can tell, the likely cause of the last miss [10]. A healthy agent loop serves most of its input from the cache. If the share drops, look for what changed at the top of the context: a tool or connector added mid-session, a model switch, or a memory file that differs between sessions.

You seeLikely causeFix
A low cache share from the first requestsEach session starts with different instructions or a different memory fileOne memory file for all agents; put task details in the brief, not the memory file
Misses after a pauseThe cache expired (the standard lifetime is minutes)Normal after a break; avoid long idle gaps inside a task
A miss after adding a toolTool definitions are part of the cached prefixSet up tools before the session starts
Requests growing past 150,000 tokensOne session reused for many tasksClear between tasks; use subagents for exploration
Exercise · Write the memory file and measure the cache35 minutes

You need: Your repository; Claude Code (or another agent tool that reads AGENTS.md); one small task from your wave plan

You will give every future agent the same short memory file, run one task, and read what the cache did.

Outcome: A short, stable memory file in your repository, and your own numbers showing what clearing and caching do.

Knowledge check

Why should the project memory file stay short and identical for every agent?

Knowledge check

An agent must search a large codebase for every place a function is used, then make a small change. How do you keep its context clean?

Knowledge check

An agent's cache share falls sharply right after you connect a new tool mid-session. Why?

References

  1. Claude Code docs: Explore the context window. https://code.claude.com/docs/en/context-window
  2. Anthropic: Effective context engineering for AI agents. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
  3. Liu, N. F. et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. https://arxiv.org/abs/2307.03172
  4. Claude Code docs: How Claude remembers your project (CLAUDE.md). https://code.claude.com/docs/en/memory
  5. AGENTS.md: a simple, open format for guiding coding agents. https://agents.md/
  6. Claude Code docs: How Claude Code uses prompt caching. https://code.claude.com/docs/en/prompt-caching
  7. Claude Code docs: Manage costs effectively (reduce token usage). https://code.claude.com/docs/en/costs
  8. Claude Code docs: Subagents. https://code.claude.com/docs/en/sub-agents
  9. Anthropic: Claude Code best practices. https://www.anthropic.com/engineering/claude-code-best-practices
  10. Claude Code docs: Manage costs effectively (track your costs with /usage). https://code.claude.com/docs/en/costs#track-your-costs

Chapter 3 · Cost per merged change

Fleet economics

Count the cost of a fleet the way a plant counts cost: per good part. For a fleet the good part is a merged change that passed its tests. Tokens are one line of that cost; your review time is usually the biggest.

30 minCost per merged changeCheapest model that passesHard caps on every runFree way first

By the end of this chapter you can

  • Calculate the cost of one merged change, including retries and your review time.
  • Match each task to the cheapest model that passes its acceptance checks.
  • Cap spending on every run and every workspace, and know what happens when a cap is hit.
  • Choose between a subscription and paying per token for your fleet, free way first.

Cost per merged change, not per token

A plant does not judge a press by its electricity bill; it judges cost per good part. Judge a fleet the same way. A night's token bill means little on its own: what matters is how many changes passed their tests, survived your review and merged, and what each one cost in total.

Take the worked night from element F01: eight Sonnet 5.5 workers with good caching cost about $115.78 in tokens, plus about $11.14 for the orchestrator [1]. Suppose six of the eight pull requests merge after review, one needs another night, and one is thrown away.

Cost of one merged changeTokens, workers (Sonnet 5.5, cached): $19.3; Tokens, orchestrator share: $1.86; Retries and CI allowance (25%): $5.29; Your review, 20 min at $60 an hour: $20 One merged change costs about $46, and the biggest single line is your own review timeTokens, workers (Sonnet 5.5, cached)$19.30Tokens, orchestrator share$1.86Retries and CI allowance (25%)$5.29Your review, 20 min at $60 an hour$20.00
One merged change, worked through and drawn to scale. Tokens ÷ 6 merged changes, plus a 25% allowance for retries and CI, plus 20 minutes of your review at a $60-an-hour loaded labour rate.
LineWorkingPer merged change
Worker tokens$115.78 ÷ 6$19.30
Orchestrator tokens$11.14 ÷ 6$1.86
Retries and CI allowance25% of the token lines$5.29
Your review20 minutes at $60 an hour$20.00
Total$46.45

Two conclusions. First, the merge rate moves the cost more than the price per token: if only three of eight merge, every change costs twice as much in tokens. Better briefs and tests written first raise the merge rate. Second, your review time is the biggest single line, so make review fast: small pull requests, tests that decide, and a checklist (Session 16).

The cheapest model that passes

Model prices differ by about four times from the small to the largest tier [1]. Most fleet tasks, built from a clear brief with tests written first, do not need the largest model. Run a task on a cheaper model; if it fails its acceptance checks twice, move it up a tier. Keep the largest model for planning, hard bugs and reviewing other agents' work.

Matching the model to the taskThree tiers of model, small, mid and largest, each with the tasks it suits and its list price. Pick the cheapest model that passes the task's acceptance checksSmall, fast modelHaiku-class · $1 in / $5 outrename, format, lint fixessummarise logs and test outputfirst pass over a long fileMid-tier modelSonnet-class · $2 in / $10 outmost build tasks from a clear brieftests, views, widgets, formsthe default for workersLargest modelOpus-class · $4 in / $20 outplanning and splitting the workhard bugs and design choicesreviewing other agents' workList prices per million tokens, Anthropic, 3 October 2026; other providers have similar tiers
Matching the model to the task. The acceptance tests decide whether the cheaper model was good enough.

Caps on every run

A fleet runs while you are not watching it, so every run needs a hard ceiling. There are three layers, and you want all of them:

LayerHow (Claude Code and the Claude API as the example)What happens at the cap
Per runHeadless runs take --max-budget-usd and --max-turns, for example claude -p --max-budget-usd 8 --max-turns 80 "…" [2]The run stops; its committed and pushed work is kept
Per workspaceOn the API, set a spend limit on the workspace the fleet's key belongs to, in the Claude Console [3]Requests are refused until the limit is raised or the month resets
Per planOn a subscription, the plan's usage windows are the ceiling; usage credits beyond them have their own monthly spend limit [3]Work pauses until the window resets, or bills credits up to the limit
Terminal: a bounded worker run
# One worker, bounded: a budget, a turn limit and only the tools its task needs. Run inside the worker's worktree.
claude -p "$(cat briefs/agent-3.md)" \
  --max-budget-usd 8 --max-turns 80 \
  --allowedTools "Read" "Edit" "Bash(npm test *)" "Bash(git add *)" "Bash(git commit *)" "Bash(git push *)" \
  --output-format json > logs/agent-3.json

Rate limits are shared

API rate limits (requests and tokens per minute) apply to the whole organisation, not to each agent [4]. Eight agents starting at the same second all ask for their big first request at once. Start workers a minute or two apart, and give the fleet its own workspace so it cannot starve anything else that uses the same account [3].

Subscription or pay per token

Way to payGood forWatch out for
Free tiers and local open modelsLearning, small experiments, private light tasksToo small or too slow for an overnight fleet
A subscription (for example Claude Pro or Max)One person running one to a few agents; a predictable monthly priceUsage windows are shared across all your sessions; a fleet hits them fast
Pay per token on the APIFleets, unattended runs and CI; exact caps and per-run cost recordsNo ceiling unless you set one (set all three layers)

Prices and plan limits change; check claude.com/pricing and the API pricing page on the day you choose, and record the choice and the reason on your platform's costs page [5][1].

Keep a fleet ledger

Each night, write one line per task: model, turns, tokens, cost from the run's output, merged or not, and your review minutes. After a week you will know your own merge rate and cost per merged change, which beats any estimate in this element. The JSON output of a headless run includes the run's cost, so a short script can fill in most of the line [6].

Exercise · Price your fleet and run two bounded workers40 minutes

You need: Your wave plan; a spreadsheet or docs/fleet-ledger.md; Claude Code or your agent tool; the Claude Console if you pay per token

You will set caps, run two workers on real tasks from your wave plan, and work out your own cost per merged change.

Outcome: Caps at every layer, a fleet ledger with real numbers, and your own cost per merged change.

Knowledge check

Eight pull requests cost $120 in tokens; three merge. Which change lowers the cost per merged change most?

Knowledge check

A task fails its acceptance tests twice on the mid-tier model. What does this chapter suggest?

Knowledge check

Why start fleet workers a minute or two apart?

References

  1. Anthropic: Claude pricing. https://platform.claude.com/docs/en/about-claude/pricing
  2. Claude Code docs: CLI reference (--max-budget-usd, --max-turns, --allowedTools). https://code.claude.com/docs/en/cli-reference
  3. Claude Code docs: Manage costs effectively (workspace spend limits, usage credits). https://code.claude.com/docs/en/costs
  4. Anthropic: API rate limits. https://platform.claude.com/docs/en/api/rate-limits
  5. Claude plans and pricing. https://claude.com/pricing
  6. Claude Code docs: Run Claude Code programmatically (headless, JSON output). https://code.claude.com/docs/en/headless

Chapter 4 · 12 questions · 80% passes

Final assessment

Twelve questions across the element. Score 80% (10 of 12) to pass. Your LMS records your score and each answer; you can review the chapters and try again.

15 min12 questions≈ 15 minutesRetake allowed

Choose one answer for each question, then submit. You will see the right answer and why for every question.

1. What makes a task suitable for a fleet worker?
2. Which tasks form the first wave?
3. Who should own package.json and shared configuration during a wave?
4. With a 5% serial share, Amdahl's law gives eight agents a speed-up of about:
5. How should you size a fleet?
6. What is 'context rot'?
7. Claude Code's documentation recommends keeping each CLAUDE.md file under about:
8. Which order keeps the prompt cache effective?
9. When should an agent's session be cleared?
10. In the worked example, what is the largest single line of a merged change's cost?
11. Which flags bound a headless Claude Code worker run?
12. A task fails its acceptance tests twice on a mid-tier model. What should not happen?