Chapter 1 · Session 15
Splitting the build for up to eight agents
Eight agents do not make you eight times faster. They make you faster only when the work is split into tasks that do not touch the same files, in an order that respects what depends on what, and when you can check what comes back.
35 minTasks in wavesOne owner per fileAmdahl's lawYou are the bottleneck
By the end of this chapter you can
- Break one workflow into tasks an agent can finish and prove on its own.
- Arrange tasks into waves by their dependencies, and give every file exactly one owner.
- Use Amdahl's law to estimate what a fleet will really save, and size the fleet to your review time.
- Set up each worker's isolated workspace: its own folder, branch, database and ports.
Where Session 15 starts
By now one agent has built a workflow with you (Session 13) and a dashboard on top of it (Session 14). You have seen how to brief an agent and read what it hands back (Session 12). Session 15 multiplies that: up to eight agents working at the same time, each on its own task. Element F01, chapter 11, explained how a fleet is put together and what a night of it costs. This element is the hands-on version: how to split the work, how to keep each agent's context useful, and how to keep the bill in proportion to what you get.
A good task for a fleet
A fleet task is a brief one agent can finish and prove on its own, without waiting for another agent and without touching files another agent is changing. A good task has five properties:
| Property | Test | Example |
|---|---|---|
| Small | One pull request a person can read in under 30 minutes | The review queue's reject button with its reason field |
| Independent | Needs nothing from another task in the same wave | The operator form does not wait for the dashboard |
| Owned files | Writes only to the folders the brief names | app/review/ and test/review/ |
| Provable | Has acceptance tests written before the work starts | test/review/reject.test.mjs, which the agent may not edit |
| Bounded | Has a turn limit or a budget | Stop after 80 turns or $8, whichever comes first |
Waves: respect the dependencies
Some tasks need others first. Nothing can query a table before the migration that creates it has merged. Draw the tasks as boxes with an arrow from each task to the ones that need it. Tasks with no arrows into them form the first wave; when they merge, the next wave can start [1].
Notice that eight tasks did not become eight agents at once. The widest wave here has four tasks. Starting more agents than a wave has tasks only produces agents waiting, or worse, agents inventing work that collides.
One owner per file
The most common cause of a broken night is two agents writing the same file. Even if each change is right, the second merge conflicts with the first, or quietly undoes it. Give every folder in a wave exactly one owner, and keep the shared files (the package list, shared configuration, the database schema) with the lead: the person or orchestrator agent who merges.
Isolation, set up once per worker
Each worker gets its own copy of the repository on its own branch (a git worktree is the cheapest way: a second folder sharing one repository), its own test database and its own port range, so its tests cannot reset another agent's data or stop another agent's server [2][3]. Claude Code can create worktrees for its sessions and subagents; any agent tool works with a worktree you make by hand [4].
# Wave 2: four workers, each with its own folder, branch, database and ports. Run from the repository folder.
for N in 2 3 4 5; do
git worktree add ../agent-$N -b wave2/agent-$N
createdb "platform_agent_$N"
printf 'DATABASE_URL=postgres://localhost:5432/platform_agent_%s\nPORT=%s\n' "$N" "$((3000 + N * 100))" > ../agent-$N/.env.test
done
git worktree list # four folders, four branchesHow much faster, really?
In 1967 Gene Amdahl pointed out that the speed-up from running work in parallel is capped by the part that stays serial [5]. In a fleet, the serial part is mostly you: reading each pull request, merging them one at a time, and sorting out anything that conflicts. If one fifth of the total effort is serial, eight agents give at best about 3.3 times the speed of one.
The formula: speed-up = 1 ÷ (s + (1 − s) ÷ n), where s is the serial share and n the number of agents. Two practical lessons follow. First, shrink s: small pull requests, tests that decide for you, and a merge order fixed in advance. Second, size the fleet to your review time. If you can carefully review four pull requests in the morning, run four agents, not eight.
You need: Your Session 13 workflow; paper or a Markdown file, docs/waves.md, in your repository; Git
You will turn the next piece of your platform into tasks, waves and file owners, and set up the workspaces.
Outcome: A written wave plan with owners, a fleet size you chose from your own review time, and isolated workspaces ready for Session 16.
Knowledge check
Two tasks in the same wave both need to add a package to package.json. What is the fix?
Knowledge check
Your review and merging are 20% of the total effort. Roughly what speed-up can eight agents give?
Knowledge check
Which task is ready to hand to a fleet worker?
References
- Anthropic: Building effective agents (the orchestrator-workers pattern). https://www.anthropic.com/engineering/building-effective-agents
- Git documentation: git-worktree. https://git-scm.com/docs/git-worktree
- PostgreSQL documentation: createdb. https://www.postgresql.org/docs/current/app-createdb.html
- Claude Code docs: Worktrees. https://code.claude.com/docs/en/worktrees
- Amdahl, G. M. (1967). Validity of the single processor approach to achieving large scale computing capabilities. AFIPS Spring Joint Computer Conference. https://doi.org/10.1145/1465482.1465560
- Anthropic: How we built our multi-agent research system. https://www.anthropic.com/engineering/multi-agent-research-system
Chapter 2 · Context management
Context: what each agent knows
An agent knows only what is in its context window: its instructions, your project's memory file, the brief, and everything it has read so far. Managing that context well makes agents more accurate and cheaper at the same time.
35 minProject memory under 200 linesA stable start gets cachedClear between tasksSubagents keep it clean
By the end of this chapter you can
- Describe what fills an agent's context window and which part the prompt cache can reuse.
- Write a short project memory file that every agent reads, and keep it stable.
- Use clearing, compaction and subagents to keep each agent's context relevant.
- Read an agent's cache statistics and spot what is breaking the cache.
The context window
A model has no memory between requests. Every request an agent makes sends the whole conversation so far: the system instructions and tool definitions, your project's memory file, the brief, every file it has read, every command's output, and every earlier turn. That bundle is the context, and the most a model can take in one request is its context window. Current models offer windows from about 200,000 tokens up to a million [1].
More context is not better
A bigger window does not mean you should fill it. Anthropic's engineering team describes context rot: as the number of tokens in the window grows, a model's ability to recall what is in it accurately goes down [2]. Earlier research found models use information at the start and end of a long input better than information in the middle [3]. For a fleet, the lesson is practical: give each agent what its task needs, and no more.
Project memory: one short file every agent reads
Agent tools read a project memory file at the start of every session: CLAUDE.md for Claude Code, AGENTS.md for many other tools [4][5]. It is where the rules that never change live: how to run the tests, the folder layout, the rules from your matrix and your clamps, what never to do. Claude Code's documentation recommends keeping each such file under about 200 lines, because longer files cost context on every request and are followed less reliably [4].
# Plant platform: rules for every agent
## Run
- npm ci, then npm test (unit and end to end against a fresh local database). Both must pass.
- Verify from a fresh clone before any pull request.
## Never
- Never edit a test under test/acceptance/ (the lead owns them).
- Never put a key in a file, a prompt or a commit. Keys live in Vercel and Supabase settings.
- Never write outside the folders your brief names.
## Rules
- Every table has row-level security generated from docs/matrix.md. A new route is closed until the matrix opens it.
- Every screen passes the phone check at 360 and 390 px: 44 px tap targets, 16 px fields, 15 px text.
- Money is integer cents. Times are stored in UTC and shown in plant time.
## Layout
- app/ screens · lib/ shared rules · db/migrations/ schema · test/ tests · docs/ intent, matrix, measuresKeeping each agent's context relevant
| Technique | What it does | When to use it |
|---|---|---|
| Clear between tasks | Starts a fresh session (/clear in Claude Code); the memory file and brief load again, the old task's files do not | Every time an agent moves to an unrelated task |
| Compaction | Replaces the history with a summary (/compact, with instructions about what to keep) | A long task that must continue, close to the window's limit |
| Subagents | A helper with its own fresh context does a side job (search the code, read a long log) and returns only a short answer | Exploration that would otherwise flood the main context |
| Point, don't paste | Give file paths and let the agent read what it needs, instead of pasting whole files into the brief | Every brief |
| Short tool output | Run tests in a mode that prints failures, not every passing line | Long test suites and build logs |
All five come from the same idea: the agent's working context should hold the task, not the history of everything it has done [7][8][9].
Reading the cache statistics
Claude Code's /usage command shows, for the current session, the share of input tokens served from the cache, the number of cache misses and, where it can tell, the likely cause of the last miss [10]. A healthy agent loop serves most of its input from the cache. If the share drops, look for what changed at the top of the context: a tool or connector added mid-session, a model switch, or a memory file that differs between sessions.
| You see | Likely cause | Fix |
|---|---|---|
| A low cache share from the first requests | Each session starts with different instructions or a different memory file | One memory file for all agents; put task details in the brief, not the memory file |
| Misses after a pause | The cache expired (the standard lifetime is minutes) | Normal after a break; avoid long idle gaps inside a task |
| A miss after adding a tool | Tool definitions are part of the cached prefix | Set up tools before the session starts |
| Requests growing past 150,000 tokens | One session reused for many tasks | Clear between tasks; use subagents for exploration |
You need: Your repository; Claude Code (or another agent tool that reads AGENTS.md); one small task from your wave plan
You will give every future agent the same short memory file, run one task, and read what the cache did.
Outcome: A short, stable memory file in your repository, and your own numbers showing what clearing and caching do.
Knowledge check
Why should the project memory file stay short and identical for every agent?
Knowledge check
An agent must search a large codebase for every place a function is used, then make a small change. How do you keep its context clean?
Knowledge check
An agent's cache share falls sharply right after you connect a new tool mid-session. Why?
References
- Claude Code docs: Explore the context window. https://code.claude.com/docs/en/context-window
- Anthropic: Effective context engineering for AI agents. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
- Liu, N. F. et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. https://arxiv.org/abs/2307.03172
- Claude Code docs: How Claude remembers your project (CLAUDE.md). https://code.claude.com/docs/en/memory
- AGENTS.md: a simple, open format for guiding coding agents. https://agents.md/
- Claude Code docs: How Claude Code uses prompt caching. https://code.claude.com/docs/en/prompt-caching
- Claude Code docs: Manage costs effectively (reduce token usage). https://code.claude.com/docs/en/costs
- Claude Code docs: Subagents. https://code.claude.com/docs/en/sub-agents
- Anthropic: Claude Code best practices. https://www.anthropic.com/engineering/claude-code-best-practices
- Claude Code docs: Manage costs effectively (track your costs with /usage). https://code.claude.com/docs/en/costs#track-your-costs
Chapter 3 · Cost per merged change
Fleet economics
Count the cost of a fleet the way a plant counts cost: per good part. For a fleet the good part is a merged change that passed its tests. Tokens are one line of that cost; your review time is usually the biggest.
30 minCost per merged changeCheapest model that passesHard caps on every runFree way first
By the end of this chapter you can
- Calculate the cost of one merged change, including retries and your review time.
- Match each task to the cheapest model that passes its acceptance checks.
- Cap spending on every run and every workspace, and know what happens when a cap is hit.
- Choose between a subscription and paying per token for your fleet, free way first.
Cost per merged change, not per token
A plant does not judge a press by its electricity bill; it judges cost per good part. Judge a fleet the same way. A night's token bill means little on its own: what matters is how many changes passed their tests, survived your review and merged, and what each one cost in total.
Take the worked night from element F01: eight Sonnet 5.5 workers with good caching cost about $115.78 in tokens, plus about $11.14 for the orchestrator [1]. Suppose six of the eight pull requests merge after review, one needs another night, and one is thrown away.
| Line | Working | Per merged change |
|---|---|---|
| Worker tokens | $115.78 ÷ 6 | $19.30 |
| Orchestrator tokens | $11.14 ÷ 6 | $1.86 |
| Retries and CI allowance | 25% of the token lines | $5.29 |
| Your review | 20 minutes at $60 an hour | $20.00 |
| Total | $46.45 |
Two conclusions. First, the merge rate moves the cost more than the price per token: if only three of eight merge, every change costs twice as much in tokens. Better briefs and tests written first raise the merge rate. Second, your review time is the biggest single line, so make review fast: small pull requests, tests that decide, and a checklist (Session 16).
The cheapest model that passes
Model prices differ by about four times from the small to the largest tier [1]. Most fleet tasks, built from a clear brief with tests written first, do not need the largest model. Run a task on a cheaper model; if it fails its acceptance checks twice, move it up a tier. Keep the largest model for planning, hard bugs and reviewing other agents' work.
Caps on every run
A fleet runs while you are not watching it, so every run needs a hard ceiling. There are three layers, and you want all of them:
| Layer | How (Claude Code and the Claude API as the example) | What happens at the cap |
|---|---|---|
| Per run | Headless runs take --max-budget-usd and --max-turns, for example claude -p --max-budget-usd 8 --max-turns 80 "…" [2] | The run stops; its committed and pushed work is kept |
| Per workspace | On the API, set a spend limit on the workspace the fleet's key belongs to, in the Claude Console [3] | Requests are refused until the limit is raised or the month resets |
| Per plan | On a subscription, the plan's usage windows are the ceiling; usage credits beyond them have their own monthly spend limit [3] | Work pauses until the window resets, or bills credits up to the limit |
# One worker, bounded: a budget, a turn limit and only the tools its task needs. Run inside the worker's worktree.
claude -p "$(cat briefs/agent-3.md)" \
--max-budget-usd 8 --max-turns 80 \
--allowedTools "Read" "Edit" "Bash(npm test *)" "Bash(git add *)" "Bash(git commit *)" "Bash(git push *)" \
--output-format json > logs/agent-3.jsonRate limits are shared
API rate limits (requests and tokens per minute) apply to the whole organisation, not to each agent [4]. Eight agents starting at the same second all ask for their big first request at once. Start workers a minute or two apart, and give the fleet its own workspace so it cannot starve anything else that uses the same account [3].
Subscription or pay per token
| Way to pay | Good for | Watch out for |
|---|---|---|
| Free tiers and local open models | Learning, small experiments, private light tasks | Too small or too slow for an overnight fleet |
| A subscription (for example Claude Pro or Max) | One person running one to a few agents; a predictable monthly price | Usage windows are shared across all your sessions; a fleet hits them fast |
| Pay per token on the API | Fleets, unattended runs and CI; exact caps and per-run cost records | No ceiling unless you set one (set all three layers) |
Prices and plan limits change; check claude.com/pricing and the API pricing page on the day you choose, and record the choice and the reason on your platform's costs page [5][1].
Keep a fleet ledger
Each night, write one line per task: model, turns, tokens, cost from the run's output, merged or not, and your review minutes. After a week you will know your own merge rate and cost per merged change, which beats any estimate in this element. The JSON output of a headless run includes the run's cost, so a short script can fill in most of the line [6].
You need: Your wave plan; a spreadsheet or docs/fleet-ledger.md; Claude Code or your agent tool; the Claude Console if you pay per token
You will set caps, run two workers on real tasks from your wave plan, and work out your own cost per merged change.
Outcome: Caps at every layer, a fleet ledger with real numbers, and your own cost per merged change.
Knowledge check
Eight pull requests cost $120 in tokens; three merge. Which change lowers the cost per merged change most?
Knowledge check
A task fails its acceptance tests twice on the mid-tier model. What does this chapter suggest?
Knowledge check
Why start fleet workers a minute or two apart?
References
- Anthropic: Claude pricing. https://platform.claude.com/docs/en/about-claude/pricing
- Claude Code docs: CLI reference (--max-budget-usd, --max-turns, --allowedTools). https://code.claude.com/docs/en/cli-reference
- Claude Code docs: Manage costs effectively (workspace spend limits, usage credits). https://code.claude.com/docs/en/costs
- Anthropic: API rate limits. https://platform.claude.com/docs/en/api/rate-limits
- Claude plans and pricing. https://claude.com/pricing
- Claude Code docs: Run Claude Code programmatically (headless, JSON output). https://code.claude.com/docs/en/headless
Chapter 4 · 12 questions · 80% passes
Final assessment
Twelve questions across the element. Score 80% (10 of 12) to pass. Your LMS records your score and each answer; you can review the chapters and try again.
15 min12 questions≈ 15 minutesRetake allowed
Your result
CivOps AI Academy
AI Fleet: Splitting the Build, Managing Context, Counting the Cost
Element F15 complete · Learner
Your LMS records this completion. For the CivOps Foundation certificate, finish the Foundation Course at https://civops.io/learn.