Skip to the lesson
CivOps AI Academy · F16AI Fleet: The Overnight Build and the Morning Review
0%

Chapter 1 · Session 16

Before you leave: the night plan

An overnight build is decided at five in the afternoon, not at three in the morning. Before you go home, every worker has a brief, a test that decides when it is done, a budget, a list of what it may do, and a reason to stop.

35 minOne page per nightdontAsk + an exact allowlistFour guardrail layersA 30-minute dry run

By the end of this chapter you can

  • Write a one-page night plan that a colleague could run without you.
  • Choose where the fleet runs unattended: a machine you control, or a scheduled CI job.
  • Set the four guardrail layers: spend caps, branch protection, sandbox and permissions.
  • Run a short dry run and fix what it shows before you leave.

Why overnight

A fleet working while the plant team sleeps turns fourteen idle hours into build time, and you spend your working day doing what only you can do: deciding, reviewing and talking to the floor. The price is that nobody is watching. Everything in this element exists to make an unwatched night safe and its results easy to check in the morning [1].

An overnight buildFrom 17:00 to 18:00 the brief and dry run, workers start staggered by 18:30 and build until about 06:00, then morning review and merging until 09:00. One overnight build, plant time (to scale)17:0019:0021:0023:0001:0003:0005:0007:0009:00Brief,dry runWorkers build, test, commit and push in small stepsLastpushesMorning reviewand mergestaggered startNothing deploys and nothing merges while nobody is watching17:00 you write09:00 you decide
One night, drawn to scale. An hour of preparation, about eleven hours of building, two hours of review. Nothing deploys and nothing merges while nobody is watching.

The night plan

Write the night on one page in your repository, docs/nights/YYYY-MM-DD.md, so the plan is versioned with the code it produced. Session 15's wave plan supplies most of it.

SectionWhat it saysExample
GoalWhat the night should deliver, in one sentenceWave 2 of the quality workflow: form, review queue, widget views, access tests
WorkersOne line per worker: brief file, folders it owns, modelagent-3 · briefs/agent-3.md · app/review/ · mid-tier model
Acceptance testsWhich tests decide each task; who wrote themtest/acceptance/review-*.test.mjs, written and reviewed today
CapsPer-run budget and turns; workspace limit$8 and 80 turns per worker; $80 workspace limit this month remaining
Stop rulesWhen a worker must stop on its ownTests fail 3 times in a row; any file outside its folders; no push for 60 minutes
NeverWhat no worker may do tonightMerge, deploy, change acceptance tests, touch package.json, read .env files
MorningWho reviews, in which order, by whenYou, in wave-plan order, 07:00 to 09:00

Where the fleet runs

OptionHowGood forWatch out for
Your own machine or a spare PCA short script starts one headless run per worktree (claude -p with a brief) [2]Full control; local test databasesSleep settings, Windows updates and power cuts; keep it plugged in and awake
Scheduled CI jobA GitHub Actions workflow on a schedule or started by hand runs the Claude Code GitHub Action per task [3]Clean machine every run; nothing on your laptopCI minutes cost money beyond the free allowance; the key is a GitHub secret a person enters on GitHub's site
Cloud agent sessionsYour agent tool's hosted sessions, one per taskNo machine of your ownCheck what each session can reach and what it costs

The four guardrails

Guardrails for an unattended fleetFour nested layers around a worker: spend caps, branch protection, sandbox, permissions. Spend capsworkspace limit · --max-budget-usd · --max-turnsBranch protectionpull requests only · CI required · no force-pushSandboxfile writes in the worktree · network to listed hosts onlyPermissionsdontAsk mode · an exact allowlist · deny rulesOne worker, in its own worktree, on its own branchEach layer stops a different failure; a night needs all four
Four layers, each stopping a different failure. A worker that goes wrong runs into the innermost layer first; if that fails, the next one holds.

1. Permissions: dontAsk with an exact allowlist

Unattended runs cannot answer permission prompts, so decide in advance. Claude Code's dontAsk mode allows only reads and the tools you pre-approve, and denies everything else instead of asking. Deny rules apply in every mode [5][6]. The mode that skips all checks is meant only for isolated containers and virtual machines; never use it on a machine that holds keys or company files [5].

.claude/settings.json in each worker's worktree: what a worker may do tonight
{
  "permissions": {
    "defaultMode": "dontAsk",
    "allow": [
      "Read", "Edit",
      "Bash(npm test *)", "Bash(npm run lint *)",
      "Bash(git status)", "Bash(git diff *)", "Bash(git add *)", "Bash(git commit *)", "Bash(git push origin wave2/*)"
    ],
    "deny": [
      "Read(./.env*)", "Bash(git push --force *)", "Bash(npm install *)", "Bash(curl *)", "Edit(./test/acceptance/**)", "Edit(./package.json)"
    ]
  }
}

2. Sandbox: the worktree and listed hosts only

Claude Code's sandbox runs the agent's shell commands inside a boundary: writes are limited to the working folder, and network traffic goes through a proxy that allows only the hosts you list [7]. Turn it on for unattended work, and list only what the tests need (for example the npm registry if packages are restored, and nothing else).

3. Branch protection: pull requests only

Session 1 protected main: changes arrive by pull request, CI must pass, nobody force-pushes [8]. That is what makes a night recoverable. The worst a worker can do is open a bad pull request, which you close in the morning.

4. Spend caps

Every run gets --max-budget-usd and --max-turns; the fleet's API workspace gets a monthly spend limit (Session 15) [2]. Add the caps together before you leave: the most the night can cost is the sum of the per-run budgets, and you should be comfortable with that number.

The dry run

Thirty minutes before you leave, start every worker with a budget of one dollar. Watch the first few minutes: does each one find its brief, run its tests, commit and push to its own branch? Do the deny rules hold (try asking one to read .env)? Then stop them, delete the dry-run branches, and start the night for real, a minute or two apart.

Terminal: start the night (Linux or macOS; on Windows, run it in WSL)
# Start the night: one bounded headless worker per worktree, staggered. Run from the folder that holds the worktrees.
for N in 2 3 4 5; do
  ( cd agent-$N && claude -p "$(cat ../platform/briefs/agent-$N.md)" \
      --permission-mode dontAsk --max-budget-usd 8 --max-turns 80 \
      --output-format json > ../logs/agent-$N.json 2> ../logs/agent-$N.err ) &
  sleep 120
done
wait; echo "Night finished: $(date)"
Exercise · Write tonight's plan and dry-run it40 minutes

You need: Your wave plan from Session 15; your worktrees; Claude Code or your agent tool; your repository

You will write a night plan, set the guardrails, and prove them with a one-dollar dry run.

Outcome: A night plan in your repository, guardrails you have seen hold, and a known maximum cost for the night.

Knowledge check

Which permission setup fits an unattended worker on your own PC, which also holds company files?

Knowledge check

What is the most a night can cost in model usage with four workers capped at $8 each?

Knowledge check

Why does branch protection make an overnight build recoverable?

References

  1. Anthropic: Claude Code best practices. https://www.anthropic.com/engineering/claude-code-best-practices
  2. Claude Code docs: Run Claude Code programmatically (headless mode). https://code.claude.com/docs/en/headless
  3. Claude Code docs: Claude Code GitHub Actions. https://code.claude.com/docs/en/github-actions
  4. GitHub Docs: About billing for GitHub Actions. https://docs.github.com/en/billing/managing-billing-for-your-products/managing-billing-for-github-actions/about-billing-for-github-actions
  5. Claude Code docs: Permission modes. https://code.claude.com/docs/en/permission-modes
  6. Claude Code docs: Configure permissions. https://code.claude.com/docs/en/permissions
  7. Claude Code docs: Configure the sandboxed Bash tool. https://code.claude.com/docs/en/sandboxing
  8. GitHub Docs: About protected branches. https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-protected-branches/about-protected-branches
  9. OWASP Top 10 for LLM Applications 2025. https://genai.owasp.org/llm-top-10/

Chapter 2 · Running unattended

While you sleep

During the night each worker is in one of a few states. Good workers push small commits, stop themselves when a rule says so, and leave a log that explains why. Nobody needs to be woken up, because nothing irreversible can happen.

30 minPush small, push oftenStop rules, not hopeLogs explain every stopNo 3 a.m. phone calls

By the end of this chapter you can

  • Describe the states a worker moves through overnight and how each one ends.
  • Write stop rules that end a run before it wastes money or does harm.
  • Use pushes and logs as a heartbeat, so a silent worker is obvious in the morning.
  • Recover a worker after a restart without losing work.

A worker's night

Each worker starts queued, then works in a loop: change some code, run its tests, fix, commit, push. It ends in one of two ways: it opens a pull request because its acceptance tests pass, or it stops because a rule told it to. Either way it writes down why.

Worker states overnightQueued, then working in a test-fix-commit-push loop, ending in a pull request or a stop with its reason logged. QueuedWorkingPull requestStoppedDone: waits for morningStopped: reason in its logtest · fix · commit · pushStop rules: budget or turn cap reached · tests fail three times in a row ·a file outside its brief changed · no push for 60 minutes
Worker states overnight. There is no state in which a worker merges or deploys: those wait for a person.

Stop rules

A worker left to itself will keep trying, and every try costs money. Write the rules that end a run into the brief, and back the important ones with hard limits outside the agent:

Stop when…Enforced byWhy
The budget or turn limit is reached--max-budget-usd and --max-turns: the run ends whatever the agent thinks [1]Caps the cost of a task that turned out bigger than planned
The tests fail three times in a row on the same errorThe brief, plus the turn limit as a backstopRepeating the same fix wastes money; a person should look
It wants to change a file outside its foldersDeny rules and the sandbox refuse it; the brief tells it to stop and explain [2]Scope creep collides with other workers
It wants to change an acceptance testA deny rule on test/acceptance/A worker that edits its own test has failed the task
It cannot find something the brief promisedThe brief: write the question in the log and stopGuessing produces confident, wrong work

A heartbeat you can read

Each push to a worker's branch is a heartbeat. Pushing small commits also means a restart loses minutes, not hours. In the morning, the pattern of pushes tells you at a glance which worker kept going, which finished early and which went quiet.

Pushes per hour as a heartbeatFour workers' commits per hour; three push steadily then finish, one stops after two hours. Commits pushed per hour, an example night (each square is one hour from 18:00)agent-2342321agent-323323221agent-431agent-51223223221agent-4 went silent at 20:00: its log shows it stopped after three failed test runsthe others stop pushing when their work is done, then open their pull requests
Commits pushed per hour on an example night. Three workers push steadily and then finish; one went silent after two hours, and its note says why.

You can collect richer signals too. Claude Code can export usage and cost metrics through OpenTelemetry to a monitoring tool you already run, and hooks can run your own script when a session stops or needs attention, for example to append a line to a log [3][4]. Keep it simple at first: pushes, the JSON result of each run (which includes its cost) and the workers' notes are enough [5].

Who gets woken up?

Nobody, if the guardrails are right. Site reliability engineers have a useful rule: alert a person only for something that needs a person now [6]. An overnight build has nothing like that: it cannot touch production, cannot merge and cannot spend past its caps. A worker that stops simply waits for the morning. If you find yourself wanting an alert, a guardrail is missing; add it instead.

When something breaks

When a worker stopsFour reasons a worker stops overnight, what each usually means, and what to do in the morning. A worker stoppedBudget or turn capTask too big or brief unclearsplit it or rewrite the briefTests failed 3 timesRead the last failure in its logfix the brief (or a wrong test)Touched files outside its briefScope creptdiscard; tighten the file listMachine or container restartedWork is safe up to the last pushresume from the branch
When a worker stops. The reason in its log points to the fix; most fixes are a better brief.
  • The machine restarted. Everything up to the last push is safe on the branch. In the morning, start a fresh run in the same worktree with the same brief plus one line: continue from the commits already on this branch.
  • The rate limit was hit. Workers slow down or pause and retry. Stagger the starts more next time, or run fewer workers at once [7].
  • The cache went cold. A worker that waited a long time pays more for its next request. That is expected; it is not worth waking up for.
  • A worker finished early. Good. Do not hand it new work at night: unplanned work is unreviewed work.
Exercise · Break a worker on purpose30 minutes

You need: One worker's worktree; Claude Code or your agent tool; the night plan from exercise 1

You will prove your stop rules work before a real night depends on them.

Outcome: Evidence that your workers stop themselves, explain why, and lose nothing when the machine or network drops.

Knowledge check

A worker has failed the same test three times with the same error at 01:00. What should happen?

Knowledge check

Why push small commits often during the night?

Knowledge check

You want an alarm to wake you if a worker fails overnight. What does this chapter suggest instead?

References

  1. Claude Code docs: CLI reference (--max-budget-usd, --max-turns). https://code.claude.com/docs/en/cli-reference
  2. Claude Code docs: Configure permissions. https://code.claude.com/docs/en/permissions
  3. Claude Code docs: Monitoring (OpenTelemetry). https://code.claude.com/docs/en/monitoring-usage
  4. Claude Code docs: Hooks reference. https://code.claude.com/docs/en/hooks
  5. Claude Code docs: Run Claude Code programmatically (JSON output includes cost). https://code.claude.com/docs/en/headless
  6. Google: Site Reliability Engineering, chapter 6, Monitoring Distributed Systems. https://sre.google/sre-book/monitoring-distributed-systems/
  7. Anthropic: API rate limits. https://platform.claude.com/docs/en/api/rate-limits

Chapter 3 · Review and merge

The morning: checking what it built

The night's output is a set of proposals. The morning decides which become part of the plant's platform. Check in the cheapest order, read every change that survives, try it on a phone, and merge one at a time.

35 minCheapest check firstNever trust a changed testRead every lineRebase, re-run, merge

By the end of this chapter you can

  • Triage a night's pull requests in order, from automatic checks to a careful read.
  • Spot the red flags in an agent's pull request: changed tests, out-of-scope files, oversized diffs, new open routes.
  • Merge in sequence, rebasing and re-running CI before each merge.
  • Turn each failure into a better brief, and record the night in the fleet ledger.

Cheapest check first

Do not start by reading the biggest pull request. Run every pull request through checks in order of cost: the automatic ones first, your careful reading last. A pull request that fails an early check goes back without you spending time reading it.

Morning triageEight pull requests: one fails CI, one changed its acceptance test, leaving six that pass every check. Eight pull requests through the morning checks, cheapest check first (example night)1CI green from a fresh clone?7 of 8 pass2Only the files in its brief?7 of 7 pass3Acceptance tests untouched?6 of 7 pass4Diff readable (under ≈ 400 lines)?6 of 6 pass5Works on the phone preview?6 of 6 passGreen: carry on to the next check. Red: stop there and send it back with a rewritten brief. Six reach a careful read and merge.
Eight pull requests through the morning checks, on an example night. Two fail before anyone reads a line; six reach a careful read.
CheckHowIf it fails
1. CI green from a fresh cloneThe pull request's checks on GitHub [1]Back to the worker with the failure; do not read further
2. Only its own filesThe pull request's Files changed list against the briefDiscard or split; tighten the file list
3. Acceptance tests untouchedNo changes under test/acceptance/ (a deny rule should have stopped it)Discard: a worker that changed its own test has not done the task
4. Readable sizeUnder about 400 changed lines; review quality falls with size [2]Ask for it to be split
5. Works where people use itOpen the pull request's preview address on your phone [3]Back with what you saw

Reading the change

Every change that reaches the plant's platform is read by a person. Google's code review guide puts the reviewer's job simply: approve once the change clearly improves the health of the code, even if it is not perfect, and look at design, function, complexity, tests, naming and comments [4]. For agent work, add the spine's rules:

Look forWhy it matters on the plant platform
A new route or pageIs it closed until the Role and Exposure Matrix opens it (default deny, Session 8)?
A new table or viewDoes it have row-level security from the matrix? Do views run with the viewer's rights (Session 14)?
A new dependencyWas it in the brief? Is it maintained and widely used? Supply-chain failures are an OWASP Top 10 category [5]
Anything that looks like a keyThere should be none. If there is, rotate that key today, whatever else you decide
Tests that only check the happy pathIs there a refused request, an empty result and a wrong input?
Code the brief did not ask forUnasked-for work is unreviewed scope; ask for it to be removed

Merging in sequence

Each merge changes main, so a pull request that passed CI last night may not pass against this morning's main. Merge one at a time, in the order of the wave plan: rebase the next pull request on the new main, let CI run again, then merge [6]. GitHub's merge queue automates the same discipline when many pull requests land at once [7].

Merging in sequencePick a pull request, rebase it on main, let CI run again, merge, and repeat; a red run after rebasing stops the line. Pick the next PRin the order of the wave planRebase on maingit rebase origin/mainCI runs againon the rebased branchMergesquash, then delete the branchrepeat for the next pull request: every merge changes mainRed after rebase? Stop.Never merge two pull requests on the strength of CI runs made before either was merged
Merging in sequence. A red run after rebasing stops the line until it is fixed.
Terminal: rebase, re-test, merge
# Merge the next pull request in sequence (its branch: wave2/agent-3). Run in your main clone.
git fetch origin
git switch wave2/agent-3 && git rebase origin/main     # stop here if there are conflicts
npm ci && npm test                                     # the same checks CI will run
git push --force-with-lease origin wave2/agent-3      # updates the pull request; CI runs again
# When CI is green on GitHub, merge the pull request there (squash) and delete the branch.

Every failure improves the next night

A pull request you send back is information. Most failures trace to the brief: a missing file path, an acceptance test that did not pin the behaviour, a limit that was not stated. Rewrite the brief and keep both versions: Session 16's homework (assignment 4) asks for a brief you rewrote after a bad result, and why.

What went wrongBrief fix for next time
Built the right thing in the wrong folderName the exact folders; add a deny rule for the rest
Passed the tests but missed the pointAdd an acceptance test for the behaviour that was missed
Ran out of budget half waySplit the task; give each half its own test
Added a library nobody asked forState the allowed dependencies; deny npm install
Ignored the phone layoutMake the phone check part of the acceptance tests

Close the night

  1. Write one line per worker in the fleet ledger: model, turns, cost, merged or not, review minutes (Session 15).
  2. Add up the cost per merged change and compare it with last night.
  3. Delete merged branches and worktrees you no longer need: git worktree remove ../agent-3.
  4. Write the next night's plan while the lessons are fresh.
Exercise · Run your first morning review45 minutes

You need: The pull requests from your night (or from exercise 1's dry run); GitHub; your phone; the fleet ledger

You will take a night's pull requests through the checks, merge the good ones in order and rewrite one brief.

Outcome: Merged, reviewed work on main, a rewritten brief for homework 4, and the night's real cost per merged change.

Knowledge check

A pull request's CI is green, but it changed a file under test/acceptance/. What do you do?

Knowledge check

Two pull requests both passed CI last night. Why not merge both straight away?

Knowledge check

In which order should the morning checks run?

References

  1. GitHub Docs: About status checks. https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/collaborating-on-repositories-with-code-quality-features/about-status-checks
  2. SmartBear: Best practices for peer code review. https://smartbear.com/learn/code-review/best-practices-for-peer-code-review/
  3. Vercel docs: Deployments (preview deployments for every pull request). https://vercel.com/docs/deployments
  4. Google Engineering Practices: The standard of code review. https://google.github.io/eng-practices/review/reviewer/standard.html
  5. OWASP Top 10:2025. https://owasp.org/Top10/2025/
  6. Git documentation: git-rebase. https://git-scm.com/docs/git-rebase
  7. GitHub Docs: Managing a merge queue. https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/configuring-pull-request-merges/managing-a-merge-queue

Chapter 4 · 12 questions · 80% passes

Final assessment

Twelve questions across the element. Score 80% (10 of 12) to pass. Your LMS records your score and each answer; you can review the chapters and try again.

15 min12 questions≈ 15 minutesRetake allowed

Choose one answer for each question, then submit. You will see the right answer and why for every question.

1. Where should the night plan live?
2. Which permission mode suits an unattended worker on a machine that holds company files?
3. What does Claude Code's sandbox limit by default?
4. Four workers each run with --max-budget-usd 8. The night's model cost cannot exceed:
5. What is the purpose of the dry run before leaving?
6. Which is a sound stop rule?
7. Why are pushes a useful heartbeat?
8. According to the site reliability rule in this element, when should a person be woken?
9. In the morning triage, which check comes first?
10. A pull request is green but changed a file under test/acceptance/. The right action is:
11. Before merging the second of two green pull requests, you should:
12. A worker built the right feature in the wrong folder. What is the brief fix for next time?