Chapter 1 · Session 16
Before you leave: the night plan
An overnight build is decided at five in the afternoon, not at three in the morning. Before you go home, every worker has a brief, a test that decides when it is done, a budget, a list of what it may do, and a reason to stop.
35 minOne page per nightdontAsk + an exact allowlistFour guardrail layersA 30-minute dry run
By the end of this chapter you can
- Write a one-page night plan that a colleague could run without you.
- Choose where the fleet runs unattended: a machine you control, or a scheduled CI job.
- Set the four guardrail layers: spend caps, branch protection, sandbox and permissions.
- Run a short dry run and fix what it shows before you leave.
Why overnight
A fleet working while the plant team sleeps turns fourteen idle hours into build time, and you spend your working day doing what only you can do: deciding, reviewing and talking to the floor. The price is that nobody is watching. Everything in this element exists to make an unwatched night safe and its results easy to check in the morning [1].
The night plan
Write the night on one page in your repository, docs/nights/YYYY-MM-DD.md, so the plan is versioned with the code it produced. Session 15's wave plan supplies most of it.
| Section | What it says | Example |
|---|---|---|
| Goal | What the night should deliver, in one sentence | Wave 2 of the quality workflow: form, review queue, widget views, access tests |
| Workers | One line per worker: brief file, folders it owns, model | agent-3 · briefs/agent-3.md · app/review/ · mid-tier model |
| Acceptance tests | Which tests decide each task; who wrote them | test/acceptance/review-*.test.mjs, written and reviewed today |
| Caps | Per-run budget and turns; workspace limit | $8 and 80 turns per worker; $80 workspace limit this month remaining |
| Stop rules | When a worker must stop on its own | Tests fail 3 times in a row; any file outside its folders; no push for 60 minutes |
| Never | What no worker may do tonight | Merge, deploy, change acceptance tests, touch package.json, read .env files |
| Morning | Who reviews, in which order, by when | You, in wave-plan order, 07:00 to 09:00 |
Where the fleet runs
| Option | How | Good for | Watch out for |
|---|---|---|---|
| Your own machine or a spare PC | A short script starts one headless run per worktree (claude -p with a brief) [2] | Full control; local test databases | Sleep settings, Windows updates and power cuts; keep it plugged in and awake |
| Scheduled CI job | A GitHub Actions workflow on a schedule or started by hand runs the Claude Code GitHub Action per task [3] | Clean machine every run; nothing on your laptop | CI minutes cost money beyond the free allowance; the key is a GitHub secret a person enters on GitHub's site |
| Cloud agent sessions | Your agent tool's hosted sessions, one per task | No machine of your own | Check what each session can reach and what it costs |
The four guardrails
1. Permissions: dontAsk with an exact allowlist
Unattended runs cannot answer permission prompts, so decide in advance. Claude Code's dontAsk mode allows only reads and the tools you pre-approve, and denies everything else instead of asking. Deny rules apply in every mode [5][6]. The mode that skips all checks is meant only for isolated containers and virtual machines; never use it on a machine that holds keys or company files [5].
{
"permissions": {
"defaultMode": "dontAsk",
"allow": [
"Read", "Edit",
"Bash(npm test *)", "Bash(npm run lint *)",
"Bash(git status)", "Bash(git diff *)", "Bash(git add *)", "Bash(git commit *)", "Bash(git push origin wave2/*)"
],
"deny": [
"Read(./.env*)", "Bash(git push --force *)", "Bash(npm install *)", "Bash(curl *)", "Edit(./test/acceptance/**)", "Edit(./package.json)"
]
}
}2. Sandbox: the worktree and listed hosts only
Claude Code's sandbox runs the agent's shell commands inside a boundary: writes are limited to the working folder, and network traffic goes through a proxy that allows only the hosts you list [7]. Turn it on for unattended work, and list only what the tests need (for example the npm registry if packages are restored, and nothing else).
3. Branch protection: pull requests only
Session 1 protected main: changes arrive by pull request, CI must pass, nobody force-pushes [8]. That is what makes a night recoverable. The worst a worker can do is open a bad pull request, which you close in the morning.
4. Spend caps
Every run gets --max-budget-usd and --max-turns; the fleet's API workspace gets a monthly spend limit (Session 15) [2]. Add the caps together before you leave: the most the night can cost is the sum of the per-run budgets, and you should be comfortable with that number.
The dry run
Thirty minutes before you leave, start every worker with a budget of one dollar. Watch the first few minutes: does each one find its brief, run its tests, commit and push to its own branch? Do the deny rules hold (try asking one to read .env)? Then stop them, delete the dry-run branches, and start the night for real, a minute or two apart.
# Start the night: one bounded headless worker per worktree, staggered. Run from the folder that holds the worktrees.
for N in 2 3 4 5; do
( cd agent-$N && claude -p "$(cat ../platform/briefs/agent-$N.md)" \
--permission-mode dontAsk --max-budget-usd 8 --max-turns 80 \
--output-format json > ../logs/agent-$N.json 2> ../logs/agent-$N.err ) &
sleep 120
done
wait; echo "Night finished: $(date)"You need: Your wave plan from Session 15; your worktrees; Claude Code or your agent tool; your repository
You will write a night plan, set the guardrails, and prove them with a one-dollar dry run.
Outcome: A night plan in your repository, guardrails you have seen hold, and a known maximum cost for the night.
Knowledge check
Which permission setup fits an unattended worker on your own PC, which also holds company files?
Knowledge check
What is the most a night can cost in model usage with four workers capped at $8 each?
Knowledge check
Why does branch protection make an overnight build recoverable?
References
- Anthropic: Claude Code best practices. https://www.anthropic.com/engineering/claude-code-best-practices
- Claude Code docs: Run Claude Code programmatically (headless mode). https://code.claude.com/docs/en/headless
- Claude Code docs: Claude Code GitHub Actions. https://code.claude.com/docs/en/github-actions
- GitHub Docs: About billing for GitHub Actions. https://docs.github.com/en/billing/managing-billing-for-your-products/managing-billing-for-github-actions/about-billing-for-github-actions
- Claude Code docs: Permission modes. https://code.claude.com/docs/en/permission-modes
- Claude Code docs: Configure permissions. https://code.claude.com/docs/en/permissions
- Claude Code docs: Configure the sandboxed Bash tool. https://code.claude.com/docs/en/sandboxing
- GitHub Docs: About protected branches. https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-protected-branches/about-protected-branches
- OWASP Top 10 for LLM Applications 2025. https://genai.owasp.org/llm-top-10/
Chapter 2 · Running unattended
While you sleep
During the night each worker is in one of a few states. Good workers push small commits, stop themselves when a rule says so, and leave a log that explains why. Nobody needs to be woken up, because nothing irreversible can happen.
30 minPush small, push oftenStop rules, not hopeLogs explain every stopNo 3 a.m. phone calls
By the end of this chapter you can
- Describe the states a worker moves through overnight and how each one ends.
- Write stop rules that end a run before it wastes money or does harm.
- Use pushes and logs as a heartbeat, so a silent worker is obvious in the morning.
- Recover a worker after a restart without losing work.
A worker's night
Each worker starts queued, then works in a loop: change some code, run its tests, fix, commit, push. It ends in one of two ways: it opens a pull request because its acceptance tests pass, or it stops because a rule told it to. Either way it writes down why.
Stop rules
A worker left to itself will keep trying, and every try costs money. Write the rules that end a run into the brief, and back the important ones with hard limits outside the agent:
| Stop when… | Enforced by | Why |
|---|---|---|
| The budget or turn limit is reached | --max-budget-usd and --max-turns: the run ends whatever the agent thinks [1] | Caps the cost of a task that turned out bigger than planned |
| The tests fail three times in a row on the same error | The brief, plus the turn limit as a backstop | Repeating the same fix wastes money; a person should look |
| It wants to change a file outside its folders | Deny rules and the sandbox refuse it; the brief tells it to stop and explain [2] | Scope creep collides with other workers |
| It wants to change an acceptance test | A deny rule on test/acceptance/ | A worker that edits its own test has failed the task |
| It cannot find something the brief promised | The brief: write the question in the log and stop | Guessing produces confident, wrong work |
A heartbeat you can read
Each push to a worker's branch is a heartbeat. Pushing small commits also means a restart loses minutes, not hours. In the morning, the pattern of pushes tells you at a glance which worker kept going, which finished early and which went quiet.
You can collect richer signals too. Claude Code can export usage and cost metrics through OpenTelemetry to a monitoring tool you already run, and hooks can run your own script when a session stops or needs attention, for example to append a line to a log [3][4]. Keep it simple at first: pushes, the JSON result of each run (which includes its cost) and the workers' notes are enough [5].
Who gets woken up?
Nobody, if the guardrails are right. Site reliability engineers have a useful rule: alert a person only for something that needs a person now [6]. An overnight build has nothing like that: it cannot touch production, cannot merge and cannot spend past its caps. A worker that stops simply waits for the morning. If you find yourself wanting an alert, a guardrail is missing; add it instead.
When something breaks
- The machine restarted. Everything up to the last push is safe on the branch. In the morning, start a fresh run in the same worktree with the same brief plus one line: continue from the commits already on this branch.
- The rate limit was hit. Workers slow down or pause and retry. Stagger the starts more next time, or run fewer workers at once [7].
- The cache went cold. A worker that waited a long time pays more for its next request. That is expected; it is not worth waking up for.
- A worker finished early. Good. Do not hand it new work at night: unplanned work is unreviewed work.
You need: One worker's worktree; Claude Code or your agent tool; the night plan from exercise 1
You will prove your stop rules work before a real night depends on them.
Outcome: Evidence that your workers stop themselves, explain why, and lose nothing when the machine or network drops.
Knowledge check
A worker has failed the same test three times with the same error at 01:00. What should happen?
Knowledge check
Why push small commits often during the night?
Knowledge check
You want an alarm to wake you if a worker fails overnight. What does this chapter suggest instead?
References
- Claude Code docs: CLI reference (--max-budget-usd, --max-turns). https://code.claude.com/docs/en/cli-reference
- Claude Code docs: Configure permissions. https://code.claude.com/docs/en/permissions
- Claude Code docs: Monitoring (OpenTelemetry). https://code.claude.com/docs/en/monitoring-usage
- Claude Code docs: Hooks reference. https://code.claude.com/docs/en/hooks
- Claude Code docs: Run Claude Code programmatically (JSON output includes cost). https://code.claude.com/docs/en/headless
- Google: Site Reliability Engineering, chapter 6, Monitoring Distributed Systems. https://sre.google/sre-book/monitoring-distributed-systems/
- Anthropic: API rate limits. https://platform.claude.com/docs/en/api/rate-limits
Chapter 3 · Review and merge
The morning: checking what it built
The night's output is a set of proposals. The morning decides which become part of the plant's platform. Check in the cheapest order, read every change that survives, try it on a phone, and merge one at a time.
35 minCheapest check firstNever trust a changed testRead every lineRebase, re-run, merge
By the end of this chapter you can
- Triage a night's pull requests in order, from automatic checks to a careful read.
- Spot the red flags in an agent's pull request: changed tests, out-of-scope files, oversized diffs, new open routes.
- Merge in sequence, rebasing and re-running CI before each merge.
- Turn each failure into a better brief, and record the night in the fleet ledger.
Cheapest check first
Do not start by reading the biggest pull request. Run every pull request through checks in order of cost: the automatic ones first, your careful reading last. A pull request that fails an early check goes back without you spending time reading it.
| Check | How | If it fails |
|---|---|---|
| 1. CI green from a fresh clone | The pull request's checks on GitHub [1] | Back to the worker with the failure; do not read further |
| 2. Only its own files | The pull request's Files changed list against the brief | Discard or split; tighten the file list |
| 3. Acceptance tests untouched | No changes under test/acceptance/ (a deny rule should have stopped it) | Discard: a worker that changed its own test has not done the task |
| 4. Readable size | Under about 400 changed lines; review quality falls with size [2] | Ask for it to be split |
| 5. Works where people use it | Open the pull request's preview address on your phone [3] | Back with what you saw |
Reading the change
Every change that reaches the plant's platform is read by a person. Google's code review guide puts the reviewer's job simply: approve once the change clearly improves the health of the code, even if it is not perfect, and look at design, function, complexity, tests, naming and comments [4]. For agent work, add the spine's rules:
| Look for | Why it matters on the plant platform |
|---|---|
| A new route or page | Is it closed until the Role and Exposure Matrix opens it (default deny, Session 8)? |
| A new table or view | Does it have row-level security from the matrix? Do views run with the viewer's rights (Session 14)? |
| A new dependency | Was it in the brief? Is it maintained and widely used? Supply-chain failures are an OWASP Top 10 category [5] |
| Anything that looks like a key | There should be none. If there is, rotate that key today, whatever else you decide |
| Tests that only check the happy path | Is there a refused request, an empty result and a wrong input? |
| Code the brief did not ask for | Unasked-for work is unreviewed scope; ask for it to be removed |
Merging in sequence
Each merge changes main, so a pull request that passed CI last night may not pass against this morning's main. Merge one at a time, in the order of the wave plan: rebase the next pull request on the new main, let CI run again, then merge [6]. GitHub's merge queue automates the same discipline when many pull requests land at once [7].
# Merge the next pull request in sequence (its branch: wave2/agent-3). Run in your main clone.
git fetch origin
git switch wave2/agent-3 && git rebase origin/main # stop here if there are conflicts
npm ci && npm test # the same checks CI will run
git push --force-with-lease origin wave2/agent-3 # updates the pull request; CI runs again
# When CI is green on GitHub, merge the pull request there (squash) and delete the branch.Every failure improves the next night
A pull request you send back is information. Most failures trace to the brief: a missing file path, an acceptance test that did not pin the behaviour, a limit that was not stated. Rewrite the brief and keep both versions: Session 16's homework (assignment 4) asks for a brief you rewrote after a bad result, and why.
| What went wrong | Brief fix for next time |
|---|---|
| Built the right thing in the wrong folder | Name the exact folders; add a deny rule for the rest |
| Passed the tests but missed the point | Add an acceptance test for the behaviour that was missed |
| Ran out of budget half way | Split the task; give each half its own test |
| Added a library nobody asked for | State the allowed dependencies; deny npm install |
| Ignored the phone layout | Make the phone check part of the acceptance tests |
Close the night
- Write one line per worker in the fleet ledger: model, turns, cost, merged or not, review minutes (Session 15).
- Add up the cost per merged change and compare it with last night.
- Delete merged branches and worktrees you no longer need:
git worktree remove ../agent-3. - Write the next night's plan while the lessons are fresh.
You need: The pull requests from your night (or from exercise 1's dry run); GitHub; your phone; the fleet ledger
You will take a night's pull requests through the checks, merge the good ones in order and rewrite one brief.
Outcome: Merged, reviewed work on main, a rewritten brief for homework 4, and the night's real cost per merged change.
Knowledge check
A pull request's CI is green, but it changed a file under test/acceptance/. What do you do?
Knowledge check
Two pull requests both passed CI last night. Why not merge both straight away?
Knowledge check
In which order should the morning checks run?
References
- GitHub Docs: About status checks. https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/collaborating-on-repositories-with-code-quality-features/about-status-checks
- SmartBear: Best practices for peer code review. https://smartbear.com/learn/code-review/best-practices-for-peer-code-review/
- Vercel docs: Deployments (preview deployments for every pull request). https://vercel.com/docs/deployments
- Google Engineering Practices: The standard of code review. https://google.github.io/eng-practices/review/reviewer/standard.html
- OWASP Top 10:2025. https://owasp.org/Top10/2025/
- Git documentation: git-rebase. https://git-scm.com/docs/git-rebase
- GitHub Docs: Managing a merge queue. https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/configuring-pull-request-merges/managing-a-merge-queue
Chapter 4 · 12 questions · 80% passes
Final assessment
Twelve questions across the element. Score 80% (10 of 12) to pass. Your LMS records your score and each answer; you can review the chapters and try again.
15 min12 questions≈ 15 minutesRetake allowed
Your result
CivOps AI Academy
AI Fleet: The Overnight Build and the Morning Review
Element F16 complete · Learner
Your LMS records this completion. For the CivOps Foundation certificate, finish the Foundation Course at https://civops.io/learn.