Skip to the lesson
CivOps AI Academy · F12Directing an AI Agent: Briefs, Checks First, and Reading the Change
0%

F12 · Foundation Session 12

You direct, the agent builds

From this session on, an AI coding agent writes most of the code. Your job changes: you say what to build and how you will know it is right, and you check what comes back. This element teaches the three skills that job needs.

8 min5 chapters≈ 1½ hours3 hands-on exercises12-question assessment · 80% passes

By the end of this chapter you can

  • Describe the directing loop: brief, build, check, merge or rewrite.
  • Describe how this element is scored.
  • Know the change you will direct: adding a field to the first-off check.

A new job

Unit 4 is where the platform grows quickly. An agent can write in an hour what took a week, and it will happily write the wrong thing in an hour too. It does not know your plant, it does not remember last week's decisions unless they are written down, and it will say it is done when it thinks it is done. None of that is a reason not to use it. It is a reason to direct it well.

Directing well takes three skills, one per chapter: writing a brief the agent can act on, writing the acceptance checks before the work starts, and reading the change it hands back before it ships. The ten quality clamps from Session 9 catch the mistakes a machine can catch. These three skills cover the rest.

The directing loopYou write a brief and checks, the agent works on a branch, opens a pull request where clamps run, you read the diff, then merge or rewrite the brief. Directing an agent is a loop; the brief improves each time roundYou writebrief + checksAgent workson its own branchPull requestclamps runYou readthe diff, in orderMergea person approvesnot right: rewrite the brief, say why, run againKeep every version of the brief with the pull request: it is the record of what you asked for.
The directing loop. Every change goes round it once. When the result is wrong, the fix is usually a better brief.

The change you will direct

Acme's quality engineer wants burr height measured on the first-off check on Line 3, not just a yes or no: a number in millimetres, limit 0 to 0.10. It touches every layer of the platform: a migration, an access rule, the operator screen, the review card and the board. That makes it a good first brief: small enough for one pull request, big enough to go wrong in interesting ways.

How it is scored

5chapters, each opened at least once
3exercises with your own agent
12assessment questions
80%to pass (10 of 12)

The element is complete when you have opened every chapter and taken the assessment, and passed at 80% or more. Homework 4 (a workflow built by agents, due at the end of week 6) asks for your briefs, the changes and your checks; start keeping them now.

Knowledge check

An agent reports that a change is done. What is your next step?

Chapter 1 · Goal, limits and the finish line

A brief an agent can act on

An agent does what the brief says, fills the gaps with guesses, and stops when it believes it is finished. A good brief leaves no important gaps and says exactly what finished looks like.

28 min6 parts1 pull request of work0 keys in any brief

By the end of this chapter you can

  • Write a brief with a goal, context, limits, acceptance checks, a done-when and a report.
  • Keep standing rules in a file the agent reads every time, not in each brief.
  • Size a brief to one reviewable pull request.

What an agent needs

Anthropic's guidance for its own coding agent puts it plainly: the more specific the instructions, the better the first attempt, and the fewer corrections later [1]. Specific does not mean long. It means the agent can answer, from the brief alone, four questions: what outcome do you want, where are the edges, how will we both know it is done, and what should I tell you when I stop.

Anatomy of a briefA brief card with six parts from top to bottom: goal, context, limits, acceptance checks, done when, and report back. brief: first-off check, add burr heightGoalthe outcome, tied to the intentContextfiles to read, rules that standLimitswhat it may touch, and may notAcceptance checkstests you wrote firstDone whenan observable resultReport backwhat changed, what it could not doWhat and whyWhere the edges areHow you will knowShort: one pullrequest of workChecks you own;the agent cannot edit
Anatomy of a brief. Six parts, usually fewer than fifteen lines. The checks are yours; the agent may not change them.
PartSaysExample for burr height
GoalThe outcome, in plant words, and whyRecord burr height in mm on the first-off check, so the quality engineer can see burr trends by die
ContextWhat to read first; which rules standRead docs/intent.md and the first_off_checks migration. Follow CLAUDE.md
LimitsWhat it may change, and what it may notAdd one migration and change the check screen, its API and the review card. Do not edit tests, applied migrations or the workflow
Acceptance checksThe tests that prove it, written by you firsttest/burr.spec.ts and test/burr-rls.sql must pass
Done whenOne observable resultBoth checks pass and all ten clamps are green on the pull request
Report backWhat you want to hear when it stopsFiles changed, decisions made, anything it could not do or was unsure of
Vague and actionable briefsLeft: a vague request with no goal, limits or finish line. Right: the same request with a goal, limits, a done-when check and a report. Vague"Improve the first-off screenand make it better for theoperators. Add whatever youthink is missing."No goal you can testNo limits: may touch anythingNo way to know it is doneActionableGoal: add burr height (mm) to thefirst-off check, limits 0 to 0.10.Touch: the check screen, its API,one new migration. Not: tests.Done when: test/burr.spec passesand all ten clamps are green.Report: files changed, anythingyou could not do.
Vague and actionable. The same request, written two ways. Only one can be checked.

Standing rules go in one file

Some rules apply to every change: money in whole cents, row-level security on every table, phone first, secrets never in code, the clamps. Do not repeat them in every brief. Coding agents read a standing-instructions file at the start of every session: Claude Code reads CLAUDE.md [2], and many other agents read AGENTS.md, an open format for the same purpose [3]. Keep it in the repository, short, and written as rules. When an agent makes the same mistake twice, add a rule.

CLAUDE.md (or AGENTS.md): standing rules
# Rules for agents working in this repository
1. Every table has row-level security, generated from docs/matrix.md. Never add a table without it.
2. A route is closed until docs/matrix.md opens it.
3. Money is integer cents. Measurements are stored in the unit named in the column (for example burr_mm).
4. Never edit a test to make it pass, never edit an applied migration, never add || true to a script.
5. Every page passes the phone check at 360 and 390 px.
6. No keys, passwords or tokens in code, chats or commits. Ask a person to set them in Vercel or GitHub.
7. Stop and report when a limit in the brief would have to be broken to finish.

Size: one pull request

A brief should produce one change a person can read in about fifteen minutes. If you cannot list its acceptance checks on one screen, split it. Agents do better on small, well-defined tasks with a clear check than on large open ones, and so do reviewers. Anthropic's own advice on building with agents is to start with the simplest approach that works and add complexity only when it is needed [4].

Let it plan before it builds

For anything bigger than a few lines, ask the agent to read the relevant files and write a short plan first, then stop. Read the plan; it takes two minutes and catches the misunderstanding before it becomes two hundred lines of code. Anthropic recommends this explore-plan-code pattern for its agent [1]. Approve the plan or correct it, then let the agent build.

Knowledge check

Which part of a brief tells the agent what it must not change?

Knowledge check

Where should the rule 'every table has row-level security' live?

Knowledge check

A brief needs three screens to list its acceptance checks. What should you do?

Exercise · Write the burr-height brief20 minutes

You need: A text editor; your repository

Write the brief for your own platform's equivalent of the burr-height change: one new measured field on your first task.

Outcome: A six-part brief of under fifteen lines, and a standing-instructions file with the platform's rules.

References

  1. Anthropic: Claude Code best practices. https://www.anthropic.com/engineering/claude-code-best-practices
  2. Claude Code Docs: Manage Claude's memory (CLAUDE.md). https://code.claude.com/docs/en/memory
  3. AGENTS.md: a simple, open format for guiding coding agents. https://agents.md/
  4. Anthropic: Building effective agents. https://www.anthropic.com/engineering/building-effective-agents

Chapter 2 · Decide what right looks like

Acceptance checks first

Write the test before the agent writes the code. A check that fails now and passes after the change is the only proof that the change did what you asked.

22 minRed, then greenPlant words in every checkChecks the agent cannot edit

By the end of this chapter you can

  • Write acceptance checks before the work starts, in the plant's words.
  • Turn a check into an automated test with Playwright or SQL.
  • Make sure the agent cannot pass by changing the check.

Why first

If the checks are written after the code, they tend to describe what the code does rather than what you needed. Written first, they are your definition of done, fixed before anyone is tempted to bend it. This is test-driven development: write a failing test, make it pass, then tidy up [1]. With an agent it matters even more, because the agent can iterate against a test on its own, many times, without you [2].

Checks firstThree steps in a loop: red, a check that fails; green, the agent makes it pass; tidy, clean up while it stays green. 1 Redyou write the check; it fails2 Greenthe agent makes it pass3 Tidyclean up; still greennext checkA check that fails before the change and passes after it proves the change did the job
Checks first. The check fails before the change (red) and passes after it (green). If it passes before the change, it proves nothing.

Write them in plant words

Start in plain language, in the shape given, when, then, which the Gherkin format popularised for exactly this purpose [3]. Each line names a person, an action and an observable result. Then turn each one into an automated test.

GivenWhenThen
An operator on Line 3 at 360 px wideThey enter burr height 0.04 and submitThe check is saved with burr_mm = 0.04 and shows Pass
An operatorThey enter burr height 0.14The server records Fail: burr above 0.10, and the review card shows +0.04 above limit
An operatorThey enter burr height as text, '0,04'The screen asks for a number and nothing is saved
A signed-in user from another plantThey read first-off checksThey see none of Acme's rows
test/burr.spec.ts (Playwright)
test("burr height above its limit fails on the server and shows the difference", async ({ page }) => {
  await signInAs(page, "operator-line3");
  await page.setViewportSize({ width: 360, height: 800 });
  await page.goto("/operator/first-off?press=L3-P02");
  await page.getByLabel("Burr height (mm)").fill("0.14");
  // ... fill the other fields within limits ...
  await page.getByRole("button", { name: "Submit check" }).click();
  await expect(page.getByText("Fail: burr above 0.10")).toBeVisible();
  const [row] = await sql`select burr_mm, result from first_off_checks order by entered_at desc limit 1`;
  expect(row).toEqual({ burr_mm: "0.14", result: "fail" });     // decided by the server, not the browser
});

Playwright drives a real browser at phone width, so the same check proves the screen, the server and the record [4]. A second, short SQL test signs in as a role from another plant and expects zero rows, proving the new column did not open a gap in the access rules.

Checks the agent cannot bend

An agent told to make a test pass may make it pass by changing the test. It is not malice; it is the shortest path to what it was asked. Close that path three ways:

  • Say it in the limits: the brief and the standing rules both say tests may not be edited.
  • Commit the checks first, on main or on the branch before the agent starts, so any change to them shows up as a separate, obvious diff.
  • Make test files need your review: a CODEOWNERS file can name you as the required reviewer for the test/ folder, so GitHub asks for your approval whenever it changes [5].

Knowledge check

Your new check passes before the agent has changed anything. What does that mean?

Knowledge check

Why is a Playwright check at 360 px a good acceptance check for an operator field?

Knowledge check

Which control makes GitHub require your review when test files change?

Exercise · Write the checks, watch them fail25 minutes

You need: Your repository; Playwright; your local database

Write the acceptance checks for your brief before giving it to the agent.

Outcome: Failing acceptance checks committed before any code, with test changes requiring your review.

References

  1. Martin Fowler: Test Driven Development. https://martinfowler.com/bliki/TestDrivenDevelopment.html
  2. Anthropic: Claude Code best practices. https://www.anthropic.com/engineering/claude-code-best-practices
  3. Cucumber: Gherkin reference. https://cucumber.io/docs/gherkin/reference/
  4. Playwright documentation: Writing tests. https://playwright.dev/docs/writing-tests
  5. GitHub Docs: About code owners. https://docs.github.com/en/repositories/managing-your-repositorys-settings-and-features/customizing-your-repository/about-code-owners

Chapter 3 · The riskiest lines first

Reading the change it hands back

A green pull request means the machine checks passed. It does not mean the change is right. Read it in a fixed order, riskiest first, try it on a phone, and decide: merge, fix, or rewrite the brief.

28 min6 steps15 minutes1 phone

By the end of this chapter you can

  • Read a pull request in a fixed order, from the summary to the code.
  • Spot the red flags that agents commonly introduce.
  • Decide whether to merge, ask for a fix, or rewrite the brief, and record why.

Green is necessary, not sufficient

The clamps prove the change compiles, follows the rules a machine can check, and passes the tests. They cannot tell you whether it solves the right problem, whether it added something nobody asked for, or whether the agent quietly worked around a limit. That is the reviewer's job. Google's published guidance for its own reviewers says the same: look first at whether the change is designed well and does what was needed, then at the details [1].

What an agent may reachThe agent works inside a workspace with its own branch and test database; main, secrets and the live database are outside its reach. The agent's workspaceIts branchLocal test DBRead the repository and the docsmainonly through a reviewed PRKeys and secretsnever in its reachLive databaseneverYoubrief, review, merge
What an agent may reach. Its own branch, a test database and the repository. Main, secrets and the live database stay with people.

Read in this order

Reading order for a changeSix stacked bars, narrowing: summary, tests, migrations, access rules and routes, packages and config, then the code. Read a change in this order: the riskiest parts first, the code last1The agent's summaryagainst your brief: did it do what you asked?2Testschanged or deleted tests are the first red flag3Migrationsnew tables, and RLS on each4Access rules and routesanything opened that was closed5Packages and confignew dependencies, workflow, env6The code itselfnow, knowing what it should do
Reading order. The parts that can do the most damage come first; the code itself comes last, when you know what it should do.
  1. The summary against the brief. Did it do what you asked, all of it, and nothing else? Anything it says it could not do is the most important line in the report.
  2. Tests. In the pull request's Files changed tab, look for any change to test files [2]. A weakened assertion or a deleted test is the first thing to stop on.
  3. Migrations. Every new table has row-level security switched on and a policy from the matrix. No applied migration was edited.
  4. Access rules and routes. Nothing that was closed is now open. Middleware, policies and the matrix changed only as the brief allowed.
  5. Packages and configuration. Any new dependency is known, maintained and needed. The workflow files and environment settings are unchanged unless the brief said so.
  6. The code. Now read it, knowing what it should do. Errors handled, not swallowed; names that match the naming standard; no copy-pasted blocks.
Red flags in a diffFive diff lines with a flag each: a weakened test, a table without row-level security, a route opened outside the matrix, an unknown dependency, an error swallowed silently. Five lines worth stopping fortest/first-off.spec.ts- expect(rows).toBe(1) + expect(rows).toBeGreaterThan(0)a test was weakeneddb/migrations/0012_burr.sql+ create table burr_limits (...)no row-level securitymiddleware.ts+ if (path.startsWith('/api/burr')) return next()a route opened outside the matrixpackage.json+ "left-pad-plus": "^0.0.3"an unknown new dependencyapp/api/burr/route.ts+ } catch { }an error swallowed silently
Red flags. Each of these has passed a green pipeline somewhere. Each is a reason to stop and ask.

Try it yourself

Open the pull request's preview address on your phone [3] and do the task as an operator would: with gloves, at 360 px, with a value outside the limits. Then sign in as the supervisor and as another plant's user. Five minutes on a phone finds what fifteen minutes of reading misses.

Merge, fix, or rewrite the brief

What you foundDo thisWrite down
Everything matches the brief and the checksApprove and mergeNothing extra
A small slip within the brief (a label, a missed case)Ask the agent for the fix on the same branchThe comment you left
The agent misread the goal or crossed a limitClose it, rewrite the brief with the missing limit or check, run againThe old brief, the new brief and why
A test was weakened or a rule broken to get greenClose it, add the rule to CLAUDE.md, rewrite the briefThe rule you added

A brief that produced a bad result is not a failure; it is the most useful thing you will learn from this unit. Homework 4 rewards showing one brief you rewrote after a bad result, and why. Keep both versions next to the pull requests.

Knowledge check

A pull request is green, but a test assertion changed from 'exactly 1 row' to 'more than 0 rows'. What do you do?

Knowledge check

Why read the code last?

Knowledge check

The agent misread the goal and built the wrong thing. What is the best next step?

Exercise · Direct, read, decide35 minutes

You need: Your AI coding agent; github.com; your phone

Give your agent the brief from exercise 1 and review what comes back in the order from this chapter.

Outcome: A change directed, read in order, tried on a phone and merged or sent back, with the brief and the review kept.

References

  1. Google Engineering Practices: How to do a code review. https://google.github.io/eng-practices/review/reviewer/
  2. GitHub Docs: Reviewing proposed changes in a pull request. https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/reviewing-changes-in-pull-requests/reviewing-proposed-changes-in-a-pull-request
  3. Vercel Docs: Environments and preview deployments. https://vercel.com/docs/deployments/environments
  4. OWASP Gen AI Security Project: Top 10 for LLM Applications. https://genai.owasp.org/llm-top-10/

Chapter 4 · 12 questions · 80% passes

Final assessment

Twelve questions across the element. Score 80% (10 of 12) to pass. Your LMS records your score and each answer; you can review the chapters and try again.

15 min12 questions≈ 15 minutesRetake allowed

Choose one answer for each question, then submit. You will see the right answer and why for every question.

1. Which set makes a brief an agent can act on?
2. Where do rules that apply to every change belong?
3. How big should one brief be?
4. The work needs an API key. What does the brief say?
5. Why ask the agent for a plan before it builds?
6. When should acceptance checks be written?
7. A new check already passes before any change. What is wrong?
8. Which file makes GitHub require your review when tests change?
9. What do you read first in an agent's pull request?
10. Which of these is a red flag in a green pull request?
11. The agent misread the goal. What is the best next step?
12. Why can prompt injection not be fixed by a better brief alone?