F12 · Foundation Session 12
You direct, the agent builds
From this session on, an AI coding agent writes most of the code. Your job changes: you say what to build and how you will know it is right, and you check what comes back. This element teaches the three skills that job needs.
8 min5 chapters≈ 1½ hours3 hands-on exercises12-question assessment · 80% passes
By the end of this chapter you can
- Describe the directing loop: brief, build, check, merge or rewrite.
- Describe how this element is scored.
- Know the change you will direct: adding a field to the first-off check.
A new job
Unit 4 is where the platform grows quickly. An agent can write in an hour what took a week, and it will happily write the wrong thing in an hour too. It does not know your plant, it does not remember last week's decisions unless they are written down, and it will say it is done when it thinks it is done. None of that is a reason not to use it. It is a reason to direct it well.
Directing well takes three skills, one per chapter: writing a brief the agent can act on, writing the acceptance checks before the work starts, and reading the change it hands back before it ships. The ten quality clamps from Session 9 catch the mistakes a machine can catch. These three skills cover the rest.
The change you will direct
Acme's quality engineer wants burr height measured on the first-off check on Line 3, not just a yes or no: a number in millimetres, limit 0 to 0.10. It touches every layer of the platform: a migration, an access rule, the operator screen, the review card and the board. That makes it a good first brief: small enough for one pull request, big enough to go wrong in interesting ways.
How it is scored
The element is complete when you have opened every chapter and taken the assessment, and passed at 80% or more. Homework 4 (a workflow built by agents, due at the end of week 6) asks for your briefs, the changes and your checks; start keeping them now.
Knowledge check
An agent reports that a change is done. What is your next step?
Chapter 1 · Goal, limits and the finish line
A brief an agent can act on
An agent does what the brief says, fills the gaps with guesses, and stops when it believes it is finished. A good brief leaves no important gaps and says exactly what finished looks like.
28 min6 parts1 pull request of work0 keys in any brief
By the end of this chapter you can
- Write a brief with a goal, context, limits, acceptance checks, a done-when and a report.
- Keep standing rules in a file the agent reads every time, not in each brief.
- Size a brief to one reviewable pull request.
What an agent needs
Anthropic's guidance for its own coding agent puts it plainly: the more specific the instructions, the better the first attempt, and the fewer corrections later [1]. Specific does not mean long. It means the agent can answer, from the brief alone, four questions: what outcome do you want, where are the edges, how will we both know it is done, and what should I tell you when I stop.
| Part | Says | Example for burr height |
|---|---|---|
| Goal | The outcome, in plant words, and why | Record burr height in mm on the first-off check, so the quality engineer can see burr trends by die |
| Context | What to read first; which rules stand | Read docs/intent.md and the first_off_checks migration. Follow CLAUDE.md |
| Limits | What it may change, and what it may not | Add one migration and change the check screen, its API and the review card. Do not edit tests, applied migrations or the workflow |
| Acceptance checks | The tests that prove it, written by you first | test/burr.spec.ts and test/burr-rls.sql must pass |
| Done when | One observable result | Both checks pass and all ten clamps are green on the pull request |
| Report back | What you want to hear when it stops | Files changed, decisions made, anything it could not do or was unsure of |
Standing rules go in one file
Some rules apply to every change: money in whole cents, row-level security on every table, phone first, secrets never in code, the clamps. Do not repeat them in every brief. Coding agents read a standing-instructions file at the start of every session: Claude Code reads CLAUDE.md [2], and many other agents read AGENTS.md, an open format for the same purpose [3]. Keep it in the repository, short, and written as rules. When an agent makes the same mistake twice, add a rule.
# Rules for agents working in this repository
1. Every table has row-level security, generated from docs/matrix.md. Never add a table without it.
2. A route is closed until docs/matrix.md opens it.
3. Money is integer cents. Measurements are stored in the unit named in the column (for example burr_mm).
4. Never edit a test to make it pass, never edit an applied migration, never add || true to a script.
5. Every page passes the phone check at 360 and 390 px.
6. No keys, passwords or tokens in code, chats or commits. Ask a person to set them in Vercel or GitHub.
7. Stop and report when a limit in the brief would have to be broken to finish.Size: one pull request
A brief should produce one change a person can read in about fifteen minutes. If you cannot list its acceptance checks on one screen, split it. Agents do better on small, well-defined tasks with a clear check than on large open ones, and so do reviewers. Anthropic's own advice on building with agents is to start with the simplest approach that works and add complexity only when it is needed [4].
Let it plan before it builds
For anything bigger than a few lines, ask the agent to read the relevant files and write a short plan first, then stop. Read the plan; it takes two minutes and catches the misunderstanding before it becomes two hundred lines of code. Anthropic recommends this explore-plan-code pattern for its agent [1]. Approve the plan or correct it, then let the agent build.
Knowledge check
Which part of a brief tells the agent what it must not change?
Knowledge check
Where should the rule 'every table has row-level security' live?
Knowledge check
A brief needs three screens to list its acceptance checks. What should you do?
You need: A text editor; your repository
Write the brief for your own platform's equivalent of the burr-height change: one new measured field on your first task.
Outcome: A six-part brief of under fifteen lines, and a standing-instructions file with the platform's rules.
References
- Anthropic: Claude Code best practices. https://www.anthropic.com/engineering/claude-code-best-practices
- Claude Code Docs: Manage Claude's memory (CLAUDE.md). https://code.claude.com/docs/en/memory
- AGENTS.md: a simple, open format for guiding coding agents. https://agents.md/
- Anthropic: Building effective agents. https://www.anthropic.com/engineering/building-effective-agents
Chapter 2 · Decide what right looks like
Acceptance checks first
Write the test before the agent writes the code. A check that fails now and passes after the change is the only proof that the change did what you asked.
22 minRed, then greenPlant words in every checkChecks the agent cannot edit
By the end of this chapter you can
- Write acceptance checks before the work starts, in the plant's words.
- Turn a check into an automated test with Playwright or SQL.
- Make sure the agent cannot pass by changing the check.
Why first
If the checks are written after the code, they tend to describe what the code does rather than what you needed. Written first, they are your definition of done, fixed before anyone is tempted to bend it. This is test-driven development: write a failing test, make it pass, then tidy up [1]. With an agent it matters even more, because the agent can iterate against a test on its own, many times, without you [2].
Write them in plant words
Start in plain language, in the shape given, when, then, which the Gherkin format popularised for exactly this purpose [3]. Each line names a person, an action and an observable result. Then turn each one into an automated test.
| Given | When | Then |
|---|---|---|
| An operator on Line 3 at 360 px wide | They enter burr height 0.04 and submit | The check is saved with burr_mm = 0.04 and shows Pass |
| An operator | They enter burr height 0.14 | The server records Fail: burr above 0.10, and the review card shows +0.04 above limit |
| An operator | They enter burr height as text, '0,04' | The screen asks for a number and nothing is saved |
| A signed-in user from another plant | They read first-off checks | They see none of Acme's rows |
test("burr height above its limit fails on the server and shows the difference", async ({ page }) => {
await signInAs(page, "operator-line3");
await page.setViewportSize({ width: 360, height: 800 });
await page.goto("/operator/first-off?press=L3-P02");
await page.getByLabel("Burr height (mm)").fill("0.14");
// ... fill the other fields within limits ...
await page.getByRole("button", { name: "Submit check" }).click();
await expect(page.getByText("Fail: burr above 0.10")).toBeVisible();
const [row] = await sql`select burr_mm, result from first_off_checks order by entered_at desc limit 1`;
expect(row).toEqual({ burr_mm: "0.14", result: "fail" }); // decided by the server, not the browser
});Playwright drives a real browser at phone width, so the same check proves the screen, the server and the record [4]. A second, short SQL test signs in as a role from another plant and expects zero rows, proving the new column did not open a gap in the access rules.
Checks the agent cannot bend
An agent told to make a test pass may make it pass by changing the test. It is not malice; it is the shortest path to what it was asked. Close that path three ways:
- Say it in the limits: the brief and the standing rules both say tests may not be edited.
- Commit the checks first, on main or on the branch before the agent starts, so any change to them shows up as a separate, obvious diff.
- Make test files need your review: a
CODEOWNERSfile can name you as the required reviewer for thetest/folder, so GitHub asks for your approval whenever it changes [5].
Knowledge check
Your new check passes before the agent has changed anything. What does that mean?
Knowledge check
Why is a Playwright check at 360 px a good acceptance check for an operator field?
Knowledge check
Which control makes GitHub require your review when test files change?
You need: Your repository; Playwright; your local database
Write the acceptance checks for your brief before giving it to the agent.
Outcome: Failing acceptance checks committed before any code, with test changes requiring your review.
References
- Martin Fowler: Test Driven Development. https://martinfowler.com/bliki/TestDrivenDevelopment.html
- Anthropic: Claude Code best practices. https://www.anthropic.com/engineering/claude-code-best-practices
- Cucumber: Gherkin reference. https://cucumber.io/docs/gherkin/reference/
- Playwright documentation: Writing tests. https://playwright.dev/docs/writing-tests
- GitHub Docs: About code owners. https://docs.github.com/en/repositories/managing-your-repositorys-settings-and-features/customizing-your-repository/about-code-owners
Chapter 3 · The riskiest lines first
Reading the change it hands back
A green pull request means the machine checks passed. It does not mean the change is right. Read it in a fixed order, riskiest first, try it on a phone, and decide: merge, fix, or rewrite the brief.
28 min6 steps15 minutes1 phone
By the end of this chapter you can
- Read a pull request in a fixed order, from the summary to the code.
- Spot the red flags that agents commonly introduce.
- Decide whether to merge, ask for a fix, or rewrite the brief, and record why.
Green is necessary, not sufficient
The clamps prove the change compiles, follows the rules a machine can check, and passes the tests. They cannot tell you whether it solves the right problem, whether it added something nobody asked for, or whether the agent quietly worked around a limit. That is the reviewer's job. Google's published guidance for its own reviewers says the same: look first at whether the change is designed well and does what was needed, then at the details [1].
Read in this order
- The summary against the brief. Did it do what you asked, all of it, and nothing else? Anything it says it could not do is the most important line in the report.
- Tests. In the pull request's Files changed tab, look for any change to test files [2]. A weakened assertion or a deleted test is the first thing to stop on.
- Migrations. Every new table has row-level security switched on and a policy from the matrix. No applied migration was edited.
- Access rules and routes. Nothing that was closed is now open. Middleware, policies and the matrix changed only as the brief allowed.
- Packages and configuration. Any new dependency is known, maintained and needed. The workflow files and environment settings are unchanged unless the brief said so.
- The code. Now read it, knowing what it should do. Errors handled, not swallowed; names that match the naming standard; no copy-pasted blocks.
Try it yourself
Open the pull request's preview address on your phone [3] and do the task as an operator would: with gloves, at 360 px, with a value outside the limits. Then sign in as the supervisor and as another plant's user. Five minutes on a phone finds what fifteen minutes of reading misses.
Merge, fix, or rewrite the brief
| What you found | Do this | Write down |
|---|---|---|
| Everything matches the brief and the checks | Approve and merge | Nothing extra |
| A small slip within the brief (a label, a missed case) | Ask the agent for the fix on the same branch | The comment you left |
| The agent misread the goal or crossed a limit | Close it, rewrite the brief with the missing limit or check, run again | The old brief, the new brief and why |
| A test was weakened or a rule broken to get green | Close it, add the rule to CLAUDE.md, rewrite the brief | The rule you added |
A brief that produced a bad result is not a failure; it is the most useful thing you will learn from this unit. Homework 4 rewards showing one brief you rewrote after a bad result, and why. Keep both versions next to the pull requests.
Knowledge check
A pull request is green, but a test assertion changed from 'exactly 1 row' to 'more than 0 rows'. What do you do?
Knowledge check
Why read the code last?
Knowledge check
The agent misread the goal and built the wrong thing. What is the best next step?
You need: Your AI coding agent; github.com; your phone
Give your agent the brief from exercise 1 and review what comes back in the order from this chapter.
Outcome: A change directed, read in order, tried on a phone and merged or sent back, with the brief and the review kept.
References
- Google Engineering Practices: How to do a code review. https://google.github.io/eng-practices/review/reviewer/
- GitHub Docs: Reviewing proposed changes in a pull request. https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/reviewing-changes-in-pull-requests/reviewing-proposed-changes-in-a-pull-request
- Vercel Docs: Environments and preview deployments. https://vercel.com/docs/deployments/environments
- OWASP Gen AI Security Project: Top 10 for LLM Applications. https://genai.owasp.org/llm-top-10/
Chapter 4 · 12 questions · 80% passes
Final assessment
Twelve questions across the element. Score 80% (10 of 12) to pass. Your LMS records your score and each answer; you can review the chapters and try again.
15 min12 questions≈ 15 minutesRetake allowed
Your result
CivOps AI Academy
Directing an AI Agent: Briefs, Checks First, and Reading the Change
Element F12 complete · Learner
Your LMS records this completion. For the CivOps Foundation certificate, finish the Foundation Course at https://civops.io/learn.