Skip to the lesson
CivOps AI Academy · F17Prove It Works: OWASP, Conformance and the Tests That Keep It True
0%

Chapter 1 · Session 17

The OWASP Top 10 as a test suite

IT will ask whether the platform is secure. 'We were careful' is not an answer. A suite of tests mapped to the OWASP Top 10, run on every change, is an answer, and it keeps being true after the people who wrote it move on.

35 minOWASP Top 10:2025One check per categoryFree scanners firstRuns on every pull request

By the end of this chapter you can

  • Name the ten OWASP Top 10:2025 categories and what each looks like on a plant platform.
  • Map each category to an automated check that runs in CI.
  • Write the most important one yourself: every role against every route, allowed or refused.
  • Add free scanners for dependencies, keys and a running preview, and read their results.

Why a list from OWASP

The Open Worldwide Application Security Project (OWASP) is a non-profit whose Top 10 is the most widely used list of web application security risks. The 2025 edition is built from data on real applications and a community survey, and IT departments and auditors recognise it [1]. Mapping your tests to it gives you and IT a shared language: "A01 is covered by these 25 tests, and they pass."

OWASP Top 10 mapped to checksTen categories, from A01 Broken Access Control to A10 Mishandling of Exceptional Conditions, each with the check that covers it. OWASP Top 10:2025, and the check that proves each on your platformA01Broken Access Controlrole × route matrix testA02Security Misconfigurationheaders test · database advisorA03Software Supply Chain Failuresnpm audit · Dependabot · lockfileA04Cryptographic FailuresHTTPS-only test · secret scanA05Injectionhostile-input tests · parameters onlyA06Insecure Designthreat model → testsA07Authentication Failuressign-in limits · session expiryA08Software or Data Integrity Failuresprotected main · audit trailA09Security Logging and Alerting Failuresaudit-log testA10Mishandling of Exceptional Conditionsfail-closed testsGreen: an automated test in CI · Cyan: a tool or setting checked by a test · Amber: a design step that produces tests
The OWASP Top 10:2025 and the check that covers each. Most are automated tests; a few are tools or settings that a test confirms; one, insecure design, is a step that produces tests.

Each category on a plant platform

CategoryWhat it looks like hereThe check
A01 Broken Access ControlAn operator opens /admin; a supervisor at one site reads another site's records by changing an id in the addressThe role × route grid below, plus a cross-site test per table
A02 Security MisconfigurationA table without row-level security; a view that runs with its owner's rights; missing security headersA test that reads every table's security setting; the Supabase advisor; a headers test [2]
A03 Software Supply Chain FailuresAn agent adds an unmaintained package with a known vulnerabilitynpm audit in CI; Dependabot alerts; the lockfile committed [3][4]
A04 Cryptographic FailuresA page served over plain HTTP; a key committed to the repositoryA test that HTTP redirects to HTTPS with HSTS; a secret scan on every push
A05 InjectionA downtime reason containing SQL or a script tagTests that submit hostile input and expect it stored as plain text; queries use parameters only
A06 Insecure DesignNo limit on how many records an operator can submit a minuteA short threat model per workflow; each threat becomes a test
A07 Authentication FailuresUnlimited password guesses; sessions that never expireTests for sign-in rate limits and for an expired session being refused
A08 Software or Data Integrity FailuresCode reaching main without review; a record changed with no traceBranch protection checked; a test that every change writes an audit row
A09 Security Logging and Alerting FailuresA refused request leaves no trace, so nobody notices someone probingA test that a refused request writes a log entry with who, what and when
A10 Mishandling of Exceptional ConditionsWhen the database is slow, the app skips the access check and shows the page anywayTests that force errors and expect a refusal, never an open door

Category names as published in the OWASP Top 10:2025 [1]. For a fuller checklist when IT asks for one, OWASP's Application Security Verification Standard (ASVS) lists hundreds of verifiable requirements in levels [5].

A01 first: every role against every route

Broken access control has been first on the list since 2021 [1]. It is also the easiest to test thoroughly, because your Role and Exposure Matrix already says exactly what each role may reach. Generate one test per cell: sign in as the role, request the route, expect allowed or refused. Then add a row for another site's supervisor, which should be refused everywhere, and a check that fails when a route exists that the matrix does not mention, so default deny stays true [6].

Role and route test gridFive roles by five routes; only the cells the matrix opens are allowed, every other request must be refused, including all requests from another site. 25 tests generated from the Role and Exposure Matrix: each cell is a request the suite makes/entry/review/dashboard/admin/api/recordsanonymousrefusedrefusedrefusedrefusedrefusedoperator200 allowedrefusedrefusedrefused200 allowedsupervisor200 allowed200 allowed200 allowedrefused200 allowedmanagerrefusedrefused200 allowedrefusedrefusedother siterefusedrefusedrefusedrefusedrefusedA refused cell is a passing test only when the request really is refused; a new route with no row fails the suite
The matrix as a test grid. Twenty-five requests, each one a test. The anonymous and other-site rows must be refused everywhere.
Node.js test runner: the access grid (run against your local app or a preview)
// test/acceptance/access-grid.test.mjs: one test per role and route, from docs/matrix.json.
import { test } from "node:test";
import assert from "node:assert/strict";
import matrix from "../../docs/matrix.json" with { type: "json" };
import { signIn, BASE } from "../helpers.mjs"; // signs in a seeded test user and returns its cookie

for (const role of [...matrix.roles, "anonymous", "other-site-supervisor"]) {
  for (const route of matrix.routes) {
    const allowed = (matrix.allow[role] || []).includes(route);
    test(`${role} ${allowed ? "may" : "may not"} open ${route}`, async () => {
      const cookie = role === "anonymous" ? "" : await signIn(role);
      const res = await fetch(BASE + route, { headers: { cookie }, redirect: "manual" });
      if (allowed) assert.equal(res.status, 200);
      else assert.ok([302, 401, 403, 404].includes(res.status), `got ${res.status}`);
    });
  }
}

Free scanners, then paid if you need them

ScannerWhat it findsCost
npm auditKnown vulnerabilities in your dependencies, from the lockfile [3]Free, built into npm
Dependabot alerts and updatesVulnerable dependencies; opens pull requests to update them [4]Free on GitHub
GitleaksKeys and passwords committed to the repository, in CI or before a commit [7]Free, open source
GitHub secret scanningKeys in the repository, with push protection [8]Free for public repositories; for private ones it is part of a paid GitHub plan, so start with Gitleaks
ZAP baseline scanMissing headers and common misconfigurations on a running site; it looks but does not attack [9]Free, open source; run it against a preview, never production
Supabase advisorsTables without row-level security, views that bypass it, and other risky settings [2]Free in every Supabase project
Exercise · Build your OWASP suite45 minutes

You need: Your repository; your Role and Exposure Matrix; your AI coding agent; GitHub; the Supabase dashboard

You will brief the agent to build the access grid, add the free scanners and map every category to a check.

Outcome: Every OWASP Top 10:2025 category mapped to a check that runs automatically, with the access grid passing.

Knowledge check

A supervisor changes the record id in the address bar and sees another site's record. Which category is this?

Knowledge check

Why add a check that fails when a route exists that the matrix does not mention?

Knowledge check

Where should you run a ZAP baseline scan?

References

  1. OWASP Top 10:2025. https://owasp.org/Top10/2025/
  2. Supabase docs: Performance and security advisors. https://supabase.com/docs/guides/database/database-advisors
  3. npm Docs: npm audit. https://docs.npmjs.com/cli/commands/npm-audit
  4. GitHub Docs: Dependabot. https://docs.github.com/en/code-security/dependabot
  5. OWASP Application Security Verification Standard (ASVS). https://owasp.org/www-project-application-security-verification-standard/
  6. OWASP Cheat Sheet Series: Authorization. https://cheatsheetseries.owasp.org/cheatsheets/Authorization_Cheat_Sheet.html
  7. Gitleaks. https://github.com/gitleaks/gitleaks
  8. GitHub Docs: About secret scanning. https://docs.github.com/en/code-security/secret-scanning/introduction/about-secret-scanning
  9. ZAP: Baseline scan. https://www.zaproxy.org/docs/docker/baseline-scan/

Chapter 2 · Intent to evidence

Conformance: does it do what the intent says?

Secure is not enough. The platform must also do what your intent promised: the right people make the right decision from the right data. Conformance tests turn every line of the intent into a check you can show.

30 minEvery line tracedGiven · when · thenHand-worked numbersWatch it fail first

By the end of this chapter you can

  • Trace every line of your intent and matrix to an acceptance criterion and a test.
  • Write acceptance criteria in a given-when-then form an agent can implement and a manager can read.
  • Test computed numbers against values worked by hand.
  • Prove a test can fail by breaking the rule it checks.

From intent to evidence

Your intent (Session 2) says what the platform is for. The matrix (Session 7) says who may do what. Conformance means each of those promises has a test, and each test has a result. Lay it out as a table in your repository, one row per promise: the line of intent, its acceptance criterion, the test file and the latest result. This is called a traceability matrix, and it is exactly what an auditor or IT reviewer wants to see.

Traceability from intent to evidenceFour linked boxes: an intent line, its acceptance criterion, the test file, and the passing CI run. IntentSupervisor approves orrejects each record, witha reasonCriterionGiven a pending record,when the supervisorrejects it without areason, then it is refusedTesttest/acceptance/review-reject.test.mjsEvidenceCI run on the pullrequest: passedEvery line of the intent traces to a criterion, a test and a result you can show ITA line with no test is a promise nobody checks; a test with no line is work nobody asked for
From a line of intent to a green check. The same chain, repeated for every promise, is the conformance table.

Criteria people and agents can both read

Write each criterion as given (the starting situation), when (what someone does) and then (what must happen). This form, from behaviour-driven development, is plain enough for a supervisor to confirm and precise enough for an agent to turn into a test [1].

Intent lineGiven · when · then
Operators record downtime on the line, even when the network dropsGiven an operator offline, when they save a downtime entry, then it is kept on the device and sent when the network returns, exactly once
Supervisors approve or reject each record, with a reasonGiven a pending record, when the supervisor rejects it without a reason, then the reject is refused and the record stays pending
Managers see OEE per line each shiftGiven the seeded shift for Press 02, when the manager opens the dashboard, then OEE shows 75.1%
Nobody outside the site sees its recordsGiven a supervisor from another site, when they request any record list, then they get no rows

Numbers worked by hand

A dashboard that shows the wrong number with confidence is dangerous. For every measure, keep one case worked by hand from the seeded data, as Session 14 did for OEE (75.1% for the seeded Press 02 shift), and assert the platform shows the same value. If the formula changes, the hand-worked case must be re-worked by a person, not updated by the agent to match the new output.

The access matrix, again

The access grid from chapter 1 is also a conformance test: it proves the matrix you wrote is the matrix that runs. Keep one source of truth (docs/matrix.json), generate both the row-level security and the tests from it, and the two can never drift apart.

Watch every test fail once

A test that has never failed may not be testing anything. Before you trust a new acceptance test, break the rule it checks on a scratch branch (remove the role check, change the formula) and run it. It must go red. Then throw the scratch branch away. Mutation-testing tools automate the same idea at scale by making small changes to your code and checking the tests notice [2].

A test that can failFour steps: test passes, break the rule, the test must fail, restore the rule; a test still green when the rule is broken is not testing it. Before you trust a test, watch it fail for the right reason1Test passesrule in place2Break the rulee.g. remove the role check3Test must failred proves it checks therule4Restoregreen again; commit nothingfrom step 2Still green at step 3? The test was not testing the rule.
A test that can fail. Red at step 3 proves the test is tied to the rule; green there means it is not.
Terminal: break the rule, see red, throw it away
# Prove the reject-reason test checks the rule. Run in your repository folder.
git switch -c scratch/break-reject-rule
# Edit app/review/actions.ts: comment out the line that refuses a reject with no reason, then:
npm test -- --test-name-pattern="reject"     # must FAIL
git switch main && git branch -D scratch/break-reject-rule
npm test -- --test-name-pattern="reject"     # passes again

Secure development frameworks ask for the same discipline: NIST's Secure Software Development Framework expects that software is checked against its security requirements and that the evidence is kept [3]. Your conformance table and its CI results are that evidence.

Exercise · Write your conformance table40 minutes

You need: Your intent page; docs/matrix.json; your measure sheet from Session 14; your AI coding agent

You will turn every promise in your intent into a criterion and a test, and prove two of the tests can fail.

Outcome: A conformance table where every promise in the intent has a criterion, a test that can fail, and a current result.

Knowledge check

Which is a well-formed acceptance criterion?

Knowledge check

The OEE formula changes and the hand-worked test fails. What should happen?

Knowledge check

You break the rule a test is meant to check, and the test still passes. What does that tell you?

References

  1. Cucumber: Gherkin reference (Given, When, Then). https://cucumber.io/docs/gherkin/reference/
  2. Stryker Mutator: mutation testing. https://stryker-mutator.io/
  3. NIST SP 800-218, Secure Software Development Framework (SSDF) Version 1.1. https://csrc.nist.gov/pubs/sp/800/218/final

Chapter 3 · CI, clamps and the audit

Tests that keep it true

A platform that passed its tests once is not secure or correct forever. Agents will change it every week. The tests stay true only if they run on every change, block any merge that fails, grow with every bug, and are reported in a form IT can read.

30 minThe test pyramidRequired checksA test for every bugThe clamp audit

By the end of this chapter you can

  • Shape a suite as a pyramid: many fast tests, few slow ones, all run on every pull request.
  • Make the checks required, so nothing merges past a red gate.
  • Add a regression test for every bug, and deal with flaky tests without switching them off.
  • Produce a clamp audit: each clamp, its evidence and the action for anything not green.

The shape of the suite

Fast, focused tests catch most mistakes cheaply; slow browser tests prove the journeys work end to end. Keep many of the first and few of the second. This shape is usually drawn as a pyramid [1]. Google's testing team gives the reason: end-to-end tests are slow and their failures are hard to trace, so they should cover the main journeys while smaller tests carry the detail [2].

The test pyramidEnd to end (browser, phone sizes): 24 tests, ≈ 3 min; Integration (API + database + access rules): 70 tests, ≈ 40 s; Unit (rules, money, measures): 180 tests, ≈ 8 s An example suite for the plant platform: many fast tests at the bottom, a few slow ones on topEnd to end (browser, phone sizes)24 tests≈ 3 minper runIntegration (API + database + access rules)70 tests≈ 40 sper runUnit (rules, money, measures)180 tests≈ 8 sper runEvery layer runs on every pull request; the slow layer stays small by testing journeys, not every case
An example suite for the plant platform. The whole pyramid runs in under five minutes, so it can run on every pull request.
LayerTests things likeTool
UnitOEE and measure formulas, money in cents, date and shift rulesNode's built-in test runner
IntegrationAPI routes with the database and its row-level security; the access grid; audit rowsNode's test runner against a local database
End to endAn operator records downtime on a phone; the supervisor rejects it; the manager sees the changePlaywright at 360 and 390 px wide [3]

Required checks: no merge past a red gate

In GitHub's branch protection for main, mark each CI job as a required status check. Then a pull request cannot merge until every one is green, whoever opened it, agent or person [4]. Run the checks from a fresh clone, so files that exist only on someone's laptop cannot make a test pass.

CI gatesSeven required checks in sequence, from a fresh clone to the clamps, plus a weekly scheduled scan. Required checks on every pull request: any red gate blocks the mergeFresh clonenpm ciBuildno warningsUnit +integrationAccessmatrix gridEnd to end360 + 390 pxSecurityaudit · secretsClampsall ten passWeekly schedule: full dependency audit, baselinescan of the preview, restore drill (Session 18)Merge allowed only when every gate is green,run from a fresh clone, on the latest main
The gates every change passes. Weekly scheduled jobs catch what changes without a pull request: new vulnerabilities in old dependencies.
.github/workflows/weekly-security.yml: the scheduled scan
name: weekly-security
on:
  schedule:
    - cron: "17 9 * * 1"   # Mondays 09:17 UTC (05:17 Eastern in summer, 04:17 in winter)
  workflow_dispatch: {}
jobs:
  scan:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: npm ci && npm audit --audit-level=high
      - uses: gitleaks/gitleaks-action@v2
        env: { GITHUB_TOKEN: "${{ secrets.GITHUB_TOKEN }}" }

GitHub runs scheduled workflows from the default branch; in public repositories it pauses a schedule after 60 days without activity [5]. Scheduled minutes count against the Actions allowance like any other job.

Every bug gets a test

When something breaks in use, the fix is not finished until a test exists that would have caught it. Write the test first, watch it fail against the broken code, then fix. Over time the suite becomes a record of every mistake the platform will never make again. The CivOps platforms run on this rule: when a change fails, the test that would have caught it is added along with the fix.

Flaky tests

A flaky test passes and fails on the same code. Google found that a meaningful share of its tests showed some flakiness, and that flaky tests teach people to ignore red [6]. Never fix flakiness by deleting the test or marking it optional. Find the cause (usually timing, shared data or test order), fix it within the week, and give each test its own data so order does not matter.

The clamp audit

Session 9 set out the quality clamps: the ten checks every change must pass. The Clamp Audit is the report of them: one page listing each clamp, the evidence that it passes today (a test name, a scan result, a count), and an action and date for anything not green. Generate it from the CI results rather than writing it by hand, so it is always current. The same page goes into the IT approval package in Session 19.

Clamp audit reportSix example clamp rows with evidence; five pass, one dependency finding to fix this week. Clamp audit, example: each clamp, its evidence, and the action for anything not greenDefault denyaccess grid: 25 of 25passRow-level security on every tableadvisor: 0 issuespassNo keys in the repositorysecret scan: 0 findingspassPhone rules360 + 390 px: 0 problemspassDependenciesaudit: 1 moderatefix this weekAudit log on every changelog test: passedpassThe full audit lists all ten clamps from Session 9; IT reads the same page in the approval package (Session 19)
A clamp audit, example rows. Evidence comes from the latest CI run; anything amber has an owner and a date.
Exercise · Make the checks required and produce the clamp audit35 minutes

You need: GitHub (repository settings); your CI workflow; the suite from chapters 1 and 2; your AI coding agent

You will make every check block merges, schedule the weekly scan, and generate your first clamp audit.

Outcome: Required checks that block a bad merge, a weekly scan, a regression test for a real bug, and a clamp audit ready for IT.

Knowledge check

Why keep end-to-end browser tests few and the unit tests many?

Knowledge check

A test fails on one run and passes on the next with no code change. What should you do?

Knowledge check

What makes a clamp audit trustworthy to IT?

References

  1. Martin Fowler: The Practical Test Pyramid. https://martinfowler.com/articles/practical-test-pyramid.html
  2. Google Testing Blog: Just Say No to More End-to-End Tests. https://testing.googleblog.com/2015/04/just-say-no-to-more-end-to-end-tests.html
  3. Playwright documentation. https://playwright.dev/docs/intro
  4. GitHub Docs: About protected branches (required status checks). https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-protected-branches/about-protected-branches
  5. GitHub Docs: Events that trigger workflows (schedule). https://docs.github.com/en/actions/writing-workflows/choosing-when-your-workflow-runs/events-that-trigger-workflows
  6. Google Testing Blog: Flaky Tests at Google and How We Mitigate Them. https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html

Chapter 4 · 12 questions · 80% passes

Final assessment

Twelve questions across the element. Score 80% (10 of 12) to pass. Your LMS records your score and each answer; you can review the chapters and try again.

15 min12 questions≈ 15 minutesRetake allowed

Choose one answer for each question, then submit. You will see the right answer and why for every question.

1. Which category is first in the OWASP Top 10:2025?
2. How many requests does the access grid make for 5 roles and 5 routes?
3. A new route exists that the matrix does not mention. What should the suite do?
4. Which free tool finds keys committed to a private repository?
5. When the database is slow, the app skips the access check and shows the page. Which category?
6. Where should the ZAP baseline scan point?
7. Which part of 'given, when, then' states the observable result?
8. Why keep a hand-worked value for each dashboard measure?
9. You break the rule a new test checks and it stays green. The test:
10. What does a required status check on main do?
11. A bug reaches users. When is the fix finished?
12. What makes a clamp audit trustworthy?