Chapter 1 · From a change to the floor
The deployment pipeline
Your platform is about to carry real work. From today every change, whether you wrote it or an agent did, takes one path to production, and every step on that path is checked by a machine or a person before the floor sees it.
35 min8 gates3 environments$0 on free plans to start
By the end of this chapter you can
- Trace a change from a branch to production through each gate of the pipeline.
- Keep local, preview and production apart, each with its own database.
- Change the database without breaking the version that is still running (expand, migrate, contract).
- Measure the pipeline with the four DORA measures.
Up to now you have deployed when you were ready. Once the floor depends on the platform, a deploy at 6:55 on a Monday that breaks the operator screen stops the start-up meeting. The answer is not to deploy less. It is to make every deploy small, checked and easy to undo, so it stops being an event.
One path for every change
You built most of this path in Sessions 1 and 17: the protected main branch, the required CI checks, the OWASP Top 10 suite and Vercel's preview addresses. Today you put it in writing and make sure nothing can skip it. GitHub's branch protection is what makes the path compulsory: with Require a pull request before merging and Require status checks to pass switched on, and bypassing switched off, not even an owner can push straight to main [1].
| Gate | What runs | Who or what decides | If it fails |
|---|---|---|---|
| 1–2 Branch and pull request | The change and a sentence on why | You or the agent | No pull request, no deploy |
| 3 CI checks | Build, unit tests, end-to-end tests, the OWASP suite, the migration dry run | GitHub Actions [2] | The merge button stays grey |
| 4 Preview | A full copy of the app on its own address, pointed at the staging database | Vercel, automatically on every push [3] | You see it before the floor does |
| 5 Review | Read the change; open the preview on your phone | A person: the owner of the change | Request changes |
| 6–7 Merge and production | Vercel builds main and moves the production address to it | Automatic | Roll back (chapter 2) |
| 8 Smoke check | A script opens the health page and one real screen | GitHub Actions | Alert, then roll back |
Three environments, three databases
A preview that points at the production database is a production deploy with a different address: a half-finished migration or a test that deletes rows hits real records. Keep a staging database for previews, loaded with the synthetic data from Session 4.
| Way to get a staging database | Cost | Notes |
|---|---|---|
| A second Supabase project on the Free plan | $0 (the Free plan allows two active projects) [4] | Pauses after a week without use; you apply migrations to it yourself or from CI |
| Supabase branching: a database per pull request | Paid: each branch is billed for the hours it runs [5] | Every preview gets a fresh database with your migrations and seed data |
Start with the free way. In Vercel, set the database address for the Preview environment to the staging project and for Production to the production project; Vercel keeps a separate value per environment [6]. You type both values into Vercel's settings page yourself. They never go into a chat, a prompt or a file.
Database changes travel as migrations
Every change to the schema is a numbered SQL file in supabase/migrations/, committed with the code that needs it. The Supabase command-line tool creates the file (supabase migration new add_scrap_reason), applies all files to a fresh local database to prove they run in order (supabase db reset) and applies new ones to a linked project (supabase db push) [7]. A migration that has run in production is never edited; a correction is a new migration.
# .github/workflows/deploy-db.yml: apply new migrations to staging on every pull request, to production on merge.
name: Database migrations
on:
pull_request:
push:
branches: [main]
jobs:
migrate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: supabase/setup-cli@v1
# The project id, access token and password are GitHub Actions secrets, typed into GitHub's settings by a person.
- run: supabase link --project-ref "$PROJECT"
env:
PROJECT: ${{ github.event_name == 'push' && secrets.PROD_PROJECT_REF || secrets.STAGING_PROJECT_REF }}
SUPABASE_ACCESS_TOKEN: ${{ secrets.SUPABASE_ACCESS_TOKEN }}
SUPABASE_DB_PASSWORD: ${{ github.event_name == 'push' && secrets.PROD_DB_PASSWORD || secrets.STAGING_DB_PASSWORD }}
- run: supabase db push
env:
SUPABASE_ACCESS_TOKEN: ${{ secrets.SUPABASE_ACCESS_TOKEN }}
SUPABASE_DB_PASSWORD: ${{ github.event_name == 'push' && secrets.PROD_DB_PASSWORD || secrets.STAGING_DB_PASSWORD }}Change the schema without breaking the running version
During a deploy, and after any rollback, the old version of the app runs against the new database. So a migration must suit both versions. The pattern is called expand and contract (also parallel change) [8]: add the new thing, move code and data across, and remove the old thing only in a later release, once nothing reads it.
- Safe in one step: add a table, add a nullable column, add an index, add a view.
- Needs expand and contract: rename or drop a column or table, change a column's type, make a column required.
- Row-level security: a new table arrives with its policies in the same migration (Session 7). A table without them never reaches production.
Measure the pipeline
Google's DevOps Research and Assessment programme (DORA) has measured software teams for more than a decade and found that speed and stability go together: teams that deploy small changes often also break things less and recover faster [9]. Four measures are enough:
| DORA measure | What it asks | Where to read it |
|---|---|---|
| Deployment frequency | How often do changes reach production? | Vercel's deployments list |
| Lead time for changes | From the first commit to production, how long? | Pull request opened to merged |
| Change failure rate | What share of deploys need a rollback or fix? | Your drill and incident log |
| Failed deployment recovery time | When a deploy fails, how long until the floor is back? | Incident log (chapter 3) |
These targets suit a plant platform with one or two owners; they are a starting point, not a DORA benchmark. Write your own in the runbook and check them monthly.
You need: Your GitHub repository, Vercel project and phone
You will confirm every gate is compulsory, including for you.
Outcome: Evidence that every change takes the same checked path, and your first lead-time measurement.
Knowledge check
A preview deployment is pointed at the production database. What is the risk?
Knowledge check
You need to rename a column the operator screen writes to. What is the safe way?
Knowledge check
Which of these is one of the four DORA measures?
References
- GitHub Docs: About protected branches. https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-protected-branches/about-protected-branches
- GitHub Actions documentation. https://docs.github.com/en/actions
- Vercel Docs: Environments and preview deployments. https://vercel.com/docs/deployments/environments
- Supabase pricing. https://supabase.com/pricing
- Supabase Docs: Branching. https://supabase.com/docs/guides/deployment/branching
- Vercel Docs: Environment variables. https://vercel.com/docs/environment-variables
- Supabase Docs: Database migrations. https://supabase.com/docs/guides/deployment/database-migrations
- Martin Fowler: Parallel Change (expand and contract). https://martinfowler.com/bliki/ParallelChange.html
- DORA: DevOps Research and Assessment. https://dora.dev/
Chapter 2 · Undo a bad deploy, recover lost data
Rollbacks and backups
Two different things go wrong. A bad deploy breaks the app while the data is fine: you roll back the code in seconds. A lost or damaged database is rarer and worse: you restore from a backup. Each has its own tool, and each is only real once you have practised it.
35 minRollback: secondsRestore: practised and timed3-2-1 copies
By the end of this chapter you can
- Roll back a bad deploy on Vercel and know what a rollback does not undo.
- Set a recovery point objective and a recovery time objective your plant can live with.
- Choose a backup plan, free way first, and keep one copy away from the platform.
- Restore a backup into a spare project and time it.
Rolling back a bad deploy
Vercel keeps every deployment it has built. The production address is a pointer to one of them, so going back is a matter of moving the pointer, with no rebuild. Vercel calls this Instant Rollback: from the project's Deployments page, open the menu on an earlier production deployment and choose Instant Rollback [1]. On the free Hobby plan you can return to the previous production deployment; paid plans can choose older ones [1].
- After a rollback, new merges do not go live by themselves until you promote a deployment again. That is deliberate: a second broken merge cannot undo your rollback [1].
- A rollback does not touch the database. If the bad release ran a migration, the old code now runs on the new schema. That is why chapter 1's expand-and-contract rule matters: it keeps the previous version compatible.
- A rollback does not change environment variables. If the problem was a wrong setting, fix the setting and redeploy; Vercel applies variable changes only to new deployments [2].
Backups: what you are promising
Before choosing a backup tool, decide two numbers with the people who use the platform. NIST's contingency planning guide defines them [3]:
| Term | Question it answers | A plant example |
|---|---|---|
| Recovery point objective (RPO) | How much recent data can we afford to lose? | One shift of entries could be re-keyed from paper: RPO 8 hours |
| Recovery time objective (RTO) | How long can the floor work without the platform? | Paper fallback covers a shift: RTO 4 hours |
Write both numbers in the runbook. They decide how often you back up (at least as often as the RPO) and how practised the restore must be (faster than the RTO, with time to spare).
Choosing a backup plan, free way first
| Option | Cost | Recovery point | Notes |
|---|---|---|---|
| Nightly export by GitHub Actions, encrypted, kept as a workflow artifact | $0: inside the free Actions minutes for a small database [4] | Up to 24 hours | You own and test it. Keep the encryption passphrase in the company password manager |
| Weekly download of that export to a company file store | $0 | Up to 7 days | Your offline copy; the 1 in 3-2-1 |
| Supabase Pro daily backups | Included in Pro (from $25 a month) [5] | Up to 24 hours, kept 7 days | Restored from the dashboard. The Free plan has no downloadable backups [6] |
| Supabase point-in-time recovery | A paid add-on to Pro; see the pricing page [5] | Minutes | Choose it when losing a shift's records is not acceptable |
Database backups do not include files in Supabase Storage, such as photos attached to a quality record [6]. If your platform stores files, back up the bucket too, or keep files somewhere that has its own backups.
The free way, written as a workflow. The three export commands follow Supabase's own backup guide [9]; underneath they run PostgreSQL's standard pg_dump tool, so the files restore into any Postgres database, not only Supabase [10].
# .github/workflows/backup.yml: a nightly, encrypted export of the production database.
name: Nightly backup
on:
schedule:
- cron: "17 7 * * *" # 07:17 UTC = 3:17 AM Eastern in summer (2:17 AM in winter)
workflow_dispatch: # a Run workflow button for drills
jobs:
backup:
runs-on: ubuntu-latest
steps:
- uses: supabase/setup-cli@v1
- name: Export roles, schema and data
env:
DB_URL: ${{ secrets.PROD_DB_URL }} # typed into GitHub's settings by a person
run: |
supabase db dump --db-url "$DB_URL" -f roles.sql --role-only
supabase db dump --db-url "$DB_URL" -f schema.sql
supabase db dump --db-url "$DB_URL" -f data.sql --use-copy --data-only
- name: Encrypt
env:
PASS: ${{ secrets.BACKUP_PASSPHRASE }} # also kept in the company password manager
run: tar czf - roles.sql schema.sql data.sql | gpg --batch --symmetric --cipher-algo AES256 --passphrase "$PASS" -o backup.tgz.gpg
- uses: actions/upload-artifact@v4
with:
name: backup-${{ github.run_id }}
path: backup.tgz.gpg
retention-days: 30A backup you have not restored is a hope
The only proof is a restore. Restore into the staging project, never into production, and time it from "start" to "a supervisor can open yesterday's records". Supabase's guide restores the three files with one command [9]:
# Run on your laptop, in a terminal, after decrypting the backup. STAGING_DB_URL is read from your password manager
# into the terminal session; it is never pasted into a chat or saved in a file.
gpg --decrypt backup.tgz.gpg | tar xzf -
psql --single-transaction --variable ON_ERROR_STOP=1 \
--file roles.sql --file schema.sql \
--command 'SET session_replication_role = replica' \
--file data.sql --dbname "$STAGING_DB_URL"- Note the time. Download last night's artifact from the workflow run page.
- Decrypt and restore into staging with the command above.
- Point a preview at staging (it already is) and open yesterday's records as a supervisor.
- Count rows in two key tables in production and in staging; they should match last night's numbers.
- Note the time. Write the date, who did it and the minutes taken in the runbook's drill log.
You need: Your Vercel project and phone
You will break the app in a harmless, visible way and undo it, so the first real rollback is not the first rollback.
Outcome: A timed rollback, done once in daylight, and production back on the normal path.
You need: GitHub, the staging Supabase project, a terminal with psql and gpg, the company password manager
You will prove the backup restores, and learn your real recovery time.
Outcome: A restore you have done yourself, with its time written down: the backup is now proven.
Knowledge check
A release broke the operator screen ten minutes ago. What do you do first?
Knowledge check
The supervisors say they could re-key up to one shift of entries from paper. What does that set?
Knowledge check
Which backup counts as proven?
References
- Vercel Docs: Instant Rollback. https://vercel.com/docs/instant-rollback
- Vercel Docs: Environment variables. https://vercel.com/docs/environment-variables
- NIST SP 800-34 Rev. 1: Contingency Planning Guide for Federal Information Systems. https://csrc.nist.gov/pubs/sp/800/34/r1/upd1/final
- GitHub pricing (Actions minutes and storage). https://github.com/pricing
- Supabase pricing. https://supabase.com/pricing
- Supabase Docs: Database backups. https://supabase.com/docs/guides/platform/backups
- CISA: Data Backup Options. https://www.cisa.gov/sites/default/files/publications/data_backup_options.pdf
- CISA: #StopRansomware Guide. https://www.cisa.gov/stopransomware/ransomware-guide
- Supabase Docs: Backup and restore using the CLI. https://supabase.com/docs/guides/platform/migrating-within-supabase/backup-restore
- PostgreSQL Documentation: pg_dump. https://www.postgresql.org/docs/current/app-pgdump.html
Chapter 3 · When something breaks at 3 AM
The runbook
A runbook is the page someone else follows when the platform misbehaves and you are not there. It is short, kept in the repository next to the code, printed by the line, and proven by drills: every procedure in it has been done once, by someone other than its author, with the time written down.
35 min8 sectionsEvery step drilledGenerated, then checked
By the end of this chapter you can
- Write a runbook with the eight sections a plant platform needs.
- Run an incident in the right order: detect, triage, mitigate, fix, review.
- Know when the platform is down before the floor calls, using a health page and a free monitor.
- Rotate a key without an outage and write the drill down.
What goes in a runbook
Google's site reliability engineers describe the value plainly: written procedures, practised in advance, roughly triple the speed of recovery compared with improvising [1]. You do not need their scale. You need one page per procedure, in plain words, that a supervisor or IT colleague can follow.
| Section | What it holds | Proven by |
|---|---|---|
| What it is | Purpose, users, address, where each part runs, links to the repository, Vercel and Supabase | A new colleague finds everything from this page |
| Who to call | Platform owner, a second owner, IT contact, Vercel and Supabase support links, in order, with hours | Names and phone numbers checked each quarter |
| Deploy | The normal path (chapter 1) and how to check a deploy worked | The pipeline exercise |
| Roll back | The Instant Rollback steps and when to choose rollback over a fix | The rollback drill, timed |
| Restore a backup | Where backups are, how to decrypt, restore to staging, then production | The restore drill, timed |
| Rotate a key | Every key, where it lives, the order of steps | The rotation drill, timed |
| Known problems | Symptom, what to check, what fixes it | Each review adds an entry |
| Drill log | Date, who, which procedure, minutes taken, what changed | Kept up to date; at least one drill a month |
Generate the first draft, then make it true
Your repository already knows most of the runbook: the workflows, the environment variable names (never their values), the migrations, the health page. Ask your coding agent to draft it from the repository, then correct it by doing every procedure. The draft saves an hour; the drills are what make it trustworthy.
Read this repository: .github/workflows, vercel.json, supabase/migrations, app/api/health and the README.
Write docs/runbook.md with these sections: What it is; Who to call (leave names as a table for me to fill in);
Deploy; Roll back (Vercel Instant Rollback); Restore a backup (from the nightly backup workflow, into staging first);
Rotate a key (list every environment variable NAME the code reads and where each is set; never print a value);
Known problems (empty table: symptom, check, fix); Drill log (empty table: date, who, procedure, minutes, changes).
Plain language for a shift supervisor. Every step names the exact page or command. Open a pull request.Know before the floor calls
A health page is a small route that checks the app can reach the database and answers ok or an error. An uptime monitor opens it every few minutes and alerts you when it fails. Free monitors exist, such as UptimeRobot's free plan [2]; a scheduled GitHub Actions workflow can also do it at no cost, though GitHub may delay scheduled runs at busy times [3].
// app/api/health/route.ts: answers 200 when the app can read from the database, 503 when it cannot.
import { createClient } from "@supabase/supabase-js";
export const dynamic = "force-dynamic";
export async function GET() {
const started = Date.now();
const db = createClient(process.env.NEXT_PUBLIC_SUPABASE_URL!, process.env.NEXT_PUBLIC_SUPABASE_PUBLISHABLE_KEY!);
// A public, harmless read: a one-row table the matrix opens to everyone, holding the current schema version.
const { error } = await db.from("platform_status").select("schema_version").limit(1);
const body = { ok: !error, ms: Date.now() - started, version: process.env.VERCEL_GIT_COMMIT_SHA?.slice(0, 7) };
return Response.json(body, { status: error ? 503 : 200, headers: { "Cache-Control": "no-store" } });
}The health page reports no records and no keys, only whether the parts talk to each other and which version is live. Add /api/health to the matrix as open to all, and point the pipeline's smoke check (gate 8) at it too.
Running an incident
NIST's incident response guide, revised in 2025, frames incident response as part of everyday risk management: prepare, detect, respond, recover, and feed what you learn back into preparation [4]. For a plant platform that comes down to five habits:
- Detect. The monitor alerts, or a supervisor calls. Write the time down.
- Triage. What is broken (one screen, everything), since when (match it to the last deploy in Vercel), how many people are affected.
- Mitigate. If the last deploy is the likely cause, roll back. If the database is the cause, switch the floor to the paper fallback and follow the restore page. Tell the floor what to do in one sentence.
- Fix. The real fix is a normal change through the pipeline, reviewed and tested. Never edit production by hand under pressure.
- Review. Within a week, write one page: what happened, the timeline, what helped, what will change. Blame the system, not the person, so people keep reporting problems early [5].
| Known problem (example) | Check | Fix |
|---|---|---|
| Every screen shows "Something went wrong" right after a deploy | Vercel Deployments: did a deploy finish in the last hour? | Instant Rollback to the previous deployment |
| Health page returns 503; screens load but nothing saves | Supabase dashboard: is the project paused or over a limit? | Restore the project or raise the limit; on the Free plan, restore a paused project from the dashboard |
| One operator cannot sign in; others can | Supabase Authentication, Users: is the account disabled or unconfirmed? | Re-send the invitation; never share another person's login |
| Entries from the tablet on Line 2 arrive late | Is the tablet offline? The offline queue shows a count (Session 10) | Reconnect Wi-Fi; the queue sends by itself |
Rotating a key without an outage
Keys leak: a laptop is lost, a contractor leaves, a key is pasted somewhere it should not be. Rotation is the routine that makes a leak a non-event. Supabase's newer API keys support several secret keys at once, so a new one can be created, deployed and proven before the old one is revoked [6]. Vercel only uses a changed variable in deployments made after the change, so a redeploy is part of every rotation [7]. OWASP's secrets guidance adds the rest: rotate on a schedule and at once after any suspected exposure, and record every rotation [8].
| Key | Where it is set | Overlap possible? | Rotate |
|---|---|---|---|
| Supabase secret key (server only) | Vercel Production; GitHub Actions if a workflow uses it | Yes: create a second secret key, then revoke the first [6] | Every 6 months and after any exposure |
| Supabase publishable key (browser) | Vercel, all environments | Yes | If abused; it is public by design and row-level security protects the data |
| Database password | Vercel and GitHub Actions secrets (backup, migrations) | No: the old password stops at once; do it at a quiet time | Every 6 months, when an owner leaves, after any exposure |
| Backup passphrase | GitHub Actions secret and the password manager | Keep the old one until its backups expire | Yearly |
You need: Your coding agent, GitHub, Vercel, Supabase, the company password manager, a colleague
You will draft the runbook from the repository, correct it, then have a colleague rotate the Supabase secret key by following it.
Outcome: A runbook that a colleague has followed successfully, with its first rotation drill logged.
Knowledge check
Why does every key rotation on Vercel include a redeploy?
Knowledge check
During an incident you suspect last night's deploy. Which step comes first?
Knowledge check
What should the runbook say about the Supabase secret key?
References
- Google SRE book: Introduction (playbooks and mean time to repair). https://sre.google/sre-book/introduction/
- UptimeRobot. https://uptimerobot.com/
- GitHub Docs: Events that trigger workflows (schedule). https://docs.github.com/en/actions/writing-workflows/choosing-when-your-workflow-runs/events-that-trigger-workflows
- NIST SP 800-61 Rev. 3: Incident Response Recommendations and Considerations for Cybersecurity Risk Management. https://csrc.nist.gov/pubs/sp/800/61/r3/final
- Google SRE book: Postmortem culture. https://sre.google/sre-book/postmortem-culture/
- Supabase Docs: Understanding API keys. https://supabase.com/docs/guides/api/api-keys
- Vercel Docs: Environment variables. https://vercel.com/docs/environment-variables
- OWASP Secrets Management Cheat Sheet. https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html
Chapter 4 · 12 questions · 80% passes
Final assessment
Twelve questions across the element. Score 80% (10 of 12) to pass. Your LMS records your score and each answer; you can review the chapters and try again.
15 min12 questions≈ 15 minutesRetake allowed
Your result
CivOps AI Academy
Deploy and Run: the Pipeline, Rollbacks, Backups and the Runbook
Element F18 complete · Learner
Your LMS records this completion. For the CivOps Foundation certificate, finish the Foundation Course at https://civops.io/learn.