Skip to the lesson
CivOps AI Academy · F14Widgets and the Dashboard: Measures Managers Act On
0%

Chapter 1 · Session 14

The measures managers act on

A dashboard is not a wall of numbers. Each number on it exists because someone makes a decision with it. This chapter turns the decision in your intent into a short list of measures, each with a formula, a target and an owner.

30 minDecision firstISO 22400 definitionsOEE worked through≈ 6 measures per role

By the end of this chapter you can

  • Trace every measure back to a decision in your intent and forward to an action.
  • Write a measure sheet: definition, formula, unit, source tables, target, owner and what to do when it is red.
  • Calculate OEE and its three factors from a shift's data, and name the loss behind each factor.
  • Tell a leading measure from a lagging one, and spot a measure nobody will act on.

Start from the decision, not the data

In Session 2 you wrote your plant's intent: the decision the platform exists to support, who makes it and the data it needs. A dashboard is where that decision gets made. So a dashboard starts from the decision too. If you start from the data instead ("we have downtime, scrap and counts, let's chart them"), you get a screen full of charts that nobody acts on.

From decision to widget to actionFive steps left to right: decision, question, measure, widget, action, with a press-shop example under each. Every widget traces back to a decision, and forward to an actionDecisionfrom the intentwho decides whatQuestionwhat they needto know firstMeasureformula · unittarget · ownerWidgetsource · visualrefresh · rolesActionwhat happenswhen it is redExample · Line 3 · press shopSchedule maintenanceon Press 02?Which press loses themost time?Downtime minutes bypress, per shiftPareto bar, refreshedevery 5 minPlanner books the dierepair
The chain every widget must complete. If you cannot name the action a red number triggers, the widget does not belong on the dashboard.

Walk the chain for each manager who will use the dashboard. Ask them: what do you decide each day or each shift, and what do you need to know first? Write the answer down in their words. Then turn each answer into a measure.

What a good measure looks like

A measure is more than a name. "Downtime" is not a measure; "minutes a press was stopped during planned production time, per press, per shift, from the downtime events table" is. Write every measure on one line of a measure sheet with these columns:

ColumnExample (Line 3, press shop)Why it matters
DecisionPlanner: book Press 02 for maintenance this week?Ties the measure to the intent
DefinitionMinutes stopped during planned production timeEveryone counts the same way
Formulasum(end − start) of downtime_events, per press, per shiftThe agent can build it exactly
Unit and periodMinutes per shiftStops comparing a shift to a week
Source tablesdowntime_events, lines, shiftsProves the spine already holds the data
Target and limitsUnder 30 min a shift; act above 45Says when it is red
OwnerMaintenance plannerSomeone answers for it
Action when redOpen a work order for the top causeIf there is no action, drop the measure

Use the standard definitions

Manufacturing already has agreed definitions for the common measures. ISO 22400-2 defines key performance indicators for manufacturing operations management, including availability, effectiveness, quality ratio, scrap ratio, OEE, mean time between failures and mean time to repair, each with its formula, unit and the times and quantities it is built from [1]. Using them means your OEE means the same as your customer's OEE, and the data elements line up with the ISA-95 model your spine already uses [2].

OEE, worked through

Overall equipment effectiveness multiplies three factors. Each factor catches different losses, so the factor that is lowest tells you where to look [3][4]. Here is one shift on Press 02:

QuantityValueFrom
Planned production time450 minshift length minus planned breaks
Downtime63 min (breakdown 38, die change 25)downtime_events
Run time387 min450 − 63
Ideal rate20 parts a minute (3-second cycle)the press's standard
Total parts6,966production counts
Good parts6,757total minus rejects from quality_checks
FactorFormulaResult
Availabilityrun time ÷ planned time = 387 ÷ 45086.0%
Performancetotal parts ÷ (run time × ideal rate) = 6,966 ÷ 7,74090.0%
Qualitygood parts ÷ total parts = 6,757 ÷ 6,96697.0%
OEE0.860 × 0.900 × 0.97075.1%

Check it a second way: good parts × ideal cycle time ÷ planned time = 6,757 × 3 s ÷ (450 × 60 s) = 75.1%. Both ways must agree. Put that check in a test, so the agent's code is proven against a number you worked by hand.

OEE and the six big lossesOEE of 75.1% splits into availability 86%, performance 90% and quality 97%; under each sit two of the six big losses. OEE 75.1%availability × performance × qualityAvailability 86.0%387 of 450 planned minBreakdownsSetups, changeoversPerformance 90.0%6,966 of 7,740 ideal partsSmall stops, idlingReduced speedQuality 97.0%6,757 good of 6,966Start-up rejectsProduction rejectsThe six big losses: each one shows up in exactly one factor, so the factor tells you where to look
OEE and the six big losses. Availability is the lowest factor here, and the breakdown minutes are the biggest loss inside it, so that is where the planner looks first.

Leading and lagging

A lagging measure tells you what already happened: yesterday's OEE, last month's scrap cost. A leading measure moves before the result does: open quality holds, overdue preventive maintenance, a die that is past its stroke count. Managers need both. Lagging measures say whether you are winning; leading measures give them time to act. A good dashboard pairs each lagging measure with at least one leading one.

Measures that mislead

  • Averages that hide the spread. A week's OEE of 75% can hide one shift at 40%. Show the shift or the line, not only the plant.
  • Percentages without the count. "Scrap 50%" on a two-part trial run is not a crisis. Show the count beside the percentage.
  • A target that becomes the goal. When people are judged on one number, they find ways to move the number instead of the work (for example, calling a breakdown a planned stop). Pair each measure with one that would expose gaming: availability with planned-stop minutes.
  • Measures from data the spine does not hold. If a measure needs a spreadsheet someone fills in by hand, fix the data first (Session 13's workflow), then build the widget.
Exercise · Write the measure sheet for your intent30 minutes

You need: Your intent page from Session 2; a spreadsheet or a Markdown table in your repository; one manager for 15 minutes

You will interview one manager and leave with a measure sheet the agent can build from.

Outcome: A measure sheet of up to six measures, each tied to a decision, with one hand-worked number to test against.

Knowledge check

A manager asks for a widget showing 'all the machine data'. What should you ask first?

Knowledge check

Planned time 450 min, run time 387 min, total 6,966 parts at an ideal 20 a minute, 6,757 good. Which factor is lowest?

Knowledge check

Which is a leading measure for scrap cost?

References

  1. ISO 22400-2:2014, Automation systems and integration: Key performance indicators (KPIs) for manufacturing operations management, Part 2: Definitions and descriptions. https://www.iso.org/standard/54497.html
  2. ISA: ISA-95 standard, enterprise-control system integration. https://www.isa.org/standards-and-publications/isa-standards/isa-95-standard
  3. Vorne: Calculating OEE. https://www.oee.com/calculating-oee/
  4. Vorne: The six big losses. https://www.oee.com/oee-six-big-losses/

Chapter 2 · The Widget Configurator

Widgets: configuration on the spine's data

A widget is not a new program. It is a few lines of configuration (which data, which measure, which target, which picture, how fresh, who may see it) drawn by one component the agent builds once. Adding a widget never opens a new path to the data.

35 minOne renderer, many widgetsViews that respect access rulesThe right visual for the questionReviewed like code

By the end of this chapter you can

  • Name the seven parts of a widget's configuration.
  • Build a database view that computes a measure and runs with the viewer's rights, so row-level security still applies.
  • Choose the visual that answers each kind of question, and say why pies and gauges are poor choices.
  • Brief an agent to add a widget, with an acceptance test that a role outside the matrix sees nothing.

Why configuration and not code

You could ask the agent to build each chart as its own page. Ten charts later you would have ten pieces of code, ten ways of fetching data, and ten places an access rule could be forgotten. Instead, the agent builds one widget component once, and every widget is a short entry in a configuration file. The curriculum calls this the Widget Configurator: in your platform it is that file, widgets.json, and the component that draws it.

  • One path to the data. The component reads only from approved database views. A new widget cannot reach a table the views do not expose.
  • Easy to review. A new widget is a few lines in a pull request. You can read it in a minute; you cannot do that with a new page of code.
  • Easy for an agent. "Add a widget for scrap ratio by press" becomes a small, checkable change instead of a new feature.
Anatomy of a widgetSeven configuration fields sit on a view that runs with the viewer's rights, which reads spine tables protected by row-level security. The spine: tables with row-level security (one rule set, generated from the Role and Exposure Matrix)production_runs · downtime_events · quality_checks · lines · sitesA view that runs with the viewer's rights (security_invoker)line_shift_oee: availability, performance, quality and OEE per line and shiftsourcethe viewmeasurethe formulatargetand limitsvisualtile · trend · barrefreshand freshnessdrillwhere a tap goesroleswho sees itwidgets.json: one entry per widget, reviewed like code
Anatomy of a widget. The configuration sits on a view; the view runs with the viewer's rights, so the row-level security generated from your matrix still decides what each person sees.

The source: a view that keeps the access rules

Each measure is computed in the database by a view: a saved query that looks like a table. In PostgreSQL 15 and later, a view created with security_invoker = on runs with the rights of the person asking, so the row-level security policies on the underlying tables still apply [1][2]. A view without it runs with its owner's rights and can show every site's rows to everyone. Supabase's database advisor flags such views as a security problem, so check it after every schema change [3].

SQL: a measure as a view (Supabase SQL editor or a migration file)
-- One measure, computed once, in the database. Runs with the viewer's rights.
create view public.line_shift_oee
  with (security_invoker = on) as
select r.site_id, r.line_id, r.shift_date, r.shift,
       sum(r.run_min)::numeric / nullif(sum(r.planned_min), 0)                as availability,
       sum(r.total_count)::numeric / nullif(sum(r.run_min * r.ideal_rate), 0) as performance,
       sum(r.good_count)::numeric / nullif(sum(r.total_count), 0)             as quality,
       sum(r.good_count * 60.0 / r.ideal_rate) / nullif(sum(r.planned_min) * 60, 0) as oee,
       max(r.recorded_at)                                                     as as_of
from public.production_runs r
group by r.site_id, r.line_id, r.shift_date, r.shift;

-- The view adds no access of its own: a supervisor still sees only their site's rows,
-- because production_runs has row-level security generated from the Role and Exposure Matrix.
grant select on public.line_shift_oee to authenticated;

The configuration

Each widget is one entry. The component reads the entry, asks the view for rows (as the signed-in person), and draws the visual. Nothing in the entry is a secret, and nothing in it can widen access: the roles field can only hide a widget from people who could already read its view.

widgets.json: three widgets
[
  { "id": "oee-now", "title": "OEE this shift", "source": "line_shift_oee", "measure": "oee",
    "filter": { "shift_date": "today", "shift": "current" }, "group": "line_id",
    "visual": "tile", "format": "percent", "target": 0.80, "red_below": 0.65,
    "refresh_s": 60, "stale_after_s": 300, "drill": "/lines/:line_id", "roles": ["supervisor", "manager"] },
  { "id": "oee-trend", "title": "OEE, last 14 days", "source": "line_shift_oee", "measure": "oee",
    "filter": { "shift_date": "-14d" }, "visual": "line", "limits": "control", "refresh_s": 900, "roles": ["manager"] },
  { "id": "downtime-pareto", "title": "Downtime by cause, this week", "source": "downtime_by_cause", "measure": "minutes",
    "filter": { "week": "current" }, "visual": "pareto", "refresh_s": 300, "roles": ["supervisor", "manager", "planner"] }
]

The visual: answer the question that was asked

Each kind of question has a visual that answers it best. In a classic study of how people read charts, Cleveland and McGill found that people judge positions along a common scale (bars, dots on a line) far more accurately than angles, areas or colour shades [4]. That is why pie charts, round gauges and 3-D effects lose to sorted bars and plain lines.

Choosing a visualFive manager questions, each mapped to a visual: number tile, line trend, sorted bars, status grid, control chart. The manager's questionThe visual that answers itHow are we doing right now?Number tile with its target and trend arrowIs it getting better or worse?Line over time, same scale every dayWhat causes most of the loss?Bars sorted largest first (Pareto)Which line or press needs me?Status grid: grey when normal, colour when notIs this change real, or noise?Control chart with its limitsAvoid pies, gauges and 3-D: people judge positions on a common scale far better than angles or areas
Choosing a visual. Start from the manager's question; the visual follows.
VisualUse it forAvoid it for
Number tileOne value now, against its target, with a small trend arrowComparing many items
LineChange over time; keep the same scale from day to dayCategories with no order
Sorted bars (Pareto)Which causes, products or presses account for most of the lossValues over time
Status gridWhich line or press needs attention nowShowing how big a problem is
Control chartTelling a real change from normal variation (chapter 3)A single value
TableLists people act on row by row: open holds, overdue work ordersSpotting a trend

A Pareto, drawn to scale

Here is one week of downtime on Line 3. Sorting the causes largest first, with a cumulative line, makes the decision obvious: two causes are well over half the minutes, so they get the engineering time this week.

Pareto of downtime causesMaterial jam 142 minutes; Die change 118 minutes; Sensor fault 64 minutes; Waiting for forklift 41 minutes; Quality hold 27 minutes; Other 18 minutes. The top two causes are 63% of the total. Line 3 downtime, one week: 410 minutes050100150142 minMaterial jam118 minDie change64 minSensor fault41 minWaiting…27 minQuality hold18 minOtherCumulative share (cyan): the top two causes are 63%
Line 3 downtime by cause, one week. Bars are minutes; the cyan line is the cumulative share. Fix the first two bars before anything else.

Adding a widget with an agent

Brief the agent the way Session 12 taught: the goal, the limits and the check that proves it is done. For a widget, the check always includes an access test.

A brief for the agent
Goal: add a widget "Scrap ratio by press, this shift" for supervisors and managers.
Limits: add one view (security_invoker = on) and one entry in widgets.json. Do not change the
widget component, any table, or any row-level security policy.
Done when, from a fresh clone, npm test passes, including the new test that:
  1. signs in as the Line 3 supervisor and sees one row per press on Line 3, with the scrap
     ratio matching the hand-worked value in docs/measures.md;
  2. signs in as a supervisor from another site and gets zero rows;
  3. signs in as an operator and does not see the widget at all.
Exercise · Configure three widgets on your own data40 minutes

You need: Your platform repository; the Supabase SQL editor or a migration file; your AI coding agent; the measure sheet from exercise 1

You will turn three rows of your measure sheet into views and widget entries, built by the agent and checked by you.

Outcome: Three working widgets on the spine's data, each proven to respect the access rules and to compute the number you worked by hand.

Knowledge check

A widget reads a view that was created without security_invoker. What is the risk?

Knowledge check

Which visual best answers 'which causes account for most of our downtime?'

Knowledge check

Why build one widget component driven by configuration, rather than a new page per chart?

References

  1. PostgreSQL documentation: CREATE VIEW (security_invoker). https://www.postgresql.org/docs/current/sql-createview.html
  2. Supabase docs: Row Level Security. https://supabase.com/docs/guides/database/postgres/row-level-security
  3. Supabase docs: Performance and security advisors. https://supabase.com/docs/guides/database/database-advisors
  4. Cleveland, W. S. and McGill, R. (1984). Graphical perception: theory, experimentation, and application to the development of graphical methods. Journal of the American Statistical Association 79 (387). https://doi.org/10.1080/01621459.1984.10478080

Chapter 3 · Layout, signal and freshness

The dashboard people act on

A good dashboard is read in five seconds, from across a room or on a phone in a glove. It keeps colour for trouble, tells real change from noise, and never shows an old number as if it were live.

30 minFive-second testGrey is normalControl limits, not hunchesEvery widget says 'as of'

By the end of this chapter you can

  • Lay out a dashboard for one role so the most important measure is read first, on a wall and on a phone.
  • Use colour only for abnormal states, and never colour alone.
  • Read an individuals control chart and decide whether a change is a signal or noise.
  • Set each widget's refresh and staleness rules, and show the time of its data.

One role, one screen, five seconds

Build one dashboard per role: the supervisor's, the manager's, the planner's. Each answers that role's questions and nobody else's. Then test it the simple way: show it to the person for five seconds, hide it, and ask what they would do next. If they cannot say, the layout is wrong, not the person.

Dashboard layout on a wall and a phoneSix numbered widgets in a three by two grid on a wall display, and the same six in one column on a phone. Wall display (16:9), read from across the roomPhone, 390 px wide1OEE now2Good parts vs plan3Line status grid4OEE trend, 14 days5Downtime Pareto6Open quality holds1OEE now2Good parts vs plan3Line status grid4OEE trend, 14 days5Downtime Pareto6Open quality holdsMost important top left; the phone keeps the same order in one column; every tile is a 44 px tap target
The same dashboard on a wall and on a phone. The order never changes: the most important measure is first on both.
  • Most important first. People read a screen top left to bottom right. The measure tied to the main decision goes top left.
  • Group by decision. Widgets that feed the same decision sit together.
  • Same scale every day. A trend whose axis rescales each morning makes a small wobble look like a cliff.
  • Phone first. On a phone the grid becomes one column in the same order, every tile is a tap target of at least 44 px, and text stays at least 15 px, as Session 10 set out [1].

Grey is normal; colour means act

Control-room displays learned this lesson the hard way. ISA-101, the standard for human-machine interfaces in process automation, describes displays where normal states are drawn in muted greys and bright colour is kept for abnormal conditions and alarms, so the eye goes straight to what needs attention [2]. The same applies to a manager's dashboard. A screen where every tile is green or red trains people to ignore colour.

  • Normal is grey. A measure inside its limits is drawn in the page's neutral colours.
  • One colour for trouble. Amber or red only when a limit is crossed, and the same meaning everywhere.
  • Never colour alone. Add a word or a shape ("Below limit", a triangle), so people with colour-blindness and screens in bright sunlight still get the message [3].
  • Enough contrast. Text needs a contrast ratio of at least 4.5 to 1 against its background [4].

Signal or noise?

Every process varies. If a manager reacts to every dip, they will change things that were fine and make the process more variable, not less. A control chart draws limits from the process's own normal variation, so you can see which points are real signals [5].

For one value per shift (an individuals chart), the limits are: the mean, plus and minus 2.66 times the average moving range (the average of the differences between one shift and the next). The 2.66 is 3 ÷ 1.128, the factor that turns the average moving range into an estimate of the standard deviation [6].

Control chart of shift OEETwenty shifts with mean 74.2%, limits 66.2% to 82.2%; only shift 16, at 64%, falls outside. Shift OEE, Line 3, last 20 shifts (%)606570758085upper limit 82.2lower limit 66.2mean 74.2shift 16: a real signal, find the causeEvery other up-and-down is normal variation: reacting to it makes things worse
Shift OEE on Line 3, with limits computed from the data. Nineteen shifts are normal variation. One is a signal worth a root-cause conversation.
Rule (react when…)What it usually means
One point outside a limitSomething specific happened that shift: find it
Eight or more points in a row on one side of the meanThe process has shifted: a new die, a new material lot, a new crew
Six points in a row steadily rising or fallingA drift: tool wear, a sensor going out of calibration

The first two are among the Western Electric rules, and the third is one of the Nelson rules, both widely used with control charts [10]. Set the widget's limits to control and have the agent compute the limits in the view, from at least 20 points, never by hand.

Fresh, or honestly stale

A number that looks live but is twenty minutes old is worse than no number: someone will act on it. Every widget therefore carries two settings from its configuration: how often it refreshes, and how old its data may get before it says so.

Data freshness and stalenessData arrives each minute for three minutes, stops for eight, and the widget turns grey with an as-of time after five minutes without data. One widget over 12 minutes: refresh every minute, stale after 5 minutes without new data0123456789101112minutesnew data each minutegateway offline: no new rowsgrey: 'as of minute 3'widget shows live valuebackA stale number that looks live is worse than no number: every widget shows its 'as of' time
Refresh and staleness. When data stops, the widget turns grey and shows the time of its last data, instead of pretending.
  • Refresh by asking again. The simplest and cheapest way: the page asks the view for new rows every refresh_s seconds. A minute is enough for most plant measures.
  • Push for events. For a widget that must change the moment something happens (a new quality hold), Supabase Realtime can push table changes to the page, and it respects row-level security [7]. It streams changes to tables, not to views, so measures computed by views are still refreshed by asking again.
  • Show 'as of'. Each view returns the time of its newest row (as_of in the example). After stale_after_s with nothing newer, the widget goes grey and says "as of 06:03".

Keep it honest over time

Dashboards collect widgets the way garages collect boxes. Record which widgets people open (a simple count per widget per week is enough). Every month, show the manager the list: a widget nobody opened, or whose red never triggered an action, comes off. Fewer, trusted widgets beat many ignored ones [9].

Exercise · Build one role's dashboard and run the five-second test35 minutes

You need: Your platform's preview address; your phone; the three widgets from exercise 2; one manager or supervisor

You will arrange your widgets for one role, make them honest about freshness, and test the result with the person who will use it.

Outcome: A one-role dashboard that passes the five-second test on a phone, shows colour only for trouble, and never shows stale data as live.

Knowledge check

On a supervisor's dashboard, what colour should a measure inside its limits be?

Knowledge check

The Line 3 individuals chart has mean 74.2% and limits 66.2% to 82.2%. Today's shift is 70%. What should the manager do?

Knowledge check

The data feed stops for 10 minutes. What should an OEE tile with stale_after_s = 300 show?

References

  1. W3C: Understanding WCAG 2.2, Success Criterion 2.5.5 Target Size (Enhanced). https://www.w3.org/WAI/WCAG22/Understanding/target-size-enhanced.html
  2. ISA: ISA-101, Human-Machine Interfaces for Process Automation Systems. https://www.isa.org/standards-and-publications/isa-standards/isa-101-standards
  3. W3C: Understanding WCAG 2.2, Success Criterion 1.4.1 Use of Color. https://www.w3.org/WAI/WCAG22/Understanding/use-of-color.html
  4. W3C: Understanding WCAG 2.2, Success Criterion 1.4.3 Contrast (Minimum). https://www.w3.org/WAI/WCAG22/Understanding/contrast-minimum.html
  5. NIST/SEMATECH e-Handbook of Statistical Methods: What are control charts?. https://www.itl.nist.gov/div898/handbook/pmc/section3/pmc31.htm
  6. NIST/SEMATECH e-Handbook of Statistical Methods: Individuals control charts. https://www.itl.nist.gov/div898/handbook/pmc/section3/pmc322.htm
  7. Supabase docs: Realtime. https://supabase.com/docs/guides/realtime
  8. Supabase pricing. https://supabase.com/pricing
  9. Perceptual Edge (Stephen Few): dashboard design articles. https://www.perceptualedge.com/articles.php
  10. NIST/SEMATECH e-Handbook of Statistical Methods: What are variables control charts? (Western Electric rules). https://www.itl.nist.gov/div898/handbook/pmc/section3/pmc32.htm

Chapter 4 · 12 questions · 80% passes

Final assessment

Twelve questions across the element. Score 80% (10 of 12) to pass. Your LMS records your score and each answer; you can review the chapters and try again.

15 min12 questions≈ 15 minutesRetake allowed

Choose one answer for each question, then submit. You will see the right answer and why for every question.

1. What must every widget on a dashboard trace back to?
2. Which is a complete measure definition?
3. Planned time 480 min, run time 432 min, performance 95%, quality 98%. OEE is closest to:
4. In the six big losses, where do small stops and reduced speed show up?
5. Which pair is a lagging measure with a leading measure for it?
6. Why do widget views use security_invoker = on?
7. What may the roles field in a widget's configuration do?
8. According to Cleveland and McGill's research, which do people judge most accurately?
9. In an ISA-101 style display, when is bright colour used?
10. For an individuals control chart, the limits are the mean plus and minus:
11. A widget's data feed stopped 12 minutes ago and stale_after_s is 300. What is correct?
12. How should an agent's brief for a new widget prove it is done?