Boards are being asked to be accountable for AI they cannot see. The usual reporting is a slide saying the project is on track, which is an answer to a different question.
Here is the artefact that would actually discharge the duty: one page, monthly, produced from the system rather than written about it.
Why the current reporting fails
It is narrative. Someone describes progress, and the board has no way to test the description. Nothing in it is falsifiable, nothing is comparable month to month, and nothing would look different in the month before an incident.
The gap is that AI oversight requires knowing what the system did, and the people writing the update are reporting on what the team is doing. Those diverge quietly, and they diverge most when things are going wrong.
What belongs on the page
Where AI is running. A list. Which systems, what each one does, and what it can act on as opposed to merely say. This alone surprises most boards the first time, because the list is longer than anyone thought.
Volume and coverage. How many interactions, and what proportion of the relevant customer journeys touched an AI feature. Growth in this number is a growth in exposure, and it should be visible.
Quality, measured the same way every month. The result of the automated evaluation set, plus the human-review rate and the escalation rate. The point is not the absolute figure. It is the trend — a quality number that moved and nobody noticed is the finding.
Incidents and near-misses. Wrong outputs that reached a customer, wrong outputs caught in review, and anything a customer complained about. Near-misses matter more than incidents, because they are the leading indicator and the only one available before something goes wrong.
Changes made. Model versions, prompt changes, new capabilities granted. Each is a change to behaviour, and the board should see that behaviour changed even where nothing broke.
Access and vendors. Who can change these systems, and which third parties see the data. Short, and it should be dull every month — the month it is not dull is the month it earns its place.
Open risks with owners and dates. The list from last month, with what moved. A risk that has appeared unchanged for four months is telling the board something.
Why it must be generated, not written
If a person compiles it by hand, three things follow: it is expensive, so it stops; it reflects what they knew, which is a subset; and it is subject to unconscious smoothing, because nobody enjoys reporting their own bad month.
Generated from the system, it is cheap, complete and awkward — which is the point. A report that can only ever say good things is not oversight.
This is also the honest test of whether the evidence layer exists. If the report cannot be generated, the reason is that the system does not record what would be needed, and that is a finding about the system rather than about the reporting.
What it is worth
Three things. Directors get something specific enough to ask a real question about. The team gets a monthly forcing function that surfaces drift before it becomes an incident. And when a regulator, auditor or acquirer asks how AI is governed here, the answer is twelve dated pages rather than an assurance.
That last one is worth more than it sounds. Evidence assembled after the question is asked is worth very little; evidence that already existed is worth a great deal.
Where to start
Ask for the report next month. Whatever comes back — including "we can't produce that" — is the most useful information you will get about your AI governance this year, and it costs nothing to ask.
Then close the gap between what came back and the list above. For FCA-regulated and compliance-heavy firms we build systems that can produce this from the first commit; where something is already live, a Reality Check establishes what could be evidenced today and what could not.
This is the part we do — the crossing from a demo to a system that survives production.