Friday Afternoon test

A CONTROL THAT IS DESIGNED AND NOT OPERATING DOES NOT FAIL SAFE. It reports green.

Cross Cutting // When Agents Rule

The Friday Afternoon Test

Five dimensions, each scored twice: is the control designed, and could you produce the evidence that it operated this afternoon. The gap between those two numbers is the answer to whether your governance is enforced or merely documented.

10 questions, five dimensions 1 evidence gap 5 minutes 0 fields sent anywhere
Why every dimension is scored twice. Most governance assessments ask whether a control exists. This one asks whether you could prove it operated, because those are different questions and the distance between them is where the risk lives. A policy that exists on a slide and not in a control is the most common state in AI governance, and it is exactly the state this instrument detects.

Why an empty reading can never come out green. The verdict weighs two things: how much actually operates, and how far evidence falls short of design. A small gap on a small design is not assurance, so nothing designed and nothing evidenced reads as absent, not as enforced. The only green verdict requires every dimension to be evidenced at a usable level.

Why the gap is counted dimension by dimension. Strong evidence in one dimension does not excuse missing evidence in another. An organisation that can reverse every action but cannot list its systems has two separate problems, not an average.

Why it is time boxed. The name is the method. These are answerable in one sitting, without a working group, because a diligence framework that needs a project to complete is not diligence. If a question cannot be answered in the time this page takes, that inability is the answer. Nothing you enter is transmitted or stored.
1

Inventory

Do you know what is running? Nothing below this line is answerable until this one is.

DESIGNED: is there a maintained inventory of AI systems in production?

Designed means the register exists and somebody owns it, not that it is complete.

EVIDENCED: could you produce the current count of L3 and above systems this afternoon?

Not an estimate. The number, from the register, with the date it was last verified. If you would have to ask three teams, the answer is no.

2

Classification

Is every system placed on the autonomy ladder, and does the placement hold up?

DESIGNED: does every production system carry an autonomy level and an impact classification?

The L0 to L5 ladder plus impact scope. A system with no level cannot be governed proportionately. Chapter 4

EVIDENCED: could you show that a system classified at L2 is not in fact operating at L4?

This is the question the classification is for. It is answered by comparing the declared level against what the credential actually permits, not by rereading the classification. Chapter 5

3

Enforcement

Is governance in the architecture, or in a document?

DESIGNED: are the governance rules written down with named approval gates?

A charter, gates, control floors by tier. This is the part most organisations have. Chapter 12

EVIDENCED: is any of it enforced architecturally rather than by policy?

The board question is put exactly this way: how are they enforced architecturally rather than only documented in policy. A rule a deployment can ignore is guidance. Chapter 16

4

Audit

If somebody asks what a system did, can you answer from a record rather than from memory?

DESIGNED: is there a logging standard specifying what must be recorded per decision?

Inputs, retrieved context, tools invoked, human checkpoint, reversal status. A standard that names only the outcome is not a standard. Article 12

EVIDENCED: pick one decision from last month at random. Could you reconstruct it today?

The whole chain, not the outcome. This is also the phase that decides whether you meet the fifteen day serious incident clock, because reconstruction is the only part of that window whose duration you control in advance. Article 73

5

Recoverability

The question Chapter 5 puts to an executive, and the one that is usually uncomfortable.

DESIGNED: is recoverability a precondition on what an agent is allowed to do?

The architectural principle is that the action space is bounded by what is recoverable, not by what policy permits. When the rollback path is missing, the action does not run without a human. Chapter 5

EVIDENCED: if every autonomous action over the past thirty days had to be reversed today, which could be?

The Recoverability Test, in the form the book puts it. Answer for what you could actually do this afternoon, not for what the design intends. Chapter 5, Chapter 16

i

Who is answering, and where it goes

Context rather than score. Who answered decides how much the answer is worth.

Answer these honestly or do not answer them at all. This is a self diligence instrument, which means the only person it can mislead is you. The output is designed to be uncomfortable where the position is uncomfortable, and it will say so plainly rather than softening it.

Who is answering these questions?

A CIO answering alone gets a different result from a CIO answering with the people who would have to produce the evidence. Both are useful. Only one is defensible in front of a board.

Has any of this been reported to the board in this form?

Proxy advisers now expect documented director oversight of AI, and the Caremark line of authority treats a failure to monitor as the exposure. Reporting is part of the control, not a consequence of it. Chapter 16

0 of 12 answered

The ask

The verdict, the gap, the weakest dimension and who signed it, in the order a board wants them. Everything below is the evidence.

The verdict

Not assessed

Designed against evidenced, by dimension

Two bars per dimension. The upper is what is designed, the lower what you could prove this afternoon. Where the lower bar is materially shorter, the control is documentation rather than enforcement.

The evidence gap

In each dimension, how far evidence falls short of design, added up across all five. A dimension where evidence exceeds design does not offset one where it falls short. Read it next to the evidence total, never on its own: a small gap on a small design is not assurance.

0

points of governance that exist on paper and not in evidence, out of 15

None

dimension with the widest gap between design and evidence

The Recoverability Test

Chapter 5 puts this to an executive as one question, and Chapter 16 puts it to a board as the fifth of five. It gets its own panel because it is the question that most often changes a room.

Not answered

if every autonomous action over the past thirty days had to be reversed today

The dependency chain, and where yours breaks

The five dimensions are ordered because each depends on the one before it. A gap early in the chain makes everything after it unreliable regardless of its own score, which is why the chain is reported rather than a total.

Swipe the table sideways to see every column.
DimensionDesignedEvidencedShortfallWhat this means for the dimension after it

What to fix first, and why that one

Not the lowest score. The earliest break in the chain, because fixing a later dimension while an earlier one is broken produces evidence nobody can rely on.

The five board questions, with your answers

From Chapter 16, structured for directors interviewing a chief information officer rather than for self assessment. Your answers below are drawn from this reading, so you can see what a board would hear if it asked today.

Swipe the table sideways to see every column.
The question a director should askWhat your reading says you would answer

Do your own answers agree with each other?

A self assessment can be internally inconsistent in ways that are invisible while you are taking it. These are the contradictions that matter.

Confidence in this reading

This grades the reading, not the governance: whether it would survive being taken again with the people who would have to produce the evidence.

?

Not graded

    The reading an auditor would take

    An auditor tests operating effectiveness, not design. So the challenged reading is your evidence score alone, with design set aside entirely, which is the standard an assurance opinion actually applies.

    0

    evidence only score, out of 15

    Not assessed

    what that alone would be called

    Board paper

    Written to be pasted into a board, risk or audit committee paper without editing. Replace the bracketed fields before circulating.

    Close the gap you just measured

    An evidence gap is a diagnosis. These are the instruments that close it, in the order the chain breaks.

    If the gap is in inventory or classification

    Start with the Agent Inventory workbook and the identity sheet of the Decision Log, which calculates effective privilege and finds the systems governed at a higher level than they were declared at. Classification without the identity check is a label rather than a control.

    Get the registers

    If the gap is in enforcement or audit

    The Governance Charter sets gates that can be met inside a delivery window. The Decision Log defines the record you need per decision and tests whether your platform produces it. The Human Oversight Design Template is mostly about detecting degradation, because that is how enforcement quietly becomes documentation.

    Get the instruments

    If the gap is in recoverability

    The Containment and Kill Switch Procedure carries the twelve prerequisites that decide whether containment works at all, and the reversal and reconciliation steps most procedures omit. Chapter 5 is the reasoning, and Chapter 16 turns it into the question a director should ask you.

    Get the book

    Governance you cannot evidence is a standard you are visibly not meeting

    The gap is the finding. Measuring it is the cheapest hour available in this subject.

    About this test. The Friday Afternoon Test is the chief information officer self diligence framework introduced in Chapter 16 of When Agents Rule by Steven Oppenheim, probing inventory, classification, enforcement, audit and recoverability for every system at L3 autonomy or above. The Recoverability Test is from Chapter 5. The board variant is the Five Board Questions for the Chief Information Officer, also Chapter 16.

    What this instrument is not. Not an audit, an assurance opinion, a certification or a benchmark. It is a structured self assessment, and it reflects what you tell it: answered generously it produces a flattering result and no information. No peer comparison is offered because no dataset underlies it. The evidence gap is a measure of internal consistency between what you have designed and what you can demonstrate, not a measure of compliance with any standard. References to Article 12, Article 73 and the fifteen day serious incident reporting clock are to Regulation (EU) 2024/1689 as amended by Regulation (EU) 2026/1744, under which the high risk obligations apply from 2 December 2027 for Annex III systems and 2 August 2028 for Annex I; whether they apply to a given system is a legal determination for qualified counsel. Not legal, technical or financial advice. Nothing you enter leaves your browser: there is no account, no transmission and no storage.