AIBOM Readiness Scorer
READINESS IS THE WEAKEST CATEGORY, NOT THE AVERAGE: a buyer asks for the file, not for a category.
AIBOM Readiness Scorer
Score your AI bill of materials across the six categories from Chapter 8. Get the category to fix first, the estimated time to produce an Annex IV file on request, and whether you would pass a 2026 vendor qualification.
And why the weakest category is the answer. An auditor asks for the file. A file with five strong categories and one absent is a file with a hole in it, and the request will land on the hole. The average of six categories is a comfortable number that predicts nothing, so it is disclosed here and then set aside.
This is the fast diagnostic. The companion AIBOM Readiness Checklist carries the full forty six check assessment and the procurement clause language. Nothing you enter is transmitted or stored.
Models
Weights, architecture, version, parent model lineage and fine tuning history.
Can you state the exact model version in production right now, for every AI system?
Not the family and not the vendor. The version, and the date it was deployed.
Do you record parent model lineage and every fine tuning run?
A fine tuned model inherits the properties and the problems of its parent. Without lineage an incident cannot be traced past the last training run.
For third party models, do you hold the provider’s own bill of materials?
Major model providers have begun publishing structured, machine readable bills of materials alongside model releases. Asking is now a reasonable request rather than an unusual one.
Could you reproduce a decision made by a superseded model version?
Six months after a complaint, the model that made the decision may no longer exist. Retention of superseded versions is the difference between explaining and apologising.
Datasets
Training data sources, licensing terms, preprocessing methodology and any personal data used.
Can you name the training data sources for every model you operate?
Including for models you did not train, where the answer may be that the provider will not say. That is an answer and it should be recorded as one.
Do you hold the licensing terms for every dataset used?
Licence terms are where a bill of materials most often turns into a legal problem, because a dataset licensed for research does not become commercial by being fine tuned.
Is personal data used in training identified, and on what lawful basis?
The Omnibus expressly permits processing special category data to detect and correct bias, which makes the fairness testing position easier and the recording duty no lighter.
Do you record the preprocessing applied to training data?
Two models trained on the same data with different preprocessing are different models. Without this, reproduction is impossible even with the data in hand.
Code
Training pipeline, evaluation harnesses and dependency tree.
Do you have a software bill of materials for the training and serving stack?
This is the one category most organisations already partly have, because software bills of materials have been a procurement requirement for years. Extending it is cheaper than starting it.
Is the evaluation harness recorded as part of the artefact?
An accuracy claim without the harness that produced it is a marketing claim. Annex IV asks for the metrics and the metrics used.
Can you tell which dependency versions were present when a given model was trained?
The library version that produced the weights is part of the provenance, and it changes under you silently.
Do you correlate AI specific risk against your components?
The correlation surface is not the common vulnerabilities feed. It is model card disclosures of failure modes, AI vulnerability databases and the agentic risk catalogue, and it is more fragmented.
Hardware
Accelerator configuration and the infrastructure needed to reproduce or operate the model.
Do you record the accelerator configuration a model was trained on?
Reproduction and cost both depend on it, and it is the category most often omitted because it feels like infrastructure detail rather than provenance.
Do you record what is required to operate the model, as against to train it?
A buyer or a regulator asking about resilience is asking this question. It is also the question that decides whether you could move the workload.
For hosted models, do you know which jurisdiction the compute sits in?
This is where the bill of materials meets the placement discipline, and where an answer of we assume so is not an answer.
Data Processing
Transformations, anonymisation techniques, and output preprocessing at inference.
Are the transformations applied to training data recorded?
Distinct from preprocessing methodology: this is what actually happened to the data, not what the method says should happen.
Are anonymisation or differential privacy techniques and their parameters recorded?
A privacy technique without its parameters is a claim. The parameter is what determines whether the claim holds.
Is the output processing applied at inference recorded?
Filters, guardrails and post processing change what the system does. A bill of materials describing only the model describes something the user never met.
Governance
Model card, bias and fairness evaluation, safety testing results and change history.
Does every production model have a model card?
Including the ones you did not build, where the provider’s card is the artefact and its gaps are your findings.
Do you hold bias and fairness evaluation results?
Article 10 expects examination for bias. The Omnibus now expressly permits processing the attributes you need in order to do it.
Do you hold safety testing results, including adversarial testing?
For an agentic system this includes goal hijacking and tool misuse, which are named risks in the agentic catalogue rather than general robustness.
Is there a change history that records what changed, when, and who approved it?
Annex IV asks for change documentation through the lifecycle. This is the category that makes the other five auditable rather than merely present.
Format, signing, generation and scope
Context rather than category score. These four decide whether the file works at all.
Which specification do you produce, or intend to produce?
Two formats reached production maturity. The choice is not neutral and it is worth making deliberately rather than by whichever tool you found first. CycloneDX ML BOM v1.7, SPDX 3.0 AI Profile
Is the artefact cryptographically signed, and is the key verifiable?
The published baseline is a document signed by the supplier with a key verifiable through known public infrastructure. Suppliers who cannot meet it are being filtered at procurement.
How is the artefact produced?
The single strongest predictor of whether the file is current. A manual artefact cannot keep pace with retraining, and the book’s whole argument is about the difference between months and days.
What proportion of production AI systems are covered?
The three step rollout in Chapter 8 starts with one high risk system, extends to all production systems, then extends to vendor procurement. Say honestly which step you are on. Chapter 8
0 of 26 answered
The ask
The decision, the category to fix, the time to produce a file and the owner, in the order a board wants them. Everything below is the evidence.
Readiness position
Not started
Time to produce the file on request
The figure Chapter 8 frames as the whole point: organisations without coverage face documentation cycles measured in months, those with it complete the same cycle in days. Derived from coverage, generation path and the number of systems in scope.
Unknown
estimated elapsed time for one high risk system
Unknown
estimated elapsed time if a regulator asks for the whole portfolio
Would you pass a vendor qualification
Scored against the published 2026 baseline: a machine readable document in a recognised specification, cryptographically signed with a verifiable key, ingestible without manual transformation, and current.
Not assessed
procurement verdict on the four gate tests
What to fix first, specifically
The lowest scoring checks inside the weakest category. These are the questions where a point is available in the only category that changes your readiness position.
Your format choice, and what it commits you to
Two specifications reached production maturity and they differ in ways that matter once a buyer tries to ingest your file.
The five questions to put to your own vendors
From Chapter 8. These are the procurement filter that separates a supplier operating at the 2026 standard from one operating at the 2024 standard. Ask them in this order.
| Question | What a good answer looks like | What the answer tells you |
|---|
Do your own answers agree with each other?
Individually plausible answers can still contradict one another, and in this area the contradiction usually shows up in front of a buyer.
Confidence in this reading
Graded on whether the score would survive somebody asking to see the artefact, rather than on how good the score looks.
Evidenced
The reading a buyer would take
Self reported coverage discounted, because coverage claimed from memory is the most inflated figure in this category of assessment.
0
challenged readiness, out of 3
None
weakest category on the challenged reading
Board paper
Written to be pasted into a board, risk or audit committee paper without editing. Replace the bracketed fields before circulating.
Turn the score into the artefact
A readiness band is a diagnosis. These are the instruments that close it.
The full forty six check assessment
This page is the fast diagnostic. The AIBOM Readiness Checklist carries forty six scored checks across the same six domains, format selection guidance, readiness bands, and the procurement clause language to require a bill of materials from your own suppliers.
Get the checklistThe cryptographic equivalent
Executive Order 14412 directs guidance on a Cryptographic Bill of Materials, a machine readable inventory of cryptographic assets. The companion workbook scores your readiness against it now, while it is cheap, and answers the FIPS 140-3 procurement question that is already live.
Get the workbookWhere the file gets used
A bill of materials is an input to the Annex IV technical documentation, not a substitute for it. The Regulatory Evidence Pack Template sets out the full structure and the one day test, and the Framework Crosswalk maps each control area to the article, function and clause a buyer questionnaire will cite.
Get the instrumentsA bill of materials is how an audit becomes ordinary operations
Without it, logging captures what happened but not what was running when it happened.
About this assessment. The AI Bill of Materials and its six categories are set out in Chapter 8 of When Agents Rule by Steven Oppenheim. This tool is a self assessment instrument, not an audit, a certification or a benchmark, and it reflects what you tell it: a generous reading produces a flattering result and no useful information. Readiness is deliberately reported as the weakest of the six categories rather than their average, for the reason given on this page. No peer comparison is offered because no dataset underlies this instrument.
On the standards referenced. CycloneDX ML BOM is maintained under the OWASP project and the SPDX AI Profile under the Linux Foundation; both continue to change, so verify the current specification before committing to a schema. The time to produce estimates are derived from your own answers and are planning figures, not commitments. Whether Article 11 and Annex IV documentation obligations apply to a given system is a legal determination for qualified counsel, and Regulation (EU) 2026/1744 moved the high risk application dates to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I. Not legal, technical or financial advice. Nothing you enter leaves your browser: there is no account, no transmission and no storage.