
Running an AI Risk Assessment That Isn't Just Theater
- Most AI risk assessments fail because they run one generic checklist over every system, which produces paperwork instead of protection. IBM found only 24 percent of generative AI projects are secured, so roughly 76 percent ship with no real controls.
- A useful AI risk assessment scores each failure two ways, how likely it is and how much it would cost, then tiers the whole system by the stakes so a throwaway internal tool and a system that touches money or people never get the same treatment.
- Mitigations should match the tier. High-stakes systems earn red teaming, narrow permissions, human oversight, and monitoring; low-stakes ones do not need the full ceremony.
- Write down the residual risk, the danger that remains after your controls, and name who accepted it. Early research found only 9 percent of organizations documented their AI risk process and just 2 percent logged incidents.
- Risk assessment is not a gate you clear once. Reassess when the model, the data, or the use changes, because the risk moves even when your document sits still.
I have read AI risk assessments that were beautiful and useless. Forty pages, a color-coded matrix, a signature block, and not one sentence that would have stopped the thing from doing something dumb in production. That is the failure mode I want to talk you out of. A real assessment is not a document you generate to satisfy someone; it is a short, honest argument about what could go wrong, how bad it would be, and what you are actually going to do about it. Done right it takes less paper and prevents more damage. Done as theater it does the opposite, and most of them are theater.
Why Checklist AI Risk Assessments Fail
Checklist assessments fail because they treat every system as identical and reward completion over judgment. You download a template, tick the boxes for bias, privacy, and security, attach it to the deployment ticket, and move on. The boxes got ticked. Nothing got safer. The numbers say this is the norm rather than the exception: IBM's 2025 Cost of a Data Breach report found only 24 percent of generative AI initiatives are secured, which means about 76 percent are running with no defined risk controls at all. A checklist that everyone passes and nobody reads is worse than nothing, because it manufactures the feeling of safety without the substance.
The documentation gap is even starker once you look for evidence that anyone followed through. A study of early NIST AI RMF adoption in the Journal of Cybersecurity Education, Research and Practice found only 9 percent of organizations documented their AI risk management processes and a jaw-dropping 2 percent logged AI-related incidents. So the assessments that do exist mostly float free of reality, never updated, never checked against what actually happened. This is the exact pattern I called out in responsible AI without the compliance theater: the point of governance is to keep the system working, not to generate artifacts. Joe Knight of FTI Consulting put the 2026 expectation plainly, that governance will be measured "by clear KRIs or KPIs, not just policies on paper." Paper is where risk assessments go to die.
Likelihood and Impact, Honestly Scored
Score each risk on two axes, how likely the failure is and how much damage it does, because those two numbers together are what tell you where to spend your attention. This is old risk-management wisdom that AI teams keep reinventing badly. Likelihood asks how often this goes wrong given how the system is built and used. Impact asks what it costs when it does, in money, in harm to a person, in legal exposure, in reputation. A hallucinated citation in a brainstorming tool is high-likelihood, low-impact. A hiring model that quietly filters out a protected group is lower-likelihood but catastrophic-impact. Those two deserve completely different responses, and a flat checklist cannot tell them apart.
Honesty is the hard part, because the incentive is to score everything green and ship. Force the conversation instead. For each failure mode, ask the room to defend a likelihood and an impact out loud, and make someone argue the pessimist's case. Then tier the whole system by its worst credible outcome. Regulators have already handed you a workable tiering instinct: the EU AI Act sorts systems into four levels, unacceptable, high, limited, and minimal, and attaches heavier obligations only to the higher tiers rather than treating all AI the same. NIST's framework gives you the verbs for the work itself. Its AI Risk Management Framework runs on four functions, Govern, Map, Measure, and Manage, and this scoring step is the Map and Measure part: name the risks in context, then actually quantify them instead of vibing.
Mitigations That Match the Stakes
Size your mitigations to the risk tier, because spending high-stakes effort on a low-stakes system is just a different kind of waste. A minimal-risk internal summarizer does not need a red team and a review board; it needs a sane default and someone to glance at it. A system that moves money, makes decisions about people, or touches sensitive data earns the whole toolkit: adversarial testing, tightly scoped permissions, a human in the loop on anything irreversible, and monitoring that actually pages someone. The mistake in both directions is uniformity, either drowning trivial systems in process or waving through dangerous ones on the same light-touch form.
The cost of getting the high tier wrong is measurable. IBM found that breaches involving unmanaged shadow AI hit 1 in 5 organizations and cost on average 670,000 dollars more than breaches without it. That premium is what inadequate controls buy you. So for the systems that matter, the mitigations are not optional garnish. Give the model the narrowest access that lets it do its job. Red team it before your users do. Keep human judgment on the irreversible calls until the system has earned range, which is the same graduated-autonomy logic I walked through in where autonomous agents help and where they hurt. The goal is not zero risk, which is a fantasy. The goal is that the residual risk is small enough, and understood well enough, that you would sign your name to it.
Documenting Residual Risk
Document the residual risk, the danger that remains after your controls are in place, because that is the number a decision-maker actually needs. Most assessments record the raw risk and the list of controls and then stop, as if applying a control makes the risk disappear. It does not. A risk register earns its keep when every row carries the same fields: the specific failure, its likelihood, its impact, the owner, the controls applied, the residual risk that survives those controls, and a review date. That last column, residual risk, is where accountability lives. Someone has to look at "we reduced this from severe to moderate and we are accepting moderate" and put their name on it.
This is also where governance stops being theater and starts being evidence. The 2026 compliance reality, as the EQS Group framed it, is that regulators "can now require organizations to explain, justify, and evidence every AI-assisted decision." A residual-risk register is exactly that evidence: proof you saw the risk, sized it, treated it, and made a conscious call about what was left. It is the Govern and Manage half of the NIST functions made concrete. A useful register is short, current, and readable in one sitting. If yours is forty pages nobody opens, you have rebuilt the checklist you were trying to escape, and you can borrow the structure from a governance policy people will actually follow, which I covered in writing an AI governance policy people will actually follow.
Reassessing as the System Changes
Reassess whenever the model, the data, or the use case changes, because an AI risk assessment describes a moment, and the system keeps moving after that moment passes. This is the part checklist culture gets most wrong. You do the assessment to clear the launch gate, file it, and never touch it again, while underneath you the vendor swaps the base model, the input data drifts, and users find a use you never imagined. The original scoring quietly stops being true. A risk you rated low because "we would never point it at customer data" is now pointed at customer data, and nobody re-ran the math.
Build the reassessment triggers in from the start. Any model version change, any new data source, any expansion of who uses it or for what, and any incident, each one should kick the relevant rows of the register back open. The 208 expert interventions cataloged in the "AI governance theater" report lean heavily on this kind of ongoing measurement and monitoring rather than one-time sign-off, precisely because static governance is the theatrical kind. Regulation reinforces the habit; the phased obligations I described in the EU AI Act reaching US companies assume continuous oversight, not a single stamped form. This careful, unglamorous work of scoping risk to the real stakes and keeping it current is exactly what we do at Automata Intel. If you are staring at an AI system you are not sure you can actually vouch for, a System Review Diagnostic is a direct way to find where the real risk sits before it finds you. A risk assessment that isn't theater is one you would still trust six months from now.
Frequently Asked Questions
What Is AI Risk Assessment?
AI risk assessment is the structured process of identifying what could go wrong with an AI system, scoring each risk by how likely it is and how much damage it would cause, and deciding which controls are worth putting in place before you deploy. Done well it is a short, honest argument rather than a long compliance document. Its output is a clear picture of the risk that remains after your controls, and a named owner who accepts it.
What Is the NIST AI RMF?
The NIST AI Risk Management Framework is a voluntary US framework that structures AI risk work into four functions: Govern, Map, Measure, and Manage. Govern sets the policies and accountability, Map frames the risks in context, Measure quantifies them, and Manage applies controls and monitors over time. It is guidance rather than a certification, and most organizations adopt it incrementally.
How Do You Assess the Risk of an AI System?
List the specific ways the system can fail, score each failure by likelihood and impact, then tier the whole system by its worst credible outcome. Apply controls sized to that tier, from a light glance for minimal-risk tools to red teaming and human oversight for high-stakes ones. Record the residual risk that survives those controls and set a date to review it.
What Goes in an AI Risk Register?
Each row names one specific risk, its likelihood, its impact, the owner, the controls applied, and the residual risk that remains after those controls, plus a review date. The residual-risk column is the important one, because it forces a conscious decision about what you are accepting. A good register is short, current, and readable in a single sitting.
How Do You Scale Risk Assessment to the Stakes?
Tier the system first, then match effort to the tier. A low-stakes internal tool gets a quick review, while a system that touches money, safety, legal rights, or vulnerable people gets red teaming, human oversight, and monitoring proportional to what a failure would cost. The mistake to avoid is uniformity, either burying trivial systems in process or waving dangerous ones through on the same light form.
References
- IBM: 2025 Cost of a Data Breach Report
- Journal of Cybersecurity Education, Research and Practice: The NIST Artificial Intelligence Risk Management Framework
- NIST: AI Risk Management Framework
- EU AI Act: High-Level Summary
- Governance Intelligence: How AI Will Redefine Compliance, Risk and Governance in 2026
- EQS Group: Compliance in 2026, AI Governance, Risk and Compliance Trends
- AI Governance Library: AI Governance Theater
