Organisations may have detailed AI risk taxonomies, policies and registers yet still leave people uncertain when a live issue appears. A category such as misleading output or insufficient human oversight gives the issue a name without revealing when it matters, who should examine it or what evidence would justify proceeding.
A map can lose the reasoning behind it
“Death by a thousand taxonomies?”: AI Risk Classification In Practice, by Berman and colleagues, examines how classifications of sociotechnical AI risk are developed and used. Drawing on 25 interviews with researchers and practitioners, the authors report that these classifications are often weakly connected to wider governance processes.
The 25 interviews do not establish how common the problem is across enterprises, but they show how a taxonomy can lose meaning as it travels. Downstream users may see the categories without the choices about which harms to include or how to group them, making an interpretive framework look like a complete inventory of risk.
Taxonomies also describe harmful outcomes without always tracing them to the actors and decisions that could produce them. An issue can be classified correctly while the basis for intervening remains unclear.
Taxonomies remain useful as shared language that improves coverage, but the study’s conclusion is narrower than a prescription for how organisations should govern: classifications need to be better integrated with governance infrastructure.
Tooling covers only part of the problem
In Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools, Alam and colleagues assessed 21 open-source tools against 32 risk-mitigation categories.
Coverage was strongest in areas suited to technical implementation, including model evaluation, adversarial testing, content filtering, runtime guardrails and observability. Governance oversight, legal intervention and financial or market remedies received much less coverage.
The study examines selected open-source tools rather than the whole governance market; its mapping was partly LLM-assisted and reviewer agreement was moderate. Because it mapped documented capabilities rather than testing their effectiveness in production, its finding is necessarily narrow. The authors propose combining tool-based controls with organisational and regulatory processes, but they do not test which governance workflows are most effective.
Marketing decisions rarely arrive in neat categories
Taken together, the studies identify a gap rather than validate a particular solution. In regulated marketing, that gap becomes visible in the evidence, responsibilities and decisions surrounding a live communication. A framework may identify misleading claims and inadequate substantiation as material risks when assessing a numerical claim in a financial promotion. The marketer still needs to know which evidence is authoritative, whether the figure can be supported and whether its qualification remains clear in the proposed format. A reviewer may also need to establish whether an earlier approval covered the same product, audience and channel.
The same issue arises when an AI governance framework identifies insufficient human oversight. That label does not determine which communications require review, who should receive them or what authority the reviewer has.
Some communications may need a specialist because they introduce a new claim or affect customers who could be vulnerable, while others may repeat established wording within approved conditions. Effective oversight depends on reviewers understanding why an item was referred, which requirements apply and what evidence supports it, as well as having the authority to approve, amend, reject or escalate.
These controls should reach earlier than final approval because risks become relevant while an audience is selected, evidence is chosen or an AI system is given permission to act. Waiting for the finished communication makes important context harder to recover and objections more expensive to address.
Scale makes selective escalation essential
Salesforce’s 2026 Agentic Enterprise Index analysed organisations running Agentforce agents in production every month between February 2025 and April 2026. Within that cohort, the average number of activated agents nearly tripled and the average agent’s available skills increased from two to six during 2025. The ratio of external action calls to generated text grew at a 15% compound monthly rate.
This is Salesforce’s analysis of its own platform rather than a picture of enterprise adoption as a whole. It nevertheless illustrates the pressure created when AI systems begin performing tasks across business systems rather than simply generating drafts.
Governance at this scale depends on explicit boundaries, allowing familiar activity to proceed when the requirements, evidence and permitted response are settled. A new claim, audience, product change or conflict between sources can send the work to someone authorised to make a fresh judgement.
Labels such as “high risk” and “low risk” are not enough. Teams need to know which circumstances move a case between them, what happens while a decision is pending and who may accept an exception.
Approval should leave a useful record
A reviewer might accept a claim for a particular audience because a named source supports it, provided a qualification appears prominently. Remove those conditions from the record and the decision becomes easy to misuse.
A useful decision history preserves what was accepted, the evidence considered and the circumstances that limited the approval, while also showing when something has changed. New product terms, revised regulation or stale substantiation can make yesterday’s sound judgement unsuitable today.
Past decisions should provide context rather than permanent permission because their value lies in showing how the organisation approached a comparable issue and where the new case differs.
As AI takes on routine work, organisations need to direct specialists towards the decisions that require them and retain the reasoning. Otherwise difficult questions will be answered from scratch at greater speed and volume.
What teams need to know
Why is a risk taxonomy not enough?
It provides shared categories, but the research found that design choices can become invisible and links to decision points or responsible actors are often missing.
What should AI governance add?
It should connect each material risk to a workflow stage, applicable requirements, supporting evidence, an accountable decision-maker and a permitted response.
Does every AI-assisted activity need human approval?
No. Established, lower-risk activity can proceed within defined boundaries while novel, ambiguous or consequential cases are escalated.
Why retain previous marketing and compliance review decisions?
Earlier reviews show how a claim, communication or AI-assisted action was assessed and what conditions shaped the decision. Keeping that evidence, context and rationale helps future teams judge whether the decision still applies or needs to be reconsidered.
When should an AI-assisted workflow escalate to an expert?
Escalation is appropriate when a case introduces a new claim, changes the audience or product context, relies on conflicting evidence or could cause material harm, with the relevant evidence and requirements included in the referral.