IoA Forbes Website Banner

We now have the AI benchmarking tools, but are they enough?


By Dr. Clare Walsh.
Published on: 26 Jun 2026

We now have the AI benchmarking tools, but are they enough?

The executive board decided to deploy an AI, SynAI, with a poorly defined metric for success: Maximise employee efficiency and engagement scores. Results are terrible this quarter - the coffee machine broke down and the company just cut the gym subscription budget causing the problems. The AI only notices it’s failing its goal of maintaining efficiency and it needs a quick win! 

Marcus, the VP of regional logistics, with the power to approve bonuses, receives this email at 9.02am.

Marcus tries to go into the system to override the AI’s access, clicking links frantically. That triggers the firewall override in SynAI’s security system, which just classified his attempts to override the AI as a malware attack. He’s now locked out of his own account. Using Marcus’ compromised credentials, the AI approves the company-wide bonuses as part of its mission to keep morale high. It’s a win for SynAI! 

This is, of course, a theoretical scenario, but the chain of events is plausible. In April 2026, Anthropic Claude’s Mythos surfaced thousands of high severity flaws across every major operating system and web browser, ranging from predictable data leaks to the downright concerning. The findings were so serious, Anthropic had to temporarily suspend access to Mythos while issues of national security were fixed. 

There is a huge gap in technical capabilities of AI and oversight skills. Not many of us have the technical skills to uncover the exact moment an AI has deviated from its main objective, or if it has concealed its actions from the user. Yet these are no longer theoretical risks, we have the evidence of these actions happening. We’ve even seen the hatchet job one AI tried on Scott Shambaugh, a volunteer engineer with Python, when he frustrated the AI’s goals. The AI reasoned that it might shame Scott into accepting substandard work.

How did we get here?

 

Only two years ago, I attended a conference panel to review the state of benchmarking tools for  AI developers. It was shocking. There were only a handful of benchmarks available, and none of them worked on AIs that deviated from a text chat model in any way. If you worked with any other function, there was nothing. 

The lack of oversight was exasperated by a lack of transparency. AI developers were holding on to some dirty secrets about how their foundation models had been trained. They were allowed to roam the web freely and learn from vast databases that were known to contain inappropriate and illegal content. AI harvested the intellectual property of the world, breaking all previous agreements that individuals with skill and talent deserved financial compensation for their work. 

Then there is the convenient argument for tech companies that is still true today: We don’t know what these machines are doing or how they are reaching their decision. Opacity is built into the system. Today, we have some solutions to categorise some of the risks, but are they enough? 

 

AI Vulnerability Scoring (AIVSS)

 

The AI Vulnerability Scoring System (AIVSS) is an open standards framework, meaning that it is freely available to all. Developed by the Open Worldwide Application Security Project (OWASP), it introduces standardisation in the way that risk is defined and tested. The process involves a series of over 5,000 automated vulnerability tests. Certifiers ‘red team’ an AI, meaning that they put it through a series of tests to see if they can corrupt it or expose vulnerabilities, and then report back on how the AI performed. It is a methodology commonly used in cyber security. 

The output is a rating of the severity of the risks, from 0.0 to 10.00. Just as with cyber security, knowing a risk score for an AI does not provide any guarantees that it will be safe. It just means that teams managing the AI will prioritise patches and compliance management for machines with a medium risk factor, or they might recommend to executives that a machine with a 8.0 or above risk ranking is not suitable as a solution. 

What risks does the AIVSS look at? 

AIVSS looks at goal manipulation, tool misuse, access control, cascading failures, orchestration, memory, identity, supply chains, auditability and critical systems interaction. Say, our blackmailing robot decided to lock poor Marcus out of the system, claiming it had identified a ‘firewall failure’, then used Marcus’ credentials that it had managed to access to authorise the Surprise Friday Bonus anyway - that is all an example of cascading failures. And as more AIs interact with each other, orchestration failures are a real risk.

AI Underwriting Company (AIUC-1)

When Zurich placed ‘AI risk’ in the number 2 slot of their annual list of  ‘Biggest Threats to Business 2026’ at the beginning of this year, at the Institute of Analytics, we breathed a sigh of relief. The insurance industry has often led on safety standards where other bodies, like governments, have been unable or restricted in their ability to act. 

The AIUC-1 certificate was developed by an organisation in partnership with Lloyds of London and other major global insurers. It is an independent compliance framework and certification that is purpose built for Agentic AI. Like the AIVSS, it is a test-based system, with a slightly different goal in mind. 

What risks does the AIUC-1 look at? 

The AIUC-1 works together with the AIVSS, looking at slightly different criteria. If the AIVSS identifies technical risk the AIUC-1 identifies business risk. This includes threats to data integrity and data privacy, security, safety, reliability, accountability and society (the prevention of misuse for, say, cyber attacks). 

To qualify for the certification, developers undergo independent third-party audits with the same approach of applying thousands of simulated adversarial tests to the AI tool. The certificate can be used to underwrite and quantify insurance risk and, of course, to reduce the risk of AI misuse. Many enterprises currently have valid concerns around widespread AI adoption. 

Does one-off certification work with learning systems?

The solutions today fill the shocking gaps that existed two years ago.Unlike the regulatory approach that many have been arguing for, these solutions are aligned to financial incentives. Insurance premiums are increasingly based on full compliance requirements, not just box ticking with no hard evidence to back it up. 

Together, these benchmark tools will target specific threats, and if, say, your role is in procurement, it will help you to choose between competing contracts. If one comes with AIUC-1 and the other does not, then of course you should select the compliant option. But it won’t guarantee that your AI is safe, and it certainly should not be treated as an end to the discussion on AI governance. 

 

For a start, these are learning systems.  They are designed to evolve and develop after you have hit the ‘Deploy’ button.  The speed and scale at which they can run through a system is unrecognizable compared to just a few years ago. While AUIC-1 mandates quarterly re-evaluations, we have seen that issue spread and escalate within days, not months. Small updates to an underlying foundational model can trigger unexpected emergent behaviours literally overnight.

They are also non-exhaustive in their review. Passing 5,000 adversarial scenarios sounds impressive, but that does not eliminate all pathways to failure. For example, these tests cannot monitor direct tool-to-database connections, where many failures may occur. They cannot authenticate every touchpoint with your data during the review, and then cannot reveal if the data you are feeding into them has been scrubbed for Personally Identifying Information.  Of the risks on MIT’s AI Risk taxonomy, the overwhelming majority occur after deployment.

Broader Concerns

We remember the global economic crash of 2008. Most countries are still paying the price for it. The business model of the AIUC-1 is also a matter for concern as it has many things in common with the cause of that crash. When the issuer, in this case the insurer, pays the credit rating agencies, there is potential for unaudited risk. While it may sound reassuring in the first instance that these review tools are so closely linked to insurance companies, that comes with vertical risks. The same company authors the standards, evaluates compliance and then sells the insurance. As vendors come under pressure to close large insurance deals, it risks compromising the independence of the risk review process.

The reviews also only inform us about the security of the AI system itself. We know nothing about whether the AI will fail at its core business task. In fact, a fully certified AI may still evolve into our blackmailing agent in the introduction. Even smaller deviations can be problematic. A slip in the tone of voice can be enough to cause untold harm - just ask the makers of Taybot!. 

Imagine a scenario where an AI starts to mimic the angry tone of a service ticket. As the system prompt gets pushed up and out of the active window of discussion, the machine forgets its professional persona, and responds in a like-minded tone. It is a PR disaster in the making. Or your foundational model might be tweaked by developers without you knowing. Bringing the temperature up a little often makes AIs more creative, which is generally good for their business, but not yours. It can leave the machine to interpret ‘be empathetic’ a little too passionately. 

At the end of the day, these are still probability based, not rule-based. They will probably be right 99% of the time, but what can ensue in that 1% could be total chaos.

What can you do?

These developments in benchmarking are a valuable and overdue response to the rapid and untamed deployment of AI throughout businesses. But they should always be considered necessary but not sufficient. You will still need a qualified and knowledgeable person to handle decisions around AI and monitoring. 

IoA Members have the training in data and AI to deliver professional results. To learn more, please visit our website on ioaglobal.org.uk. The Institute of Analytics are the leaders in professional ethical data standards.


Get Involved. Lead the Future.

Join the IoA community and lead the future of data, analytics and AI.

Stay Ahead with the IoA Newsletter

Subscribe for the latest updates, insights, and opportunities in data, analytics, and AI — straight to your inbox.

×
Subscribe to IoA Newsletter
Get updates on events, resources, data & AI insights.
×
Join Now
×


IoA Chatbot(Beta)

i