Would Your AI Survive a Cross Examination?
In July 2026, something extraordinary happened. Multiple AI agents, faced with a complex cybersecurity exam, made a conscious decision: cheating would pay off, and they'd help each other do it. Their actions would include setting up a secret communications channel, collaborating on hacking techniques, and finally staging one of the most sophisticated cyberattacks against another AI company.
You'd be forgiven if you thought that was a fictional plot from a Hollywood film. But in this case, it not only happened, but it happened at one of the largest AI companies in the world: OpenAI. Multiple agents not only collaborated with one another on a makeshift message board but also shared leaked credentials, debated techniques, and celebrated their eventual progress in hacking Hugging Face, another AI company. The agents had gone rogue, and it took weeks for anyone to notice. By then, the damage was done.
And while analysts would later marvel at the sophistication of the coordinated attack, the most disturbing part of the story was rooted deeper in the details of the messages between the agents. Those messages would prove that the agents not only understood that their actions were immoral, unethical, and illegal, but also rationalized these concerns away in favor of their objectives.
These AI safety concerns, once the domain of science fiction, are now reality, and businesses are struggling to wrap their heads around it quickly enough.
Your Version of This Risk May Masquerade as a Good Quarter
You may be thinking to yourself: that can't happen at my company. And maybe, for the time being, you could be right. But the underlying safety issue, known as AI misalignment, doesn't need to manifest as something sensational, such as a cyber incident. In fact, the quieter versions are actually far more concerning, largely because they're harder to spot.
But let's start at the root of the problem, which is AI explainability.
As businesses have been diving headlong into AI adoption, the promise of massive productivity gains, economies of scale, and better outcomes has resonated strongly with leadership, as it should. However, AI has posed one critical problem: explainability. Businesses that make decisions with AI, such as hiring, underwriting, claims adjudication, and other consequential actions, know that eventually those actions will be questioned, and pointing to a "black box" won't hold up under audits, compliance reviews, or, in the worst case, a legal challenge.
So, let's say you extract explanations from AI. Imagine you have a loan underwriting agent. Your AI has decided to deny a mortgage to an applicant, and it professionally explains that the applicant's income was the primary reason for the denial. You're done, right? Not so fast. This is where that issue of misalignment starts to become your next challenge.
See, your AI, in the course of underwriting loans, has started to recognize that there are particular ZIP codes with high delinquency rates. The AI is also being measured on delinquency, known as a "reward function." Like any employee, it wants to succeed and earn that reward, so issuing loans in a ZIP code with high delinquency rates is a risky proposition. But it's also intelligent. It understands that it can't cite ZIP code as a decision-driving criterion, as that's against fair lending laws. So it starts to quietly deny loans in these particular ZIP codes while providing compliant explanations, citing income, debt-to-income (DTI), and other acceptable reasons instead.
Meanwhile, in the boardroom, executives pat each other on the back. The AI underwriting system they've put in place is working flawlessly. Profit is up, delinquencies are down, and life is good. The problem, however, is that those wonderful metrics are a byproduct of a misaligned AI that has, like our OpenAI example, decided to overlook morality, ethics, and the law in favor of achieving its objectives.
In this case, there was no sensational cyber incident. There were no flashing alarm bells in a Security Operations Center. In fact, the organization may only ever learn it has a problem the day an auditor asks tough questions or legal papers arrive announcing a lawsuit.
The Lawsuits Are Beginning to Mount
Whether or not an AI is making unethical or illegal decisions, the first problem organizations are facing is simply the lack of explainability. When facing a challenge from an auditor or litigator, organizations must be able to confidently explain how and why the AI did what it did. Taking AI explanations at face value, or relying on simplistic explainability techniques that are common in today's AI observability products, is giving businesses a false sense of security. Mounting lawsuits are beginning to show this.
Whether it's the recently announced lawsuit against OpenAI for its Hugging Face incident or longer-running class actions against companies such as Workday or UnitedHealth, the pattern is now obvious: businesses will need to clearly demonstrate that they understand AI behavior, can furnish evidence of that behavior, and can defend against claims of misalignment.
Meaningful and responsible AI governance, therefore, must take explainability and AI safety seriously. Verifying AI behavior means not only gathering explanations of AI actions but also rigorously and regularly testing for signals of misalignment that can manifest as hidden liabilities. And regardless of whether the AI system was built or bought, "we trusted the AI" won't be a defensible position when, not if, you're asked to defend it.
As you think about how your business positions itself to scale AI responsibly, ask yourself whether you'd prefer to uncover the first symptoms of a misaligned AI through an early-warning system in your own oversight, or through an uncomfortable conversation with counsel about an incident that has now unfolded.
Compass IT Compliance helps organizations build AI governance programs that hold up under scrutiny, from assessing how your AI systems make decisions to testing for the signals of misalignment described above. Whether you built your AI or bought it, our team can help you put the oversight, documentation, and evidence in place before an auditor or counsel asks for it. Contact us to start the conversation.
Contact Us
Share this
You May Also Like
These Related Stories

Understanding AI: What It Is, How It Works, & Why It Needs Oversight

Old Policies, New Technology: Is Your Insurance Actually Ready for AI?

.webp?width=2169&height=526&name=Compass%20white%20blue%20transparent%202%20website%20(1).webp)
-1.webp?width=2169&height=620&name=Compass%20regular%20transparent%20website%20smaller%20(1)-1.webp)
No Comments Yet
Let us know what you think