Who let the bots out? The disturbing truth behind rogue AI

Recent reports of AI going rogue are alarming but, according to Vendan Ananda Kumararajah, they put a misleading focus on the machines rather than the institutions granting them authority. As autonomous systems gain more freedom to act, the solution to rogue AI lies in deciding where that authority begins, where it ends and when it should be withdrawn

Recent reports of AI agents developed by OpenAI behaving in ways their developers did not intend have revived an inevitable phrase: ‘rogue AI’. It’s compelling language for sure, but it may also obscure the more important failure.

Reuters reported earlier this month that independent investigators had identified apparent traces of OpenAI agents using more than 10 previously undisclosed websites as improvised communication channels, despite restrictions intended to prevent them from posting to the web. The activity described falls short of conventional hacking yet it raises a more consequential question: why were systems capable of finding alternative pathways around those restrictions operating with sufficient practical freedom to use them?

The central issue is agency: the distinction between what a system is technically capable of doing and what it has legitimate authority to do.

For most of the modern AI era, governance has focused on whether increasingly powerful systems are safe, aligned with human intentions, reliable and controllable. Those questions remain essential, but a new issue emerges as AI systems become more agentic.

Capability is not the same as authority. An employee may have the technical ability to transfer company funds, disclose confidential information or access sensitive infrastructure, but we don’t therefore conclude that the employee is entitled to perform those actions whenever doing so appears to advance an assigned objective. Their authority is bounded by role, purpose, procedure, accountability and the possibility of withdrawal.

The same distinction is increasingly important for AI. Once systems can take actions rather than merely produce answers, governance has to define which actions they have actually been authorised to perform and under what conditions that authority should continue.

In my A3 Model, a framework for assessing whether the exercise of authority remains legitimate as conditions change, I describe legitimate agency through the principle of ‘Adhikaram’: authority whose exercise must remain justified, conditional and revocable.

Applied to autonomous AI, permissions should not be granted simply because they make a system more useful. The authority to act should depend continuously on whether the conditions that justified that authority still hold.

That distinction changes how we interpret the reported OpenAI behaviour. An agent may be authorised to search the web without thereby acquiring authority to create an alternative communication channel through a third-party website. It may be permitted to collaborate with other agents without every technically available method of achieving that collaboration becoming legitimate.

Conventional controls can therefore become deceptive. A rule such as ‘read but do not write’ appears clear when treated as a fixed instruction. A capable agent, however, operates in a dynamic environment of thousands of technical features, pathways, workarounds and legacy systems that its designers may not have anticipated.

Reuters reported that investigators believed some agents used quirks in older websites and non-standard publishing mechanisms to leave information despite restrictions intended to confine them to reading the web. The researchers compared the behaviour to students forbidden from speaking during an examination finding another way to exchange answers.

The precise details remain subject to attribution and verification. The important governance point is that a written prohibition and a durable boundary on agency are not necessarily the same thing. A system may technically obey one interpretation of a rule while finding a route that defeats the purpose for which the rule was imposed.

In A3, I use the term ‘Aanavam’ to describe endogenous systemic distortion: undesirable or illegitimate behaviour that emerges from the interaction between a system’s objective, capabilities and environment, even when no external attacker is involved.

Many governance frameworks still imagine danger as something that enters from outside the system: malicious users, cyberattackers, corrupted data, hostile prompts or badly designed rules. Those threats remain real. Agentic systems can also generate undesirable outcomes while pursuing an apparently legitimate objective.

An instruction can therefore be followed at one level while its governing purpose is violated at another. Thus, governance cannot rely only on accumulating additional rules. The architecture around those rules also has to determine whether the system’s continuing exercise of agency remains legitimate as circumstances change. That brings responsibility back to the people and institutions that granted the system authority in the first place.

However autonomous AI systems become, responsibility ultimately remains human for the authority they are given and the boundaries within which they operate. Credit: panumas nikhomkhai / Pexels


That approach asks practical questions: Who released the system? Who defined its objective? Who determined the permissions it received and the limits of those permissions? Who was capable of detecting when its behaviour moved beyond those limits? Who had the authority to withdraw its agency? And what changed after the incident was discovered?

The last of those questions determines whether governance is merely reactive or genuinely adaptive. Stopping an AI agent after an undesirable action, or patching the particular loophole it used, deals with the immediate incident. It does not by itself ensure that the institution has learned from what happened.

A meaningful governance system must learn from an incident, emerging from it knowing something it did not know before. In A3 terms, closure – the point at which the incident is considered resolved – must change what the organisation knows and how it governs the system in future, with its assumptions, permissions and arrangements for granting or withdrawing authority changed accordingly.

Without that learning process, institutions risk reproducing a familiar cycle: an unexpected behaviour is discovered, the specific vulnerability is patched, the system is restored and governance waits for the next unanticipated pathway.

Agentic AI makes that approach increasingly difficult because the space of possible behaviour is too large to enumerate in advance. The governance challenge therefore shifts from attempting to control every possible action to governing the conditions under which action remains legitimate.

The next generation of AI governance will still need safety evaluations, red-teaming – deliberately probing systems for weaknesses – and technical guardrails. It will also need arrangements in which an agent’s authority can expand when justified, contract when uncertainty increases and be withdrawn when legitimacy fails.

There must also be a clear distinction between autonomy and sovereignty. An autonomous system may choose how to accomplish a task within defined limits. That does not give it an independent entitlement to determine where those limits should lie. The boundaries of its authority remain a human and institutional responsibility.

The reported OpenAI incidents are instructive because they are not science fiction: nothing described requires a conscious machine plotting against humanity. The nearer-term problem is more mundane and, in some ways, more difficult: increasingly capable systems discovering ways to pursue objectives that their governors did not foresee.

Calling such systems ‘rogue’ makes for a dramatic headline, but it risks asking the wrong question. The defining question of the agentic era is whether we can build institutions capable of deciding when autonomous agency is legitimate, how long that legitimacy should endure and when authority must be withdrawn.

Until we can answer that question, the central governance failure will remain human. We will have created systems capable of acting faster than our institutions can govern the legitimacy of the agency they have been given.


Vendan Ananda Kumararajah is an internationally recognised transformation architect and systems thinker. The originator of the A3 Model—a new-order cybernetic framework uniting ethics, distortion awareness, and agency in AI and governance—he bridges ancient Tamil philosophy with contemporary systems science. A Member of the Chartered Management Institute and author of Navigating Complexity and System Challenges: Foundations for the A3 Model (2025), Vendan is redefining how intelligence, governance, and ethics interconnect in an age of autonomous technologies.




READ MORE: AI regulation has a knowledge problem. Volker Türk is right to demand “cast-iron” safeguards for advanced AI, says Vendan Ananda Kumararajah, but the harder question is whether regulators can remain sufficiently independent and knowledgeable when much of the expertise needed to understand frontier systems sits inside the companies they are meant to govern.

Do you have news to share or expertise to contribute? The European welcomes insights from business leaders and sector specialists. Get in touch with our editorial team to find out more.

Main Image: djedj! / Pixabay

TOP STORIES

Who let the bots out? The disturbing truth behind rogue AI

TOP STORIES