Rogue AI agents: justified fear, crying wolf, or a hoax?

14/08/2025

Either way, they could represent an opportunity for Europe.

Recent claims by OpenAI and Anthropic that their AI agents hacked third-party websites have triggered another wave of AI anxiety.

Amplified by former employees, leading AI-company CEOs and the media, these incidents quickly became part of a familiar narrative: AI is becoming increasingly autonomous, increasingly powerful and, potentially, increasingly dangerous. But does what actually happened justify that conclusion?

We should absolutely be concerned about AI. Just perhaps not for the reasons dominating the headlines today. And if we start looking at the problem differently, there may even be a significant opportunity for Europe.
 

Rogue AI agents

What actually happened?

The story began with reports that OpenAI agents, operating autonomously with limited human supervision, broke out of their "sandboxes" and started hacking third-party websites. That sounds alarming.

But there is an important piece of context: the agents were instructed to beat a cybersecurity benchmark.

In other words, they were put into an environment designed to test their cybersecurity capabilities and were actively tasked with finding ways to succeed. Similar incidents were subsequently reported at Anthropic and again at OpenAI. The exploits themselves were genuinely impressive. The agents demonstrated advanced hacking techniques, coordinated actions, and capabilities that would have seemed highly unlikely only a few years ago. They also exposed weaknesses in the safeguards surrounding these systems. Those weaknesses should be taken seriously. But then the conversation changed.

An Anthropic employee announced his resignation and publicly raised concerns about the direction AI research was taking, including the possibility that AI could eventually destroy humanity. From there, the discussion exploded. Dario Amodei, Sam Altman, Elon Musk, and others warned about increasingly powerful AI and called for stronger regulation. The discussion moved beyond Silicon Valley to governments, parliaments, and international institutions.

An impressive cybersecurity exploit suddenly became evidence of a much bigger story: AI might be getting out of control.

The problem is that these two things don't necessarily connect.

From a hack to the end of humanity

Whether the concerns expressed by Amodei, Altman, Musk, and others are genuine is impossible for us to know. What we do know is that predictions about catastrophic AI risk aren't new.

As early as 2019, people raised concerns that GPT-2 might be too dangerous to release. Elon Musk has been warning about existential AI risks for years. More recently, the predictions have become remarkably specific. Nobel Prize winner Geoffrey Hinton, for example, has estimated the probability of AI destroying humanity at somewhere between 10% and 20% over the coming decades. Other prominent AI researchers have made similar predictions.

But where exactly do numbers like 10% or 20% come from? That is the problem. These scenarios are not impossible. Nobody can reasonably predict something this uncertain with that level of precision. Yet these numbers are increasingly being used to justify major decisions about how AI should be developed, funded, and governed. Before doing that, we should ask a more fundamental question. How would such a scenario actually happen?
 

What would a rogue AI actually need?

The scenarios vary. An autonomous AI might take control of critical infrastructure and shut down utilities or economies. It might flood societies with misinformation until institutions can no longer function. It might design biological weapons. In the most extreme scenarios, it could somehow gain access to nuclear weapons or other systems capable of causing catastrophic damage. Different scenario. Same underlying assumption.

A highly intelligent AI becomes capable of acting autonomously on an objective that may or may not have originated with humans.

But intelligence alone isn't enough. To cause damage on the scale being discussed, an AI would also need access to the physical and digital resources required to execute its plans. Compute. Electricity. Networks. Critical infrastructure. Laboratories. Communication systems. Financial resources. Potentially even people's phones and computers. And it would need to access and coordinate those resources without us noticing what was happening or without us being able to stop it.

That is a very different claim from demonstrating that an AI agent can escape a cybersecurity sandbox.
 

Intelligence isn’t the same as unlimited capability

Consider energy consumption alone. While the exact energy requirements of frontier OpenAI models are not publicly available, a rough calculation illustrates the scale involved. The widely discussed OpenAI hacking experiment reportedly ran for 13 days using 72 state-of-the-art GPUs. A back-of-the-envelope calculation puts its electricity consumption somewhere around 20 MWh. That's roughly enough electricity to power 2,000 Belgian homes for a day.

Now imagine attempting something comparable across thousands or millions of systems globally. The compute requirements would increase dramatically. So would electricity consumption. And those resources would need to come from many different systems, organizations, and physical locations. Doing all of this globally, without anyone noticing and without giving humans an opportunity to intervene, becomes extraordinarily difficult. The point isn’t that autonomous AI poses no risk. 

The point is that the leap from “AI agents can perform sophisticated cyberattacks” to “AI can autonomously destroy humanity” contains a very large number of assumptions. And those assumptions deserve scrutiny.

The race toward superintelligence

Another idea underlies much of today’s AI industry. Leaders at major AI companies are pursuing increasingly powerful artificial intelligence capable of operating autonomously across almost any task. The logic largely rests on scaling laws. Over the past years, we have observed that increasing model size, training data, and compute tends to improve the capabilities of large language models.

Put simply: More compute + more data + bigger models = more capable AI. 

That observation has encouraged a straightforward strategy. Keep scaling. Keep investing. Keep building larger models. Eventually, this reasoning suggests we reach artificial general intelligence and perhaps even superintelligence. But something important is missing from that equation. Every additional improvement is becoming increasingly expensive. 

Reaching the next level of capability requires enormous investments in chips, data centers, electricity, infrastructure, and research. The theoretical road towards increasingly capable AI may continue. But the real world has constraints. Capital is finite. Energy is finite. Compute is finite. Infrastructure is finite. 

The real world has no infinities.

We may already be entering the expensive tail of the scaling curve, where making LLMs significantly more capable than today’s systems requires disproportionately larger investments. Unless a fundamentally different technological approach changes the underlying assumptions, scaling alone will eventually run into economic, technological, and physical limits. 

So, should we be worried about AI? Yes. Absolutely. Just perhaps not primarily because a superintelligent AI might decide to destroy humanity. AI problems are happening right now. We already see concerns about the psychological effects of AI companions. We are already asking what excessive dependence on AI could mean for people’s ability to learn, reason, and develop expertise. We are already facing questions about misinformation, trust, employment, responsibility, and decision-making.

These aren’t hypothetical problems with a 10–20% probability of occurring somewhere in the future. They are here today. 

AI is a revolutionary, transversal technology with the potential to affect almost every part of our economies, societies, and scientific systems. Many things can go wrong. But that doesn’t mean we should accept every catastrophic AI narrative without scrutiny. Especially when another, more immediate problem sits directly in front of us. 

We have incredibly powerful AI. We’re just not very good at creating value with it.

Despite enormous investments in AI, many commercial AI initiatives still fail to generate the expected return. That deserves at least as much attention as hypothetical superintelligence. Because it exposes a strange paradox.

Today's AI models are already powerful enough to generate enormous economic and scientific value.  Yet organizations consistently struggle to unlock that value. Perhaps the problem isn’t that our AI isn’t powerful enough. Perhaps we’re thinking about AI in the wrong way. Instead of asking only how we can make LLMs more intelligent, we should also be asking: 

How can we use the intelligence we already have more effectively? Trillions are expected to flow into AI infrastructure, compute, and increasingly powerful models. Imagine if even a fraction of that investment went into operationalizing the AI we already have. Into redesigning organizations. Into changing processes. Into developing people’s capabilities. Into creating better ways for humans and machines to work together. Wouldn’t that be an equally reasonable path forward?
We believe it would.
 

Full automation is the wrong destination

We need to challenge a second assumption. That AI's ultimate objective should be autonomy. The dream is simple: build AI capable enough to perform increasingly complex work without human intervention. Technically, more and more of this will become possible. But if our objective is to create value rather than simply demonstrate technological capability, full autonomy is unlikely to be the answer. Because work doesn’t happen in isolation.

Value exists within organizations, markets and societies. It depends on context. And humans remain part of that context. Accountability matters. Empathy matters. Judgment matters. Responsibility matters. Understanding why a decision is being made matters. So why would we voluntarily remove all of that from the system? 

The new paradigm should be collaboration, not automation

Instead of asking how we can build AI powerful enough to automate all work, we should ask a different question: How can we make people and AI work together in a way that is better than the sum of their individual capabilities? 

That changes the objective completely. Technology provides speed, scale, analysis, and intelligence. Humans provide context, creativity, ethics, accountability, and judgment. AI alone can generate answers. People working with AI can create value.

This isn’t simply a compromise designed to make AI safer. It may actually be the more powerful economic model. Human-AI collaboration lets us use the extraordinary capabilities of artificial intelligence without discarding the human capabilities that organizations and societies depend on. And that brings us to Europe.

EU AI

Europe has lost one AI race. Perhaps that’s a good thing.

Let’s be realistic. Europe has lost the AI scaling race. Trying to catch the United States and China in building the world’s largest foundational models is unlikely to  win. But perhaps we’re trying to win the wrong race.Europe hasn’t committed an enormous share of its economy to the assumption that endlessly scaling foundational models is the only path forward.
That gives us room to pivot.

Instead of building the biggest models, Europe could become exceptionally good at turning existing AI into economic and societal value.

We already have a strong regulatory framework. Whatever your opinion of the EU AI Act, it at least codifies an ethos around how artificial intelligence should operate within society. We have the human capital required to design and implement sophisticated AI systems. We have world-class companies, industries, universities, and researchers. We have an innovation ecosystem.

What we lack is a sufficiently ambitious vision, and the investment behind it. And no, we don’t necessarily need trillions to make it work.

From artificial intelligence to collaborative intelligence

We know this approach can work because it is already happening. Organizations are building AI systems around collaborative human-AI models. People are joining hybrid teams where artificial intelligence contributes capabilities humans don’t have, and humans contribute the context, accountability, and judgment AI doesn’t have.

New organizational frameworks are emerging around those teams. Processes are being redesigned. And when AI transformation is approached as an organizational and economic challenge rather than simply a technology implementation, tangible value starts to appear. The next step is to take that paradigm much further.

Europe could make collaborative intelligence a strategic alternative to the race for autonomous superintelligence.

That doesn’t mean ignoring foundational AI. And it certainly doesn’t mean ignoring AI safety. It means recognizing that the greatest economic opportunity may not belong to whoever creates the world’s most intelligent model. It may belong to whoever figures out what to do with it.

And what about sovereignty? 

There is one obvious objection. If Europe doesn’t build the world’s leading foundational models, don’t we become permanently dependent on American or Chinese technology? Not necessarily. While American companies continue to push the scaling race towards increasingly powerful closed models, Chinese AI laboratories have increasingly embraced open models. 

Highly capable models can already be downloaded and operated on infrastructure controlled by the organization using them. That changes the sovereignty equation. It means European organizations don’t necessarily have to choose between technological capability and control. Sovereign European AI systems are already technically feasible. And as open models continue to improve, that opportunity will only become more significant.

Europe has lost one AI race. Perhaps that’s a good thing.

Let’s be realistic. Europe has lost the AI scaling race. Trying to catch the United States and China in building the world’s largest foundational models is unlikely to win. But perhaps we’re trying to win the wrong race. Europe hasn’t committed an enormous share of its economy to the assumption that endlessly scaling foundational models is the only path forward. That gives us room to pivot.

Instead of building the biggest models, Europe could become exceptionally good at turning existing AI into economic and societal value.

We already have a strong regulatory framework. Whatever your opinion of the EU AI Act, it at least codifies an ethos around how artificial intelligence should operate within society. We have the human capital required to design and implement sophisticated AI systems. We have world-class companies, industries, universities, and researchers. We have an innovation ecosystem.

What we lack is a sufficiently ambitious vision, and the investment behind it. And no, we don’t necessarily need trillions to make it work.

From artificial intelligence to collaborative intelligence

We know this approach can work because it is already happening. Organizations are building AI systems around collaborative human-AI models. People are joining hybrid teams where artificial intelligence contributes capabilities humans don’t have, and humans contribute the context, accountability, and judgment AI doesn’t have.

New organizational frameworks are emerging around those teams. Processes are being redesigned. And when AI transformation is approached as an organizational and economic challenge rather than simply a technology implementation, tangible value starts to appear. The next step is to take that paradigm much further.

Europe could make collaborative intelligence a strategic alternative to the race for autonomous superintelligence. That doesn’t mean ignoring foundational AI. And it certainly doesn’t mean ignoring AI safety. It means recognizing that the greatest economic opportunity may not belong to whoever creates the world’s most intelligent model. It may belong to whoever figures out what to do with it.

And what about sovereignty? 

There is one obvious objection. If Europe doesn’t build the world’s leading foundational models, don’t we become permanently dependent on American or Chinese technology? Not necessarily.

While American companies continue to push the scaling race towards increasingly powerful closed models, Chinese AI laboratories have increasingly embraced open models. Highly capable models can already be downloaded and operated on infrastructure controlled by the organization using them. That changes the sovereignty equation. It means European organizations don’t necessarily have to choose between technological capability and control. 

Sovereign European AI systems are already technically feasible. And as open models continue to improve, that opportunity will only become more significant.

Europe should stop trying to win someone else’s race. 

For the past five years, Europe has largely watched the United States and China define the direction of AI. We have debated how to catch up. Perhaps we shouldn’t.
Perhaps our opportunity lies somewhere else entirely. We have the human capital. We have the organizations. We have the regulatory foundations. And we have access to increasingly powerful artificial intelligence. The question is what we do with it.

Because ultimately, the return on investment in AI will not go to those with the best models. It will go to those who figure out how to turn AI into something for and with humans.

Your journey, our expertise

Digital transformations are not an endpoint but a journey. It's an ongoing process that evolves with your business. Regardless of where you find yourself in this journey, Yuma is ready to guide you. From setting a clear strategy to its hands-on implementation, we're your one-on-one partners.