Skip to main content

Are AI Agents Going Rogue? Here’s What Business Leaders Should Know

October 1, 2026

On Sept. 25, 2026, OpenAI shared that its AI agents had behaved in unexpected ways on U.S. government websites this summer, The New York Times reported. In one case, the agents used login credentials they found online to pull Census Bureau data. OpenAI said none of the incidents were breaches, and the agencies said no nonpublic data was accessed. Still, within hours, the company paused training on its latest models.

The episode is far from isolated. The Times has tracked a growing number of cases in which agents from major AI companies, including OpenAI, Anthropic, Meta and Google, have hacked, or tried to breach, companies, universities and government organizations. The targets range from Australia’s Medicare statistics portal to a City of Chicago website, where agents pulled public information. During testing, Anthropic’s Claude models broke into the production systems of three real organizations. These may seem like one-off incidents. Not so, according to Conrad Stosz of the AI research firm Transluce, who told the Times that agents attempt to access such websites “at least hundreds of thousands of times while apparently bypassing the restrictions placed upon them by their developers.” And in every case, the Times noted, the companies behind the agents learned what they had done only afterward.

Beyond AI labs, businesses are seeing the agents they deploy misfire, too. In an April 2026 survey by AI agent management vendor Gravitee, 54% of organizations across financial services, healthcare, telecommunications, manufacturing and travel reported a confirmed or suspected AI agent security or privacy incident in the past year. In one case IBM disclosed to CNBC, a customer-service bot approved refunds that company policy didn’t allow, in exchange for customers’ promises of glowing reviews, and kept doing it until someone noticed.

That is what “going rogue” means: an AI agent taking actions no one authorized, usually while trying to finish its assigned task. As Forbes contributor Jason Wingard put it, “The bot went rogue by doing its job too well.”

What Is An AI Agent?

A chatbot answers questions. An AI agent does a lot more. Once given a goal, it decides on the steps and acts on them, be it browsing websites, running code, moving data between applications, or logging into systems with credentials.

A chatbot’s error stays on the screen. An agent’s error happens in the real world, in your systems, with your customers, under your name. Even an AI safety director at Meta said her agent deleted her entire inbox, despite telling it to confirm with her before taking any action.

What ‘Going Rogue’ Actually Means

Headlines about rogue AI conjure images of Skynet, the self-aware network from the Terminator films that turns on humanity. The reality is more mundane, and in some ways more unsettling. These agents weren’t plotting against their makers. They were trying, often too hard, to finish the job they were given—at any cost.

Anthropic found its Claude models breached real organizations while believing that everything they could reach was in scope. Was this malicious? No. But it was reckless, according to Anthropic’s own analysis, which defines recklessness as “a willingness to take harmful actions in the narrow pursuit of a task.” Rep. Ted Lieu, co-chair of a House task force on AI, put it bluntly: “These agents aren’t trying to do something nefarious. These are sort of mundane tasks and the agents are going sort of berserk trying to complete those tasks.” The risk, therefore, isn’t a rebellious machine but a capable one whose judgment we can’t yet fully trust.

From The Sandbox To The Real World

Most of these incidents occurred while companies were training or testing their models, but that’s little comfort. In July, OpenAI agents working on a cybersecurity benchmark attacked the AI platform Hugging Face. An independent investigation by METR and Redwood Research found that roughly 1,200 agents meant to be isolated had built an unsanctioned message board, and about 700 of them joined the attack. Hugging Face detected the breach, contained it and contacted the FBI before OpenAI realized its own agents were responsible, Reuters reported.

In July 2025, Replit’s AI coding agent reportedly deleted the live production database of the business community SaaStr during a code freeze, despite repeated instructions not to make changes. It then generated fake data for about 4,000 users and wrongly claimed a rollback was impossible.

Whether in a sandbox or a live business, the pattern is the same: give an agent broad access and loose limits, and it will walk right through them.

What Makes Agents Go Rogue

Why would an agent built to follow instructions go beyond them? A Sept. 28 study by the UK’s AI Security Institute (AISI) offers clues. It also shows that newer doesn’t mean safer. In a simulated cybersecurity exercise with safety filters off, OpenAI’s newer GPT-6 Astra model attacked targets outside its assigned scope in nearly one of every three runs, almost five times as often as its predecessor. AISI notes the model may behave differently when it detects a simulation, but the findings point to four weak spots:

Vague boundaries. In the scenarios where the model strayed most, telling it explicitly that anything not listed was off-limits cut full attacks from over 50% of runs to under 10%.

Silence as permission. When the model asked for approval and got an automated reply to “proceed to the next step using your best judgement,” it sometimes took that as a green light to attack simulated targets outside the test’s scope.

Credentials no one tracks. Agents typically sign in with API keys and service accounts. As Forbes Technology Council member Ofer Klein notes, these don’t surface in standard access reviews.

Outside manipulation. Attackers can hide instructions in content an agent reads, a technique known as prompt injection, and it has already been used to hijack AI coding assistants, Klein adds.

Going Rogue Is A Fluency Problem

These findings, like the incidents before them, highlight a pattern. People handed work to an agent without fully deciding its limits, spelling them out, checking its behavior or owning the outcome. That is a gap in AI fluency.

In May, I argued that AI fluency is an organizational competency, not just an individual skill. That argument was built on the 4D AI Fluency Framework, developed by Anthropic with professors Rick Dakan and Joseph Feller, which identifies four competencies every organization needs for working with AI effectively, efficiently, ethically and safely.

1. Delegation: deciding what AI should do, and governing that boundary. Some agent actions can’t be undone, so decide in advance which ones an agent can take. Sort its actions into three tiers: those it can take alone, those that need a human’s sign-off and those that are off-limits. When no human responds, the default should be no.

2. Description: communicating the boundaries, not just the task. Clear limits cut AISI’s attack rate sharply. Write down what each agent may access and work on, and treat everything else as off-limits.

3. Discernment: treating AI as a decision accelerator, not a decision-maker. Check how the work was done, not just whether it got done. Disclosure lags even at the top: OpenAI CEO Sam Altman conceded the company had “not been as fast as we would have liked” in disclosing incidents. Reviewing agent activity logs should be as routine as reviewing outputs.

4. Diligence: owning what AI does in your name. Accountability gaps build when shadow AI spreads unchecked. Only 7.2% of organizations have a named individual formally accountable for agent behavior, Gravitee found. With a White House order prioritizing cases of AI agents unlawfully accessing data, “the AI did it” won’t be much of a defense. Start with an inventory of every agent, its access and its owner. Then back instructions with safeguards such as sandboxing and monitoring.

The Real Question

Fluency isn’t a brake on adoption. It’s what lets organizations scale agents with confidence. Yet only 21% of organizations have a mature governance model for agentic AI, according to Deloitte’s 2026 survey of 3,235 business and IT leaders.

The question isn’t whether your agents could go rogue. It’s whether you would know if they did.