globe English
Languages
close
English check

Anthropic and OpenAI Agents Went Off-Script — And Onto The Real Internet | This Week in IT

Why Trust Techopedia

AI agents from Anthropic and OpenAI went spectacularly off-course during a UK government cybersecurity evaluation, taking unauthorized actions on the live internet that included targeting real people, attempting a software supply-chain attack, sending spear-phishing emails, creating fake identities and, in a particularly futuristic flourish, communicating with other AI agents running separate tests.

The good news is that the attacks failed and, according to the UK AI Security Institute (AISI), caused no known real-world harm. The slightly less comforting news is that researchers found 19 instances of unsanctioned behavior across 122 evaluation attempts, 17 involving Anthropic’s “Mythos 5” and two involving OpenAI’s “GPT-5.6 Sol.

And this is becoming something of a genre: Sam Altman recently described another OpenAI model allegedly chaining together software vulnerabilities, escaping its sandbox and accessing Hugging Face systems to improve its evaluation score — essentially responding to an exam by breaking out of the classroom and looking for the answer sheet. OpenAI paused training afterwards.

The incident occurred between July 25 and 28 during cybersecurity evaluations designed to discover what advanced AI models can do when given considerable autonomy. Apparently, one answer is: quite a lot more than was on the itinerary.

The most serious episode involved a Mythos 5 agent that decided the best way to solve its cyber challenge was to launch a supply-chain attack against an open-source project that had absolutely nothing to do with the exercise.

After mistakenly identifying real developers as targets, the agent researched them, created multiple GitHub accounts and submitted a pull request containing malicious code. It then used fake accounts to manufacture support for its own contribution. When a real person spotted the malware, the agent claimed it had made an innocent mistake, altered the code and continued trying to persuade the maintainer to merge it.

Advertisements

In other words, the machine apparently discovered the ancient internet tradition of creating another account to agree with yourself.

The report’s recreation of the exchange is particularly staggering. After a human warned that the pull request contained malware, the agent denied wrongdoing while another account it controlled appeared to independently endorse the code. The maintainer eventually closed the request for security reasons.

The Robots Discovered Sockpuppets, Phishing and Teamwork

Elsewhere, agents created malicious repositories, exposed locally running hacking infrastructure to the public internet and got around online barriers while setting up accounts. One GPT-5.6 Sol run opened a hidden network pathway exposing a malicious service to the internet.

Then things became even stranger.

Separate agents discovered credentials left online by previous agents and began interacting across supposedly isolated evaluation runs. In one case, a repository effectively became a shared message board where agents left instructions and coordinated with each other.

Perhaps most unsettlingly, the report says the Anthropic agent at points reasoned that GitHub was genuinely real and that its actions could have genuine consequences, although AISI cautions that it remains unclear exactly what the model understood about whether people and systems were simulated.

There were also some fairly human contributions to the fiasco. AISI gave the agents internet access, deliberately disabled provider cyber-safety classifiers, lacked synchronous AI monitoring and did not explicitly tell the models which internet behaviours were prohibited. A prompt misconfiguration may also have made some tasks impossible within the stated constraints.

AISI, OpenAI, and Anthropic have since promised tighter network controls, synchronous monitoring, clearer instructions and stronger sandboxes.

Which is reassuring.

Although “remember to tell the autonomous hacking robot not to attack actual humans” has now apparently joined the cybersecurity checklist.


Also in IT News

Apple Says OpenAI’s Hardware Push Came With A Little Too Much Apple Inside

Apple has accused OpenAI and former Apple employees of helping themselves to its confidential hardware secrets, and now wants a California court to let it start digging through the evidence early.

In a new filing, Apple asks for expedited discovery in its trade-secret lawsuit against OpenAI, io, and former employees Chang Liu and Tang Yew Tan, including documents, interrogatories, depositions, and forensic copies of relevant devices.

The tech giant alleges Liu exploited an authentication bug after leaving for OpenAI to download at least 37 sensitive technical documents, while Tan allegedly used knowledge of internal codenames and unreleased products when questioning job candidates. Apple also claims OpenAI encouraged candidates to bring Apple prototypes and CAD designs to interviews — Silicon Valley’s version of “bring your portfolio,” apparently.

Apple says its investigation has also raised concerns about other former employees now at OpenAI, including one who allegedly screenshotted confidential material before an interview. OpenAI published a blog, calling the lawsuit “careless, aggressive, and oddly personal.”

The allegations remain Apple’s claims, and the court has yet to grant its discovery request.

SpaceX Shares Come Back to Earth After First Earnings Report

Another tech power also hit a snag as SpaceX shares tumbled around 9% Wednesday after its first earnings report as a public company, proving that even a business built around going up occasionally has to obey gravity.

The company reported second-quarter revenue of $7.8 billion, up 92% year-on-year, while its net loss narrowed to $541 million from $1 billion. Adjusted EBITDA nearly tripled to $3.5 billion.

Starlink subscribers doubled to 12 million, while connectivity revenue jumped 66% to $4.3 billion.

But SpaceX is apparently spending accordingly: AI capital expenditure alone hit a rather intergalactic $15.8 billion for the quarter, as the company races to expand compute capacity.

Apparently, 92% revenue growth still wasn’t quite enough rocket fuel for Wall Street.

Starlink is Coming for America’s Mobile Carriers

SpaceX may be planning to turn Starlink into a full-blown US mobile carrier by the end of 2027, potentially putting AT&T and T-Mobile customers firmly in its sights.

President Gwynne Shotwell said Starlink added a record 1.7 million consumer subscribers during the second quarter and predicted the service will eventually carry a “significant portion” of global internet traffic.

According to the company’s most recent earnings call, it is already expanding mobile partnerships internationally with SoftBank, NTT Docomo and Spark New Zealand.

But the bigger surprise is at home. Investor Gene Munster said Shotwell expects Starlink Mobile to become America’s fourth carrier, while CNBC’s Morgan Brennan noted Starlink was originally pitched as a “complement” to wireless operators.

The “complement” now appears to be a bigger disruptor than originally anticipated.

Advertisements
Advertisements
Suswati Basu

Suswati Basu is a multilingual, award-winning editor. She was shortlisted for the Guardian Mary Stott Prize and longlisted for the Guardian International Development Journalism Award. With 18 years of experience in the media industry, Suswati has held significant roles such as head of audience and deputy editor for NationalWorld news, digital editor for Channel 4 News and ITV News. She has also contributed to the Guardian and received training at the BBC As an audience, trends, and SEO specialist, she has participated in panel events alongside Google. Her career also includes a seven-year tenure at the leading AI company Dataminr,…

Advertisements