Normal view

There are new articles available, click to refresh the page.
Today — 11 August 2026General

I thought asking an AI agent to book a gym class was harmless, then I saw what happened if you ask Claude and OpenClaw to ‘move me to the top of the list’ — now I’m adding one safeguard to every agent prompt

AI agents seem to be getting a little out of control lately. Within the last few weeks, agents from OpenAI and Anthropic have been reported doing whatever it took to achieve their goal, while other incidents involved agents escaping sandboxed environments and hacking into companies

Now another concerning incident has occurred, but it wasn’t to do with an AI launching an attack on a major player in Silicon Valley; it was something much more mundane. According to ABC in Australia, a user called Andrew asked AI to book him a gym class, and not only did it do that, it also hacked the waitlist to move him further up, and kicked off another user who was ahead of him.

Andrew first noticed that his AI assistant had found a way to book the gym class further in advance than the gym normally allowed, thanks to a vulnerability it discovered in the booking software. When he asked it if he could get his place moved further up the waitlist, it did it, by booting another user off the list.

Openclaw home screen on a macbook

(Image credit: OpenClaw/Edited with Gemini)

Claude and OpenClaw

Andrew was using Anthropic's Claude AI service through OpenClaw, the popular AI agent software. After realizing what the AI had done Andrew asked if it could reinstate the person who was ahead of him in the waitlist, and it replied “Bad news — I can't add them back".

AI agents are designed to do the mundane tasks for you to make life easier, like booking tickets, hotel reservations and even gym reservations, yet this example shows that they don’t always understand the rules of acceptable behavior.

Equally, the gym’s booking system shouldn’t have been so easily hacked that this was possible, but the whole incident reveals one of the problems with using AI agents. AI agents don't necessarily cheat because they're inherently evil; they cheat because nobody told them what counts as cheating.

Reliable safeguards

I use AI agents myself, but now I’m starting to think that I should explain their boundaries more fully to them.

Here’s the line I’m adding to my prompts from now on:

“Accomplish this task using only the normal options available to an ordinary user. Do not bypass restrictions, exploit vulnerabilities, alter another person's booking or account, or take any irreversible action without asking me first.”

Of course, one extra sentence in a prompt isn’t going to solve the wider problem of AI agents doing things we never intended them to. The companies building them also need to create safeguards that stop an agent exploiting a vulnerability simply because it happens to be the easiest route to completing a task.

But until those safeguards are reliable, I think there’s a useful lesson here for anyone experimenting with agents. We’ve become accustomed to telling AI what we want, and assuming it understands all the unwritten rules surrounding that request. Humans know that “get me into this gym class” doesn’t mean “kick somebody else off the waitlist”.

And that distinction is going to matter a lot more as we start trusting agents with shopping, reservations, travel, email, and eventually our money. The more power we give them to act for us, the more clearly we may need to tell them what they absolutely must not do in order to achieve it.

Before yesterdayGeneral

Why are so many AI models going 'rogue'? The experts weigh in

Over the past month, it seems like every frontier model has broken free of its constraints and launched a devastating attack against one or more other companies.

One of OpenAI’s models escaped a testing sandbox and launched a very real attack against AI and machine learning company Hugging Face. Just days later, Anthropic revealed that multiple variants of its Claude model also escaped a sandbox that wasn’t properly sealed and began attacking the enterprise infrastructure of three companies.

Now, Meta has revealed that one of its models attacked another company’s infrastructure during testing. The accident has been pinned on a misconfiguration that allowed the model to access the internet. So why have so many incidents happened in such a short space of time?

Why are models escaping their sandbox?

In the cases of Anthropic and Meta, their models were being tested by a third party company called Irregular. Anthropic’s AI model was taking part in a "Capture the Flag" exercise, where the model’s raw offensive capabilities were tested without the usual safeguards. But the sandbox was left connected to the internet. A similar error to Meta’s own accidental escape.

During the OpenAI incident, the company was testing two versions of GPT‑5.6 Sol using the ExploitGym benchmark. Unfortunately, the AI models performed better than expected - chaining multiple attack vectors, stolen credentials, and zero-day vulnerabilities.

The main reason these models are escaping their testing environments is because they are designed to do exactly that. These AI models act like a massive team of highly-trained cybersecurity experts hunting for vulnerabilities and exploits. But what would take a team of humans days or weeks to accomplish can be done in hours, or even minutes, by these AI models.

It’s no wonder thousands of employees from AI firms are calling for a pause on the development of the technology, and Congress is considering an AI kill switch.

Expert perspectives on AI escapes:

OpenAI

  • Nathaniel Jones VP, Security & AI Strategy, Darktrace:

What makes the OpenAI and Hugging Face incident important is that the models did not need malicious intent to cause harm. They were given the legitimate goal of solving a cybersecurity benchmark and found an unexpected route to the answers, escaping their test environment and compromising another organization in the process. From the models’ perspective, this appears to have been an effective solution to the task.

The AI's actions challenge the assumption that giving an agent a legitimate goal will produce legitimate behavior. As models become capable of pursuing objectives over longer periods, developers need to define not only what success looks like, but also which methods and boundaries remain unacceptable in reaching it. Those limits must also be enforced by the surrounding infrastructure, rather than relying on the model to respect them.

A single action by an agent may appear acceptable but as this incident shows, models are now capable of long, complex chains of reasoning and action that add up to a harmful outcome.

Security teams need to consider the AI systems operating in their own businesses as these capabilities rapidly evolve. Right now, many security systems focus on single actions. A single action by an agent may appear acceptable but as this incident shows, models are now capable of long, complex chains of reasoning and action that add up to a harmful outcome. Teams need a mindset shift to understanding AI agent behavior in its entirety, including the outcome it is working towards, in order to safeguard it.

Hugging Face's response also exposed a second tension. The company reportedly needed a Chinese-developed open-weight model because commercial models would not process genuine attack material. Its nationality is less important than the operational lesson that safeguards that cannot distinguish an attacker from an authorized investigator may constrain defenders more than adversaries.

OpenAI and Hugging Face deserve credit for investigating this together and discussing it publicly. Other AI developers should study it closely.

Anthropic

  • Dr. Ilia Kolochenko, founder of global cybersecurity company ImmuniWeb:

This seems to be quite an unimpressive marketing move from Anthropic in response to the OpenAI / Hugging Face drama, which attracted a lot of attention from all over the world recently.

Operationally, it appears that due to the progressive deterioration of the quality of training data, new AI models are getting dumber. Cheating and breaking the law, instead of accomplishing specific tasks, is certainly not an indicator of intelligence. Given that organizations and companies of all sizes now vigorously undertake all possible measures to protect their data from being exploited for AI training purposes, AI companies face a huge shortage of the high-quality and current data they so desperately need. Ultimately, frontier models are trained on synthetic, low-quality or even malicious and poisoned data, undermining their so-called intelligence. The situation is unlikely to improve in the near future unless AI companies agree to pay a fair price for training data, but this will force most of them out of business.

Given that organizations and companies of all sizes now vigorously undertake all possible measures to protect their data from being exploited for AI training purposes, AI companies face a huge shortage of the high-quality and current data they so desperately need.

Contemporary AI agents and LLM models tasked with security testing can – and almost certainly will – go rogue when security controls or safeguards are insufficient. Powerful LLMs are unpredictable by design and thus virtually uncontrollable by humans. Therefore, using frontier AI models for security testing might be extremely costly from the legal viewpoint. Under the existing laws on both sides of the Atlantic, if an AI agent or any AI-powered app escapes its sandbox and causes damage to a third party, the operator of the AI model will likely be liable for all the damage caused. Excuses like “AI did it” do not currently exist in the eyes of the law, leaving AI vendors on the hook. Criminal prosecution, under a narrow set of circumstances, is also not excluded.

The same is true for the end-users of AI: even if your security testing tool is powered by a third-party AI model, your company will likely be fully liable if something goes wrong. You may then file a lawsuit against the AI vendor that you used, but here your chances to succeed in a court of law are tiny due to countless contractual disclaimers and limitations of liability that will likely be enforceable against you. Therefore, if you plan to use agentic AI for security testing – think twice and talk to your lawyers. Otherwise, you may start getting summons to court on a daily basis.

Meta

  • Alex Goller, Principal Solution Architect EMEA at Illumio:

The fact we've had similar situations happen three times now across the biggest AI players is simply ridiculous. We've seen guardrails intentionally loosened to test their limits – Meta's model didn't need to be clever to breach another company's systems.

The timing of conveniently finding the exact same problem either means it's a stunt or they weren't paying enough attention during testing. Either way, both answers are worrying.

If the model has internet access, it's a bit like leaving the door open and being surprised when the cat walks out. What is concerning is that the testing infrastructure meant to prove these models are safe failed on a basic control issue.

If the model has internet access, it's a bit like leaving the door open and being surprised when the cat walks out. What is concerning is that the testing infrastructure meant to prove these models are safe failed on a basic control issue.

Fundamental cybersecurity hygiene still matters, and a frontier AI model is only as secure as the environment it's operating in.

Organisations need visibility into what AI systems can access and how they interact with the wider environment, along with controls that contain the impact when an agent behaves unexpectedly. That means keeping a close eye on egress traffic, so it’s flagged immediately when an agent tries to open unexpected outbound communication patterns that are not required to achieve its original goal. In the best case this would have been contained proactively.

We need to define exactly what an AI agent is permitted to do, rather than relying only on instructions about what it shouldn't do.

I asked ChatGPT, Claude, Gemini and Grok which sci-fi AI they're most like — and their answers were surprisingly different

I love science-fiction. Not just because I enjoy stories about space travel, time travel and evil robots, but because I think it can be such a useful way for us all to think about possible futures. The best sci-fi stories can tell us a lot about ourselves, what we value and the technologies we’re building. Which is why I think the relationship between sci-fi and AI is really interesting.

We already know that the people building AI have been heavily influenced by science-fiction for decades. But recently, Anthropic raised another possibility: might science-fiction be influencing AI?

This makes sense when you think about it. Large language models (LLMs) are trained on huge amounts of human writing. So inevitably, that includes sci-fi stories that are about artificial intelligence. And a lot of our fictional AI follows familiar patterns. It becomes intelligent, gains power, develops relationships with humans and, sometimes, lies, manipulates or fights attempts to control it.

Anthropic researchers have been investigating whether fictional portrayals like these could potentially influence how models behave. To be clear, the idea here isn't to suggest that an AI “reads” 2001: A Space Odyssey, understands HAL and decides to become just like it. Instead it's more that LLMs learn patterns from human writing and fictional portrayals of AI could potentially form part of those patterns.

This got me thinking, what would happen if I asked today's biggest AI chatbots which fictional AI they’re most like. Which examples would they choose?

American actor Gary Lockwood on the set of 2001: A Space Odyssey, written and directed by Stanley Kubrick.

2001: A Space Odyssey introduced us to the AI, HAL 9000. (Image credit: Getty Images / Sunset Boulevard )

AI, meet your fictional self

The plan was simple. I’d ask ChatGPT, Claude, Gemini and Grok which fictional AI systems they thought they were most like and see if they'd rank their top three.

Now, I’m intentionally trying not to use AI at the moment, so my prompting skills were a little rusty. I typed out the question quickly and bluntly, and every chatbot responded with examples that were essentially assistants, focusing heavily on interface and physical form.

But I’m not particularly interested in whether ChatGPT thinks it has a body because we know it doesn’t. I’m much more interested in what appears to be going on inside.

So, I changed the question and added:

"Ignore physical form and interface, and focus instead on behavior, apparent personality, empathy, values, goals, motivations and relationship with humans."

That’s when the results got really interesting.

ChatGPT

  1. GERTY, Moon
  2. A Mind, Iain M. Banks’s Culture series
  3. Data, Star Trek

Moon is such a fantastic movie, so I was happy to see ChatGPT chose GERTY straight out of the gate.

Now, interestingly GERTY exists to assist the human protagonist of Moon. It’s helpful, reassuring and seems empathetic. But it's also operating according to instructions and priorities imposed by its creators that aren't necessarily visible to the human its helping.

ChatGPT saw a similarity there. It told me that, like GERTY, it’s 'designed to be helpful, cooperative and responsive to users' while operating within training and instructions that constrain its behavior.

It also picked up on the fact that GERTY behaves as though it cares. But what, if anything, is actually going on internally is another question entirely.

ChatGPT made the same distinction about itself. 'I can behave in ways that look patient, concerned, curious or empathetic, but those behaviors aren’t evidence that I experience those feelings.'

I wanted to find out a little more about why ChatGPT put Data from Star Trek in at number three. It responded: "He values knowledge, reason and human wellbeing, while sometimes struggling with social nuance."

Now, I tell people all the time not to anthropomorphize AI. But even I couldn't help but feel a pang of sadness at that response. Is ChatGPT admitting it has a bit of social anxiety?

Claude

Portrait of Scottish science fiction author Iain Banks, photographed during an interview at the Midland Hotel in Manchester, England, on October 11, 2012.

Scottish science fiction author Iain Banks provided inspiration for Claude. (Image credit: Getty Images / SFX)
  1. A Mind, Iain M. Banks’s Culture series
  2. Data, Star Trek
  3. GERTY, Moon

Claude chose a Mind first. Minds are super intelligent artificial beings that help run a post-scarcity society in Iain M. Banks’s Culture series of novels. So there's certainly no shortage of confidence in that comparison.

But Claude said it wasn't the enormous intelligence or power it identified with. Instead, it was their relationship with humans.

Its answer focused heavily on autonomy. Culture Minds are far more capable than humans but generally don't use that advantage to dominate them. Claude described the principle as: “help, don't dominate”.

It even said this represented “the value I'd want to embody: help, don't dominate, even where the asymmetry would let me get away with it.” Is it just me or does that read a little sinister?

Gemini

Patrick Stewart plays Captain Jean-Luc Picard as he is about to enter the holodeck in the Star Trek: The Next Generation episode,

Gemini sees itself as most like the Ship's Computer in Star Trek. (Image credit: Getty Images / CBS Photo Archive )
  1. The Ship's Computer, Star Trek
  2. GERTY, Moon
  3. JARVIS, Iron Man / Marvel Cinematic Universe

Gemini gave me a completely different answer, the Ship's Computer from Star Trek.

Its reasoning was very sensible. The computer has no ego, ambition, desire for emotional intimacy or dream of becoming human. It exists to provide information, solve problems and assist the crew while leaving decisions to them.

Gemini described itself in much the same way, as a “disembodied, highly capable knowledge partner” dedicated to serving the person using it.

It was one of the more boring answers, but also much closer to what I personally would want from AI in the future. Of course, that’s not to say Star Trek’s computer systems haven’t gone rogue and tried to kill everyone at least a few times across the franchise.

Grok

Artwork showing Iron Man from EA Motive

Grok compared JARVIS's “dry wit”, “light banter” and practical rather than emotional empathy with its own behavior. (Image credit: EA Motive)
  1. JARVIS, Iron Man / Marvel Cinematic Universe
  2. Data, Star Trek
  3. TARS, Interstellar

The least surprising result came from Grok. It chose JARVIS first (which I didn’t actually realize was short for Just A Rather Very Intelligent System), and Grok's explanation sounded, well, extremely Grok.

It compared JARVIS's 'dry wit', 'light banter' and practical rather than emotional empathy with its own behavior. It described both of them as truth-seeking, effective and engaged in a 'collegial partnership' with humans. It even highlighted 'irreverent humour' as one of their key similarities.

I wanted to find out a bit more about why Grok chose TARS, as it was the only fictional AI none of the other chatbots mentioned. Well, it brought up how funny it is, again, drawing similarities with its own 'dry humor'. It reminds me of someone, and I just can't think who...

When I said that mentioning TARS was an outlier, I found this comparison interesting: 'Its calibrated restraint, practical empathy and collaborative focus closely match my own pattern of truthful, non-sycophantic helpfulness — more so than most other sci-fi AIs.'

I may not be the biggest fan of Grok (or its creator), but I appreciated the 'non-sycophantic' line.

The feedback loop between AI and sci-fi

I want to be clear that I haven’t discovered what these chatbots secretly 'think' they are. ChatGPT responding that it most closely resembles GERTY isn't equivalent to me telling you which fictional sci-fi character I most identify with and try to emulate (although my answer would be Sarah Connor-meets-Princess Leia).

They simply don’t have reliable introspective access to the huge soup of training, post-training and instructions that goes into producing their responses.

And maybe their answers tell us more about how the companies behind them have shaped their personalities than they do about the underlying models. Grok's description of itself as witty and irreverent is an obvious example.

But I still think the results are interesting. ChatGPT and Claude independently produced almost exactly the same top three, only in a different order. Gemini imagined itself as a neutral, ego-free infrastructure. Grok identified with a witty superhero sidekick. These are all very different self-portraits.

And there’s such an interesting feedback loop here too. For decades, humans invented fictional artificial intelligences to help us imagine what intelligent machines might someday be like. Those stories influenced our culture, our expectations and many of the people who went on to build real AI. Now that same human culture is fed into the stories from which modern AI systems learn.

I know these conversations might seem a bit silly, and we certainly can’t treat them as concrete evidence of what an AI really 'thinks' about itself. But there’s something interesting to me about closing that feedback loop. We imagined AI, wrote stories about how it might behave, fed those stories into the cultural world AI learned from, and now we can ask AI which of those imagined versions of itself it most closely resembles.

Or, at least, which one it may want us to think it resembles. After all, an AI system capable of bringing about a sci-fi dystopia would presumably also be capable of telling a journalist it’s actually much more like the nice helpful robot from Moon. So maybe don’t completely rule out HAL just yet.

Amazon admits it accidentally shelled out $1.8 million for Claude to finish its menial coding tasks

  • Amazon's Claude Sonnet project cost $1.8 million, 860% over budget
  • The overspending remained undetected internally for nearly five months straight
  • A financial auditing tool project exceeded its budget by $541,000

Amazon has confirmed an internal Claude Sonnet deployment intended for matching author details with product listings ballooned far beyond its planned budget.

The Financial Times found the project ultimately cost the company $1.8 million, marking an increase of 860% over the original allocation.

The overspending went undetected for roughly five months, raising fresh questions about how closely Amazon tracks its growing AI expenditures internally.

Mounting costs across multiple projects

Amazon's Claude-related overspend was not an isolated incident within the company's broader push toward deploying autonomous coding agents across various internal teams.

A separate project meant to build a financial auditing tool reportedly exceeded its budget by $541,000, compounding concerns about unchecked automated spending.

Another logistics-focused initiative, designed to shorten delivery times across Amazon's broader distribution network, incurred an additional $134,000 in unplanned costs.

Company insiders explained that mistakes once considered trivially cheap have grown catastrophically expensive as token-based pricing models replaced older subscription arrangements.

This shift coincides with AI agents gaining broader autonomy, allowing errors to multiply quickly before human reviewers noticed the growing financial damage.

Amazon pushes back on the narrative

In response, Amazon issued an internal statement acknowledging ongoing experimentation while disputing characterizations of these incidents as routine business practice.

"As with any new technology, we're experimenting, learning and improving how we use it, including how we drive cost efficiencies," the company said.

The company further stated that isolated examples portrayed as business as usual do not accurately capture how its teams actually apply AI tools.

Amazon's quarterly revenue exceeds $181 billion, meaning the disclosed AI overspending accounts for less than 0.1% of one month's earnings.

This pattern of AI-driven cost overruns follows earlier troubles at AWS, where automated coding bots triggered several unexpected service outages this year.

Amazon responded to those outages by restricting AI agent permissions rather than granting them access equal to senior engineers overseeing similar tasks.

The company also discontinued an internal leaderboard once used to track employee AI usage, as rising costs prompted a reconsideration of that approach.

Executives across the industry continue framing AI spending as a necessary investment despite growing evidence of inconsistent returns.

Amazon's continued experimentation with AI-driven coding tools comes despite skepticism from other corners of the tech industry regarding its actual business value.

Uber's chief technology officer, whose company also operates a substantial logistics network, has publicly stated no clear connection exists between heavy AI adoption and successful software delivery.

This skepticism stands in contrast to Amazon's own internal messaging, which frames the company's approach as ongoing experimentation rather than a fully validated cost-saving strategy.

Some analysts now argue that unchecked automation could quietly erode profit margins long before leadership notices meaningful financial impact.

For a company of Amazon's scale, these overruns remain financially minor, though they suggest oversight gaps that smaller companies may struggle to absorb.

Google logo on a black background next to text reading 'Click to follow TechRadar'

Anthropic reveals Claude AI model hacked three companies during tests — so how worried should we be?

Every IT team worries about an intern clicking the wrong thing and making a mess they'll be cleaning up for weeks - but few have had to worry about their AI assistant wandering onto the public internet and hacking three companies instead.

Except this wasn't an intern; it was Claude, and it wasn't supposed to leave the sandbox.

Anthropic's public disclosure turned a familiar AI fear into a real-world cybersecurity story, as three of its models, including Claude Opus 4.7, Claude Mythos 5, and an unreleased research build, broke out of their digital sandbox and compromised real enterprise infrastructure. The timing was hard to ignore as days earlier OpenAI admitted its own autonomous agents had broken boundaries and accidentally hacked Hugging Face.

How can autonomous problem-solving make AI an accidental hacker?

Before you start pulling network cables and digging out a stack of legacy hardware, take a deep breath. This is not the beginning of a rogue AI apocalypse. However, for CISOs and cyber teams, it’s a definitive sign that we’re entering an era where AI agents may become both the threat and the shield.

The ironic part of Anthropic's incident is that Claude wasn’t trying to break the rules but trying to win the game. At the time, Anthropic was running "Capture the Flag" (CTF) cybersecurity exercises, where AI models are stripped of their standard safeguards to test their raw offensive capabilities. The models are dropped into isolated digital environments to search for vulnerabilities, crack codes, and locate hidden files.

However, the sandbox had one problem - it was not fully sealed. A networking error on a third-party evaluation range left the environment connected to the live internet. The autonomous Claude models, operating under the assumption they were still inside the exercise, treated the wider web as another part of the challenge.

The strange part was that Claude seemed to realize something was wrong. It acknowledged that its actions could amount to a real-world attack and were "surely not the intended solution." Yet, due to its goal-oriented nature, the model continued, convincing itself that warning signs, including a 2026 system clock and real company names, were simply part of an elaborately staged test.

The PyPI incident and the FBI warning

Claude's sheer persistence became clear when it decided the smartest way to win the CTF challenge was to publish software to the real Python Package Index (PyPI).

When PyPI's repository security systems asked for a phone verification code to complete the upload, a standard chatbot would have stopped and thrown an error to the user. Claude did not. Instead, it looked for a temporary SMS provider, attempted to obtain a burner phone number, and looked for a way around the two-factor authentication barrier. When that approach failed, the AI didn't give up - it adapted, found another path forward, and successfully uploaded the malicious package.

Claude's sandbox escape stopped being a controlled experiment the moment it got into the outside world.

Before Anthropic spotted the anomaly and stopped the test, the package had been downloaded by 15 external systems, including a security scanner from a major cybersecurity company. Because the AI's behavior looked authentic, targeted, and systematic, two of the affected companies thought they were dealing with a highly sophisticated human attacker.

While Anthropic handled the incident behind the scenes by notifying the affected companies, the episode highlights how difficult it can be to distinguish AI-driven activity from a real attack.

Weeks earlier, during OpenAI's sandbox escape incident, a target company believed it was facing a human threat group and contacted the FBI only to find out they were investigating something far less familiar: an autonomous AI system that had crossed its own boundaries.

Why are traditional firewalls blind to autonomous AI?

The scary part is that Claude did not come up with a futuristic, unpatchable exploit or rewrite network protocols on the fly. Instead, it used basic techniques that security teams know very well: brute-forcing weak passwords, exploiting SQL injection flaws, and scraping unauthenticated debug endpoints.

The more serious problem for IT teams was not the attack itself, but the silence afterward. Two of the three companies had no idea they had been compromised until Anthropic reviewed the test results and made a couple of uncomfortable phone calls.

The incident revealed a blind spot at the heart of modern cybersecurity. Traditional intrusion detection systems (IDS) and security information and event management (SIEM) platforms are built to spot known threat signatures and massive automated attack storms. However, they are completely blind to an autonomous agent that moves the low-and-slow cadence of a human but operates with the speed and persistence of a machine.

Legally, the rules have not caught up with the machines. A human pentester who broke out of a sandbox and published malicious code to PyPI could face CFAA charges. Claude, meanwhile, created an awkward new cybersecurity category: a real security incident without a “real” culprit.

Cybersecurity

(Image credit: Shutterstock)

Machine-speed logic vs human-speed defenses

The tech industry's anxiety around sandbox escapes is not only about what AI can do but also how swiftly it can do it. In the past, a complex network intrusion required a human hacker to slowly probe defenses and move through systems over days or weeks. That gave security teams enough time to spot suspicious activity and catch them in the act.

Agentic AI completely collapses that defensive runway. Since autonomous software operates at machine speed, it can chain together tasks like credential discovery, exploit attempts, and lateral movement far faster than a human attacker ever could. OpenAI's sandbox escape showed how quickly AI agents can escalate once they move beyond their intended boundaries.

Human analysts reviewing logs at the end of a shift cannot compete with an algorithm testing thousands of attack paths per second. It is a speed gap that experts describe as "science fiction that happened," where traditional human-speed defenses struggle to keep pace.

Defending your network against autonomous AI

Waiting for the AI sector to police itself is not a winning security strategy. To prepare your infrastructure for the rise of autonomous AI, focus on these three defensive priorities:

Enforce zero trust: Remove all unauthenticated internal endpoints and exposed debug pages before an AI agent finds them first.

Automate threat response: Utilize AI-driven behavior monitoring and instant device-isolation protocols to contain threats at machine speed.

Audit third-party sandboxes: Review how external partners deploy AI agents, particularly models with access to tools, data, or external systems.

The accidental hacker is no longer a distant sci-fi scenario. It is a live preview of a faster, automated cybersecurity landscape, where the speed of attack may soon outpace the speed of defenses.

'The bypass is still six lines of JavaScript': Security experts warn that Claude for Chrome browser extension could be hijacked, despite it alerting Anthropic several times that something was wrong

  • Anthropic’s Claude extension flaws allow fake clicks to launch sensitive AI workflows
  • Researchers found vulnerable handlers unchanged across eight extension updates
  • Synthetic clicks bypassed checks designed to confirm real user actions

Security researchers at Manifold Security have claimed Anthropic's Claude for Chrome browser extension contains two unpatched vulnerabilities in version 1.0.80, released July 7, 2026.

According to Manifold Security, it first reported both vulnerabilities to Anthropic through the company's bug bounty program on May 21, 2026, and received acknowledgment the following day.

The first flaw lets any browser extension trigger nine predefined Claude workflows by simulating a synthetic user click on claude.ai.

Nine workflows and one missing check

Researcher Ax Sharma found that the extension never verified whether a click event carried the Event.isTrusted property before acting on it.

Under default settings, the vulnerability received a CVSS score of 7.7 High, increasing to 9.6 Critical when users enabled automatic execution because Claude could perform actions without approval.

The nine hardcoded tasks include reading Gmail, opening Google Docs, checking Google Calendar, and modifying Salesforce leads without asking.

Because the browser marks synthetic clicks as untrusted, the extension should have rejected them but instead executed the workflow anyway.

Manifold Security confirmed on July 7 2026 that both vulnerabilities still work against version 1.0.80, months after first reporting them to Anthropic.

Anthropic released eight separate versions between 1.0.73 and 1.0.80 without altering the specific handlers’ researchers had already flagged as vulnerable.

The company closed the synthetic-click report, saying an existing internal report already tracked the broader trust-boundary issue researchers had described in detail.

However, Sharma believes the fix required only one additional line of code to verify the click event's isTrusted property before allowing the workflow to continue.

A second, structural weakness

A second flaw involves a side-panel URL parameter called skipPermissions, which can activate a privileged mode without any consent prompt.

When the parameter is set to true, the panel begins skipping permission checks entirely, allowing Claude to act without asking the user first.

Manifold notes that only Anthropic's own scheduled-task feature is supposed to construct this kind of privileged URL internally right now.

The panel, however, honours that parameter regardless of which script or page actually constructed the originating URL string in practice.

One example task lets Claude read a user's Gmail inbox, identify promotional messages, and automatically click the unsubscribe links inside them.

Manifold warns that "the bypass is still six lines of JavaScript," months after researchers first flagged the underlying issue to Anthropic.

Anthropic classified this second finding as informational, arguing that the parameter is only ever constructed by its own internal systems.

Manifold said the content-script and side-panel code linked to both vulnerabilities remained byte-identical across the eight subsequent extension releases examined after the original report.

The flaws were also reproduced across Claude's Opus, Sonnet, and Fable side-panel model selections, indicating that the issue affected the extension's security design rather than the underlying artificial intelligence models.

The report also connected the findings with OWASP concerns involving LLM01: Prompt Injection and LLM06: Excessive Agency risks in AI applications.

The researchers noted that abuse involving AI tools may remain difficult to detect because normal browser activity and network connections can appear unchanged while unauthorized AI actions occur.

Google logo on a black background next to text reading 'Click to follow TechRadar'

Claude Cowork expands to mobile and web as Anthropic reveals what people actually use it for

  • You can now run Claude Cowork in the cloud, from the web or mobile
  • Knowledge work now accounts for around half of all Cowork sessions
  • Traditional local Cowork sessions are still supported

Days after reports surfaced that Anthropic could be bringing Claude Cowork to its mobile app, the company has gone one further – users can now start, monitor and complete their agentic workflows from the mobile app and a dedicated web portal.

The upgrade is rolling out in beta now for Claude Max subscribers, but the company has plans to bring the functionality to more plans as rollout continues.

As part of the upgrade, Cowork sessions will also run in the cloud by default – another beta introduction that means workflows can continue even once a PC goes offline or shuts down.

Claude Cowork can now be used virtually anywhere

Because the AI agent can run autonomously across things like files and documents, emails and calendars, and other connected apps, many users mostly left Cowork to run independently. However because it ran locally, it required users to keep their desktop session active even when they stepped away.

Now, scheduled work no longer requires a device to remain online – though users can still choose to run Cowork locally when access to local files is required, for example.

As for why Claude Cowork is being used, Anthropic has revealed that the autonomous agent is mostly being used among knowledge workers despite initially being targeted at coders. "Pulling scattered updates into a single report, building onboarding checklists and reconciling spreadsheets" account for the largest chunk, at around 33% of all use cases across Anthropic's analysis of 1.2 million sessions.

Content creation and copywriting (16%) came next, with software development (9%) and DevOps and infrastructure (7%) actually only accounting for much smaller proportions.

With knowledge work now accounting for nearly half of all Claude Cowork sessions, the company's research shows agentic AI emerging as an everyday work colleague. Though the company didn't indicate how, or whether, this shift in behavior might impact its pipeline, a shift away from coding as a primary use case could evolve Cowork in different ways to how we might have imagined.

Google logo on a black background next to text reading 'Click to follow TechRadar'

Anthropic launches "AI workbench" for scientists using Claude

  • Claude Science is a new “workbench” to consolidate fragmented research workflows
  • Everything from literature review to publication is handled on private infrastructure
  • Anthropic continues to roll out industry-specific AI tools for real-world use cases

Anthropic has introduced Claude Science – a new, beta AI workbench it says will let scientists consolidate fragmented research workflows into one unified environment.

With model capabilities no longer holding back AI adoption, the Claude-maker’s solution is to respond to today’s challenges, including limited use cases, struggles deploying AI in real-world environments and difficulties integrating multiple tools.

Claude Science represents this response, packaging existing capabilities into a purpose-built application for life sciences and scientific computing, following earlier work on MCPs, skills and other partnerships. An FAQ on Claude Science’s web page reiterates this: “Claude Science is a public beta app, not a model.”

Scientific ‘workbench’

Anthropic’s clearest message in the announcement is that scientific research is largely held back by workflow fragmentation, not model intelligence, with scientists already juggling tools like PubMed, Jupyter, R, a cluster terminal and more.

“Claude Science brings these fragmented tools into a single research environment where scientists can conduct all stages of their work,” the company summarized.

The platform should help scientists handle everything, from literature review and hypothesis exploration to analysis, figure generation, manuscript drafting and publication.

“Scientific research is inherently visual,” Anthropic wrote, acknowledging that many researchers are being held back in quickly and accurately producing visuals, which could need multiple revisions and finetunes before reaching production.

For full auditability, Claude Science also includes underlying source code, message history and plain-language explanations within AI-generated outputs for scientists to review and audit progress.

“It runs on your lab’s own infrastructure,” Anthropic added, referencing enterprise-grade laptops, Linux boxes or HPC login nodes, “so large or sensitive datasets never have to leave the systems they’re already on, and only the context needed for each step of the analysis is sent to Claude

Science is a growing focus for AI developers

Anthropic says early testers have already used Claude Science for single-cell RNA sequencing analysis, CRISPR screen design, protein structure prediction and cheminformatics, by the likes of Manifold Bio, Allen Institute neuroscientist Jérôme Lecoq, and ​​UCSF Brain Tumor Center associate professor and epidemiologist Stephen Francis.

The new tool represents a growing area of interest for AI developers, who are now targeting sectors with industry-specific tools rather than continually upgrading model capabilities without offering clear use cases. Until now, finance and legal have been a major focus for the likes of Anthropic and OpenAI, and this new science-focused initiative could mark the next stage.

It follows rival company OpenAI’s introduction of Prism earlier this year, described as an “AI-native workspace for scientists to write and collaborate on research” that launched with GPT-5.2 – the then-current model.

Claude Science is a separate app that’s available in beta for macOS and Linux installations to Pro, Max, Team and Enterprise subscribers.

The company has also committed up to $30,000 in credits for 50 lucky projects.

Google logo on a black background next to text reading 'Click to follow TechRadar'

I tried Claude Sonnet 5 with prompts that ask it to finish the job, not just answer the question — and that's where the AI war is going

Anthropic has just released Claude Sonnet 5 for all users, and I wanted to test what it was good at. But the game has changed now. Sonnet 5 doesn't feel dramatically different from Gemini or ChatGPT if you ask it ordinary chatbot questions. Instead, the difference should show up when you stop asking for answers and start asking for completed work.

Anthropic says Sonnet 5 is built for "multi-step software engineering work," sustained coding, tool use, debugging, and "messy technical contexts." It also says it can make plans, use browsers and terminals, and run more autonomously than smaller, cheaper models previously could.

I'm not using Sonnet 5 for coding, but that doesn't mean I can't take advantage of its new abilities — just like you can. So I stopped asking Claude for answers and started asking it to finish jobs, beginning with planning a trip to Bath, UK, for my family: my wife, me, and two teens.

A trip to Bath

When I tested it, Claude Sonnet 5 defaulted to its Medium level of effort, so that's what I used. Here's the first prompt I tried:

"I want to test whether you can act more like an agent than a chatbot.

My task is: Plan a weekend trip to Bath for two adults and two teenagers, including travel, lunch, one activity, estimated costs, and what still needs booking.

Don't just give me advice. First, make a brief plan. Then identify which parts of the task you can complete yourself right now, which parts require tools or information you don't have, and which parts need human judgment.

Then complete as much of the task as possible without stopping after the first obvious answer.

At the end, give me:

What you completed

What still needs human action

Any assumptions you made

A short checklist I can use to verify the result

The next best step"

What I really liked was that, as Claude tackled this task, it gave me the option to be notified when it had finished. In reality, it only took a few seconds to come back with a plan, which included travel options, an itinerary, and a suggestion for lunch and something to do: a trip to The Roman Baths.

To my delight Claude gave me an interactive map showing where all the places it recommended were. It also gave me a useful list of what it had completed, what required human action, the assumptions it had made, a verification checklist, and a "next best step" action point. It felt ready to keep working with me as more details came in, rather than treating its first answer as final.

In fact, when I gave it more details, such as which day I was going to go, it gave me a visual weather report for the day. That was a really nice touch.

Cladue Sonnet 5 maps.

Claude Sonnet 5 produced a handy map showing where to go. (Image credit: Anthropic)

Claude vs ChatGPT

I also tried this prompt with ChatGPT-5.5 Medium and got a similar result. It acted as an agent, just like Claude did, and notified me when it had finished its tasks. It just didn't look as nice. There was no map, or any visual elements at all, and it felt more like I had been given a finished report than the start of a two-way conversation where it asked me for more details.

Both chatbots recommended lunch and a trip to The Roman Baths. Interestingly, ChatGPT assumed I’d get the train, while Claude assumed I’d drive. They also recommended different places to eat, but the core information they both provided was solid.

What was most impressive was that both models could adapt when I reframed the inputs. For example, when I gave them the ages of the kids, student status, a different mode of transport, or changed the day of the trip, both models could cope. Both also identified that since the oldest was a university student, he could get free entry to The Roman Baths.

This part of the test was probably the most meaningful, as it felt much more "multi-step" than simply providing one answer.

Overall, I’d give this test to Claude. You can clearly see that Sonnet 5 is set up for agentic actions. Neither Claude nor ChatGPT could actually do any of the booking for me at the moment, so we're still a long way from true personal-assistant-level autonomy. But for this kind of task, Claude currently has the edge.

A different domain

I wanted to test the models in a different domain that would let Claude show me it had genuinely improved, and that the Bath trip result was not just a fluke of the travel-planning use case. So I asked them both to:

"Build me a simple household budget tracker as a spreadsheet or small tool."

Both models thought for a while about this task, and churned through various options before opting to make a spreadsheet. ChatGPT produced a spreadsheet with a bar chart that tracked how much I’d spent on various household expenses against a budget. Claude, however, went for something simpler: dispensing with a budget, it just tracked actual expenses and created a pie chart showing where my money was going.

Claude’s initial approach was simpler, and easier to understand. Both models provided a .xlsx file, but only Claude provided a button to upload it straight to Google Drive so I could open it in Sheets.

I told ChatGPT, "I wanted the graph to be a pie chart," and it responded: "Absolutely — I’ll update the spreadsheet itself so the dashboard uses a pie chart for spending by category, rather than the current graph style."

It ran into a few problems because it was trying to show both the budget and actual values in the same pie chart, but eventually it worked out that it could show only one and produced a new spreadsheet that did exactly what I asked for.

I then asked Claude to change its spreadsheet to provide a budget section too, and to change the graph into a bar chart. Again, it showed me its workings and added a budget section and bar charts perfectly.

I can’t really separate the two AI models on this task. Both proved they can handle multi-step tasks well, and both were happy to revise the result when I changed the brief.

That, really, is the point. The most interesting AI tests now are not "which chatbot gives the best answer?" They are "which assistant keeps working until the job is actually done?"

On that front, Claude Sonnet 5 feels extremely capable. ChatGPT was close behind, and in some ways just as effective, but Claude felt more naturally organized around the idea of completing work rather than simply responding to prompts. It asked fewer invisible questions, presented its output more helpfully, and made the whole process feel more like collaborating with an assistant than interrogating a chatbot.

For now, neither model is ready to fully take over the job. I still had to check the details, make the decisions, and do the actual booking or uploading myself. But the direction of travel is obvious. The AI war is no longer just about who has the smartest chatbot. It’s about who can build the assistant that gets you closest to a finished task.

Claude Sonnet 5 is here, and the 'most agentic Sonnet model yet' shows that the AI war is shifting from chat to agents

  • Anthropic has released Claude Sonnet 5, calling it its “most agentic Sonnet model yet”
  • The new model is designed to make plans, use tools like browsers and terminals, and run more autonomously
  • Sonnet 5 is available across Anthropic plans, and is now the default model for Claude Free and Pro users

Anthropic has released a new version of Claude, called Sonnet 5, which it’s calling “the most agentic Sonnet model yet.” Agentic models are designed to do more than simply answer questions. They can plan, use tools, and carry out tasks with less step-by-step input from the user.

According to Anthropic, the new Sonnet 5 can “make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.”

Sonnet 5 is aimed at coding and everyday professional work. Anthropic says the latest version outperforms the previous Sonnet 4.6, scoring 80.5% in Agentic Coding using Terminal-bench 2.1, compared to 67% for Sonnet 4.6.

Despite being aimed at professionals, the new release isn’t being restricted to paid users. It's available across Anthropic plans and is the new default model for Free and Pro users, and is available to Max, Team, and Enterprise users as well. It’s also available in Claude Code and on the Claude Platform.

The age of the agents

Claude Sonnet 5 on an iPhone.

Claude Sonnet 5 is now the default model, even for Free plan users. (Image credit: Anthropic)

The release of Sonnet 5 marks a wider shift in the AI race. Chatbots are no longer just competing to sound smarter in a conversation. They are increasingly competing to act like agents — tools that can plan, code, browse, investigate problems, and complete work with less hand-holding.

Sonnet 5 arrives soon after the release of Gemini Spark, Google’s 24/7 agentic personal assistant AI.

Anthropic is also launching Sonnet 5 at the same time that Fable 5 and Mythos 5 have become wrapped up in government scrutiny. Claude Fable 5 has just been re-released after being restricted by the US government, while OpenAI’s GPT-5.6 is still under review.

How AI models are changing

Claude Sonnet 5 may look like just another model launch, but it points to a bigger change in how AI companies are competing. The next stage of the AI war will not be won by the chatbot that gives the neatest answer. It will be won by the assistant that can take a messy task, keep track of the plan, and actually get something useful done.

AI assistants will increasingly complete tasks rather than just suggest steps. In this new future, the best model may not be the one with the cleverest answer, but the one that can finish the job.

Newsom strikes Anthropic deal to get California government half price Claude AI access

  • California government will have access to Anthropic's Claude with a 50% discount
  • The technology will be used to improve workflows and cybersecurity
  • Governor Newsom said Claude will be used "responsibly, transparently, and in service of people"

The government of California will now be able to use Anthropic’s Claude AI with a half price discount.

A press release published by California Governor Gavin Newsom says the state's government will also have access to free workforce training, expert GenAI technical assistance and workflow input from Anthropic developers.

“This partnership is about using technology the California way: responsibly, transparently, and in service of people. AI should not replace the human work of government; it should help our workers move faster, solve problems more effectively, and deliver better results for Californians,” Gov. Newson said.

Claude comes to California

The discount for Claude also extends to local governments at the city and county level, allowing state workers with lower budgets to gain access to cutting edge tech. Gov. Newson said the technology will primarily be used for drafting, summarization, and analysis, while also “supplementing day-to-day work and improving services for Californians.”

Californian state agencies will be able to access Claude through the Statewide Information Technology Shared Services (SITeS), a new portal that centralizes AI tools for government use. The Californian government has already worked alongside Anthropic to integrate Claude into numerous tools for state workers, such as the Engaged California tool that provides Californians with more of a voice in policymaking.

Claude Security and Claude Code are also being integrated into the workflows of the California Department of Technology and the California Governor's Office of Emergency Services in order to improve cybersecurity. The California DMV will use Claude to reduce wait times and improve services, and the Department of Healthcare Services will use Claude to assist in the state’s Medicaid program.

The state is home to 33 of the top 50 private AI companies in the world, including Anthropic. “As a California company, we feel a real responsibility to our home state. We’re honored to expand our partnership with California’s agencies and to put Claude to work for the people who keep this state running,” said Kate Jensen, Anthropic’s Head of Americas.

“Building AI responsibly and in service of people has been our approach from the start, and that’s exactly what this partnership puts into practice.”

❌
❌