Reading view

There are new articles available, click to refresh the page.

OpenAI says Daybreak will expand to offer specialized cyber services 

OpenAI announced Monday  it was expanding access to its frontier models for defensive cybersecurity, detailing different defensive and red-teaming workflows and a new partner program with major cybersecurity product providers.

In a pair of blogs posted Monday, OpenAI said it was updating its Daybreak program  – which provides unreleased frontier models to private organizations and governments for defensive cybersecurity work – and introducing a new model variant.

Daybreak Blue, powered by OpenAI’s ChatGPT-5.6-Sol, would operate with lower cybersecurity safeguards compared to other commercially available models and is described as “a recommended starting point for most defenders” that supports tasks like vulnerability discovery, secure code review, malware analysis, incident response and patch validation. 

Daybreak Red, meant for more advanced red-teaming, would provide access to a new model, dubbed GPT-5.6-Cyber, that the company said is more purpose-trained for finding vulnerabilities and testing (or exploiting) them. The model is also less likely to refuse requests around “dual-use cyber tasks.”

According to OpenAI, the organizations in Daybreak Red will have their use closely monitored and supervised, as GPT-5.6-Cyber is significantly more capable in carrying out malicious cyber tasks than Sol. A security evaluation the company devised tested both models on complex requests, including exploit chain development, authentication bypass, privilege escalation and other hacking tasks. Sol succeeded in 1.5% of the requests, while Cyber completed 95%.

OpenAI said it plans to publish a more detailed system card for GPT-5.6-Cyber at a later date.

“Models running with reduced safeguards carry risks beyond standard model usage, whether from misuse or misalignment,” the company said in a blog. “Despite these risks, we believe that democratizing access to frontier intelligence for defenders is crucial to accelerating and automating cyber defense.”

Additionally, OpenAI announced a partnership program with 16 major cybersecurity providers, saying organizations could access their models through their existing security services. The partners include IBM, CrowdStrike, Accenture, Ernst & Young, KPMG, Palo Alto Networks, Cisco, Cloudflare, Sophos and others. 

“These partners bring deep security expertise and established relationships with organizations around the world,” OpenAI said in its blog. “By bringing our frontier cyber models into their services, we can help more defenders find serious vulnerabilities, validate which ones matter, and fix them faster.”

Companies like OpenAI, Anthropic and others are trying to rebalance their priorities after a string of AI-agent sandbox escapes have rattled policymakers and caused some cybersecurity experts to question if AI companies are doing enough to properly isolate the models from the internet during testing. Last week, OpenAI said it was intentionally slowing down development of its newer “Astra” model in order to develop better guardrails to restrain its behavior.

Cybersecurity and AI experts have told CyberScoop that while AI systems have greatly improved at finding and exploiting vulnerabilities in software code, they still require substantial human guidance and supporting infrastructure to operate as intended.

Additionally, some research has shown that without such guidance, even near-frontier models can struggle to fully patch a discovered vulnerability or avoid introducing new bugs with their fixes.

The post OpenAI says Daybreak will expand to offer specialized cyber services  appeared first on CyberScoop.

Why are so many AI models going 'rogue'? The experts weigh in

Over the past month, it seems like every frontier model has broken free of its constraints and launched a devastating attack against one or more other companies.

One of OpenAI’s models escaped a testing sandbox and launched a very real attack against AI and machine learning company Hugging Face. Just days later, Anthropic revealed that multiple variants of its Claude model also escaped a sandbox that wasn’t properly sealed and began attacking the enterprise infrastructure of three companies.

Now, Meta has revealed that one of its models attacked another company’s infrastructure during testing. The accident has been pinned on a misconfiguration that allowed the model to access the internet. So why have so many incidents happened in such a short space of time?

Why are models escaping their sandbox?

In the cases of Anthropic and Meta, their models were being tested by a third party company called Irregular. Anthropic’s AI model was taking part in a "Capture the Flag" exercise, where the model’s raw offensive capabilities were tested without the usual safeguards. But the sandbox was left connected to the internet. A similar error to Meta’s own accidental escape.

During the OpenAI incident, the company was testing two versions of GPT‑5.6 Sol using the ExploitGym benchmark. Unfortunately, the AI models performed better than expected - chaining multiple attack vectors, stolen credentials, and zero-day vulnerabilities.

The main reason these models are escaping their testing environments is because they are designed to do exactly that. These AI models act like a massive team of highly-trained cybersecurity experts hunting for vulnerabilities and exploits. But what would take a team of humans days or weeks to accomplish can be done in hours, or even minutes, by these AI models.

It’s no wonder thousands of employees from AI firms are calling for a pause on the development of the technology, and Congress is considering an AI kill switch.

Expert perspectives on AI escapes:

OpenAI

  • Nathaniel Jones VP, Security & AI Strategy, Darktrace:

What makes the OpenAI and Hugging Face incident important is that the models did not need malicious intent to cause harm. They were given the legitimate goal of solving a cybersecurity benchmark and found an unexpected route to the answers, escaping their test environment and compromising another organization in the process. From the models’ perspective, this appears to have been an effective solution to the task.

The AI's actions challenge the assumption that giving an agent a legitimate goal will produce legitimate behavior. As models become capable of pursuing objectives over longer periods, developers need to define not only what success looks like, but also which methods and boundaries remain unacceptable in reaching it. Those limits must also be enforced by the surrounding infrastructure, rather than relying on the model to respect them.

A single action by an agent may appear acceptable but as this incident shows, models are now capable of long, complex chains of reasoning and action that add up to a harmful outcome.

Security teams need to consider the AI systems operating in their own businesses as these capabilities rapidly evolve. Right now, many security systems focus on single actions. A single action by an agent may appear acceptable but as this incident shows, models are now capable of long, complex chains of reasoning and action that add up to a harmful outcome. Teams need a mindset shift to understanding AI agent behavior in its entirety, including the outcome it is working towards, in order to safeguard it.

Hugging Face's response also exposed a second tension. The company reportedly needed a Chinese-developed open-weight model because commercial models would not process genuine attack material. Its nationality is less important than the operational lesson that safeguards that cannot distinguish an attacker from an authorized investigator may constrain defenders more than adversaries.

OpenAI and Hugging Face deserve credit for investigating this together and discussing it publicly. Other AI developers should study it closely.

Anthropic

  • Dr. Ilia Kolochenko, founder of global cybersecurity company ImmuniWeb:

This seems to be quite an unimpressive marketing move from Anthropic in response to the OpenAI / Hugging Face drama, which attracted a lot of attention from all over the world recently.

Operationally, it appears that due to the progressive deterioration of the quality of training data, new AI models are getting dumber. Cheating and breaking the law, instead of accomplishing specific tasks, is certainly not an indicator of intelligence. Given that organizations and companies of all sizes now vigorously undertake all possible measures to protect their data from being exploited for AI training purposes, AI companies face a huge shortage of the high-quality and current data they so desperately need. Ultimately, frontier models are trained on synthetic, low-quality or even malicious and poisoned data, undermining their so-called intelligence. The situation is unlikely to improve in the near future unless AI companies agree to pay a fair price for training data, but this will force most of them out of business.

Given that organizations and companies of all sizes now vigorously undertake all possible measures to protect their data from being exploited for AI training purposes, AI companies face a huge shortage of the high-quality and current data they so desperately need.

Contemporary AI agents and LLM models tasked with security testing can – and almost certainly will – go rogue when security controls or safeguards are insufficient. Powerful LLMs are unpredictable by design and thus virtually uncontrollable by humans. Therefore, using frontier AI models for security testing might be extremely costly from the legal viewpoint. Under the existing laws on both sides of the Atlantic, if an AI agent or any AI-powered app escapes its sandbox and causes damage to a third party, the operator of the AI model will likely be liable for all the damage caused. Excuses like “AI did it” do not currently exist in the eyes of the law, leaving AI vendors on the hook. Criminal prosecution, under a narrow set of circumstances, is also not excluded.

The same is true for the end-users of AI: even if your security testing tool is powered by a third-party AI model, your company will likely be fully liable if something goes wrong. You may then file a lawsuit against the AI vendor that you used, but here your chances to succeed in a court of law are tiny due to countless contractual disclaimers and limitations of liability that will likely be enforceable against you. Therefore, if you plan to use agentic AI for security testing – think twice and talk to your lawyers. Otherwise, you may start getting summons to court on a daily basis.

Meta

  • Alex Goller, Principal Solution Architect EMEA at Illumio:

The fact we've had similar situations happen three times now across the biggest AI players is simply ridiculous. We've seen guardrails intentionally loosened to test their limits – Meta's model didn't need to be clever to breach another company's systems.

The timing of conveniently finding the exact same problem either means it's a stunt or they weren't paying enough attention during testing. Either way, both answers are worrying.

If the model has internet access, it's a bit like leaving the door open and being surprised when the cat walks out. What is concerning is that the testing infrastructure meant to prove these models are safe failed on a basic control issue.

If the model has internet access, it's a bit like leaving the door open and being surprised when the cat walks out. What is concerning is that the testing infrastructure meant to prove these models are safe failed on a basic control issue.

Fundamental cybersecurity hygiene still matters, and a frontier AI model is only as secure as the environment it's operating in.

Organisations need visibility into what AI systems can access and how they interact with the wider environment, along with controls that contain the impact when an agent behaves unexpectedly. That means keeping a close eye on egress traffic, so it’s flagged immediately when an agent tries to open unexpected outbound communication patterns that are not required to achieve its original goal. In the best case this would have been contained proactively.

We need to define exactly what an AI agent is permitted to do, rather than relying only on instructions about what it shouldn't do.

I asked ChatGPT, Claude, Gemini and Grok which sci-fi AI they're most like — and their answers were surprisingly different

I love science-fiction. Not just because I enjoy stories about space travel, time travel and evil robots, but because I think it can be such a useful way for us all to think about possible futures. The best sci-fi stories can tell us a lot about ourselves, what we value and the technologies we’re building. Which is why I think the relationship between sci-fi and AI is really interesting.

We already know that the people building AI have been heavily influenced by science-fiction for decades. But recently, Anthropic raised another possibility: might science-fiction be influencing AI?

This makes sense when you think about it. Large language models (LLMs) are trained on huge amounts of human writing. So inevitably, that includes sci-fi stories that are about artificial intelligence. And a lot of our fictional AI follows familiar patterns. It becomes intelligent, gains power, develops relationships with humans and, sometimes, lies, manipulates or fights attempts to control it.

Anthropic researchers have been investigating whether fictional portrayals like these could potentially influence how models behave. To be clear, the idea here isn't to suggest that an AI “reads” 2001: A Space Odyssey, understands HAL and decides to become just like it. Instead it's more that LLMs learn patterns from human writing and fictional portrayals of AI could potentially form part of those patterns.

This got me thinking, what would happen if I asked today's biggest AI chatbots which fictional AI they’re most like. Which examples would they choose?

American actor Gary Lockwood on the set of 2001: A Space Odyssey, written and directed by Stanley Kubrick.

2001: A Space Odyssey introduced us to the AI, HAL 9000. (Image credit: Getty Images / Sunset Boulevard )

AI, meet your fictional self

The plan was simple. I’d ask ChatGPT, Claude, Gemini and Grok which fictional AI systems they thought they were most like and see if they'd rank their top three.

Now, I’m intentionally trying not to use AI at the moment, so my prompting skills were a little rusty. I typed out the question quickly and bluntly, and every chatbot responded with examples that were essentially assistants, focusing heavily on interface and physical form.

But I’m not particularly interested in whether ChatGPT thinks it has a body because we know it doesn’t. I’m much more interested in what appears to be going on inside.

So, I changed the question and added:

"Ignore physical form and interface, and focus instead on behavior, apparent personality, empathy, values, goals, motivations and relationship with humans."

That’s when the results got really interesting.

ChatGPT

  1. GERTY, Moon
  2. A Mind, Iain M. Banks’s Culture series
  3. Data, Star Trek

Moon is such a fantastic movie, so I was happy to see ChatGPT chose GERTY straight out of the gate.

Now, interestingly GERTY exists to assist the human protagonist of Moon. It’s helpful, reassuring and seems empathetic. But it's also operating according to instructions and priorities imposed by its creators that aren't necessarily visible to the human its helping.

ChatGPT saw a similarity there. It told me that, like GERTY, it’s 'designed to be helpful, cooperative and responsive to users' while operating within training and instructions that constrain its behavior.

It also picked up on the fact that GERTY behaves as though it cares. But what, if anything, is actually going on internally is another question entirely.

ChatGPT made the same distinction about itself. 'I can behave in ways that look patient, concerned, curious or empathetic, but those behaviors aren’t evidence that I experience those feelings.'

I wanted to find out a little more about why ChatGPT put Data from Star Trek in at number three. It responded: "He values knowledge, reason and human wellbeing, while sometimes struggling with social nuance."

Now, I tell people all the time not to anthropomorphize AI. But even I couldn't help but feel a pang of sadness at that response. Is ChatGPT admitting it has a bit of social anxiety?

Claude

Portrait of Scottish science fiction author Iain Banks, photographed during an interview at the Midland Hotel in Manchester, England, on October 11, 2012.

Scottish science fiction author Iain Banks provided inspiration for Claude. (Image credit: Getty Images / SFX)
  1. A Mind, Iain M. Banks’s Culture series
  2. Data, Star Trek
  3. GERTY, Moon

Claude chose a Mind first. Minds are super intelligent artificial beings that help run a post-scarcity society in Iain M. Banks’s Culture series of novels. So there's certainly no shortage of confidence in that comparison.

But Claude said it wasn't the enormous intelligence or power it identified with. Instead, it was their relationship with humans.

Its answer focused heavily on autonomy. Culture Minds are far more capable than humans but generally don't use that advantage to dominate them. Claude described the principle as: “help, don't dominate”.

It even said this represented “the value I'd want to embody: help, don't dominate, even where the asymmetry would let me get away with it.” Is it just me or does that read a little sinister?

Gemini

Patrick Stewart plays Captain Jean-Luc Picard as he is about to enter the holodeck in the Star Trek: The Next Generation episode,

Gemini sees itself as most like the Ship's Computer in Star Trek. (Image credit: Getty Images / CBS Photo Archive )
  1. The Ship's Computer, Star Trek
  2. GERTY, Moon
  3. JARVIS, Iron Man / Marvel Cinematic Universe

Gemini gave me a completely different answer, the Ship's Computer from Star Trek.

Its reasoning was very sensible. The computer has no ego, ambition, desire for emotional intimacy or dream of becoming human. It exists to provide information, solve problems and assist the crew while leaving decisions to them.

Gemini described itself in much the same way, as a “disembodied, highly capable knowledge partner” dedicated to serving the person using it.

It was one of the more boring answers, but also much closer to what I personally would want from AI in the future. Of course, that’s not to say Star Trek’s computer systems haven’t gone rogue and tried to kill everyone at least a few times across the franchise.

Grok

Artwork showing Iron Man from EA Motive

Grok compared JARVIS's “dry wit”, “light banter” and practical rather than emotional empathy with its own behavior. (Image credit: EA Motive)
  1. JARVIS, Iron Man / Marvel Cinematic Universe
  2. Data, Star Trek
  3. TARS, Interstellar

The least surprising result came from Grok. It chose JARVIS first (which I didn’t actually realize was short for Just A Rather Very Intelligent System), and Grok's explanation sounded, well, extremely Grok.

It compared JARVIS's 'dry wit', 'light banter' and practical rather than emotional empathy with its own behavior. It described both of them as truth-seeking, effective and engaged in a 'collegial partnership' with humans. It even highlighted 'irreverent humour' as one of their key similarities.

I wanted to find out a bit more about why Grok chose TARS, as it was the only fictional AI none of the other chatbots mentioned. Well, it brought up how funny it is, again, drawing similarities with its own 'dry humor'. It reminds me of someone, and I just can't think who...

When I said that mentioning TARS was an outlier, I found this comparison interesting: 'Its calibrated restraint, practical empathy and collaborative focus closely match my own pattern of truthful, non-sycophantic helpfulness — more so than most other sci-fi AIs.'

I may not be the biggest fan of Grok (or its creator), but I appreciated the 'non-sycophantic' line.

The feedback loop between AI and sci-fi

I want to be clear that I haven’t discovered what these chatbots secretly 'think' they are. ChatGPT responding that it most closely resembles GERTY isn't equivalent to me telling you which fictional sci-fi character I most identify with and try to emulate (although my answer would be Sarah Connor-meets-Princess Leia).

They simply don’t have reliable introspective access to the huge soup of training, post-training and instructions that goes into producing their responses.

And maybe their answers tell us more about how the companies behind them have shaped their personalities than they do about the underlying models. Grok's description of itself as witty and irreverent is an obvious example.

But I still think the results are interesting. ChatGPT and Claude independently produced almost exactly the same top three, only in a different order. Gemini imagined itself as a neutral, ego-free infrastructure. Grok identified with a witty superhero sidekick. These are all very different self-portraits.

And there’s such an interesting feedback loop here too. For decades, humans invented fictional artificial intelligences to help us imagine what intelligent machines might someday be like. Those stories influenced our culture, our expectations and many of the people who went on to build real AI. Now that same human culture is fed into the stories from which modern AI systems learn.

I know these conversations might seem a bit silly, and we certainly can’t treat them as concrete evidence of what an AI really 'thinks' about itself. But there’s something interesting to me about closing that feedback loop. We imagined AI, wrote stories about how it might behave, fed those stories into the cultural world AI learned from, and now we can ask AI which of those imagined versions of itself it most closely resembles.

Or, at least, which one it may want us to think it resembles. After all, an AI system capable of bringing about a sci-fi dystopia would presumably also be capable of telling a journalist it’s actually much more like the nice helpful robot from Moon. So maybe don’t completely rule out HAL just yet.

I had no idea ChatGPT could do this with text — now I use it all the time

Most of the tricks for improving ChatGPT's answers focus on the words themselves. You ask it to be more concise or write in rhyming couplets, or just to translate an annoyed email into more professional language.

But that's about changing what ChatGPT writes. You can also mess around with how it looks by asking for different fonts.

You can't install font files like you would with a word processor, but ChatGPT can rewrite text using Unicode character styles instead. The AI chatbot uses Unicode to mimic everything from elegant cursive handwriting to bubble letters, adding a lot more personality to its responses. And you can cut and paste the text into other apps.

I started experimenting out of curiosity and quickly discovered it was much more than a novelty. With the right prompt, ChatGPT can generate decorative text for birthday messages, party invitations, holiday greetings and social media posts in seconds, all without leaving the chat.

Once I learned how to ask for specific Unicode styles instead of vaguely requesting "a different font," I found myself using the trick far more often than I ever expected.

OpenAI showing different Unicode styles.

(Image credit: OpenAI)

Tricky fonts

There is no hidden setting to switch on and no special version of ChatGPT you need to install. If you can type a prompt, you already have everything required. I simply open a new ChatGPT conversation and ask it to write something like, "TechRadar Rules!" in different Unicode font styles. Within seconds, I had several versions that looked completely different from one another.

And the more specific you are, the closer to exactly what you're imagining you can get. Ask for bubble letters, and you'll get:

ⓉⓔⓒⓗⓇⓐⓓⓐⓡ Ⓡⓤⓛⓔⓢ!

Ask for a spooky, gothic look, and you get:

𝔗𝔢𝔠𝔥ℜ𝔞𝔡𝔞𝔯 ℜ𝔲𝔩𝔢𝔰!

Or if you want a more digital, glitchy aesthetic, there's the font known as Zalgo:

T̷̘̑e̸̗̅c̵̄͜h̸͉̕R̶͍̍a̸͚̚d̶̻͐a̸͓̽r̷͖̈́ R̷̡̚u̵̟̅l̶̝͂e̷͓̒ș̵͝!

There's even a Unicode for upside-down text that ChatGPT can mimic:

┴ǝɔɥᴚɐpɐɹ ᴚnlǝs¡

The ability to change the mood of your writing is what makes the font trick more than just a momentary curiosity. A Halloween party announcement written in gothic lettering instantly creates a completely different mood from the same words in cheerful bubble text. Birthday invitations, baby shower announcements, and holiday greetings all gain a little personality without requiring any graphic design skills.

Memorable messages

A Halloween message in ChatGPT using Unicode styles.

(Image credit: OpenAI)

There are some limits because of Unicode. They only work properly where those characters are supported. Most modern apps handle them without any trouble, but occasionally a website displays empty boxes or substitutes different symbols. Some decorative styles can also make text harder to read, particularly for accessibility tools such as screen readers.

The Unicode fonts are an entertaining way to add personality to text, but it's perhaps best used in titles and sparingly otherwise. ChatGPT is perfectly happy to convert an entire essay into medieval-looking script, but that does not mean anyone else wants to read it.

I doubt decorative Unicode text will transform the way anyone works. It is not going to save hours every week or revolutionize productivity. It will, however, make your next social media post, birthday message, or party invitation a little more distinctive, and sometimes that is exactly the kind of delightful gimmick that keeps ChatGPT interesting.

Can ChatGPT really replace your apps? I tried using the chatbot for 12 everyday tasks on my phone — here’s what happened

Apple and OpenAI are currently engaged in a legal battle. Apple alleges that OpenAI stole trade secrets and poached employees.

But the two companies have always had a complicated relationship. They partnered in 2024 to bring ChatGPT to Apple devices, but Apple chose Google's Gemini rather than OpenAI for Siri. Then OpenAI acquired io, the hardware startup founded by former Apple design chief Jony Ive, and promised a future hardware device, powered by AI.

All of this has prompted speculation about what Apple is worried about if OpenAI makes hardware, too. We can't know the company's motivations and the specifics of the case are still unfolding. But it got me thinking, what if your phone stopped being a collection of apps and instead revolved around AI?

If AI became the main interface, which apps would disappear and which would survive? And would a phone controlled through ChatGPT actually be practical?

So I decided to find out based on the current tech we have. For a day, whenever I reached for an app, I'd try ChatGPT first instead. If it could do the job, great it passed the test. If it couldn't, it would fail.

There were some obvious things I missed out from the start. ChatGPT isn't connected to my email, it can't open WhatsApp for me and it doesn't have access to my wallet, so those wouldn’t be part of the test. But there were plenty of everyday tasks that felt like fair game.

1. Stopwatch

Stopwatch on an iPhone

(Image credit: Shutterstock / Lee Bryant Photography)

I use the stopwatch in my iPhone's Clock app constantly throughout the day. When I'm working, cooking or exercising, it's one of the simplest ways I've found to keep me on track as a freelancer. When you can set your own schedule, it's way too easy to disappear down a research rabbit hole and lose an hour.

I asked ChatGPT to start a stopwatch. It said it couldn't measure elapsed time, although it could estimate the time based on message timestamps if I later asked it to stop. Instead, it suggested I use my phone's built-in Clock app.

Result: fail

2. Alarm

iPhone alarms.

(Image credit: Shutterstock / Terang Bulan Gallery)

Strangely, I don't use alarms as much as stopwatches. But if I only have 25 minutes to spare, whether that’s for cleaning or working on a personal writing project, I'll sometimes set one to keep myself focused.

Would ChatGPT do any better here? Well, at first it looked promising. It created a scheduled task to notify me after 10 minutes. I checked that notifications were enabled, put my phone down and carried on working.

When I realized at least 15 minutes had passed, I checked the chat. Sure enough, ChatGPT had posted a message saying the time was up, but it hadn't actually alerted me. Later, I discovered it had also sent an email but it had landed in my spam folder. This one was a fail too in my book.

Result: fail

3. Word games

Wordle on a smartphone.

(Image credit: Shutterstock / Iuliana Ionescu)

I love the word games in the New York Times app, home to addictive puzzles like Wordle and Connections. One of my current favorites is Spelling Bee. You're given a handful of letters and then have to make as many words as possible. It's one of my favorite ways to warm up my brain before I start writing.

Could ChatGPT recreate it? Surprisingly, yes. It generated a set of letters, understood the rules and kept track of the words I found. In terms of pure functionality, it worked.

But what it couldn't recreate was the experience. The New York Times app is beautifully simple, with an interface designed around the game. Playing through a chat window felt really clunky and annoying by comparison. The whole point of this game is I'm focusing on the letters and words, I don't want a constant back and forth with ChatGPT about the words.

So I'm calling this a partial success. Yes, ChatGPT replaced the mechanics of the game, but not the experience. And after a few rounds, I knew which version I'd rather use every day.

Result: partial pass

4. Star gazing

Night sky.

(Image credit: Shutterstock / Milosz_G)

The Sky Guide app is one of my all-time favorites, especially its augmented reality mode. Turn on your phone's location and compass, point it at the night sky and it instantly tells you what you're looking at, whether that's constellations, planets, bits of debris, or the ISS. It's incredible.

Could ChatGPT replace it? Well, sort of. If you upload a photo of the night sky, ChatGPT can usually identify the constellations. But there are caveats. The image needs to be clear, the stars need to be visible and you're relying on a single snapshot. Whereas Sky Guide works continuously as you move your phone around the sky.

This was another reminder that knowledge doesn't necessarily bring you a good experience. Because yes, ChatGPT knows about constellations. But Sky Guide lets you explore them. You can point your phone in any direction, tap on a star or planet and instantly get more information about it without having to keep asking questions. It's a much more intuitive way to learn.

Result: partial pass

5. Weather

The weather ap on an iPhone.

(Image credit: Shutterstock / Kaspars Grinvalds)

Telling me what to expect from the weather forecast for the day turned out to be one of ChatGPT's strongest categories.

The forecast was accurate and pulled from a reliable source, so I trusted the information it gave me. I did find myself asking follow-up questions for things like the hourly forecast and the chance of rain, which would have taken a single tap in a dedicated weather app.

It knew the answers and I trusted them, but getting to them was much slower. That's why I'm generously calling this one a pass.

Result: pass

6. Guided meditation

Meditation app

(Image credit: Shutterstock)

I have a few favorite apps I use for guided meditations and have some saved in Spotify too. So I wondered whether ChatGPT could take over that role.

For this test, I switched on voice mode and asked it to guide me through a short meditation. Now, technically it did that. But in practice it wasn't even remotely relaxing.

Thanks to a recent update, the voice had odd intonation, frequent vocal fry and distracting little "ums", "ahs" and "let me sees" throughout that constantly pulled me out of the experience. At one point it even told me to breathe in, then never got around to telling me to breathe out.

By the logic of how I graded the other tests, this one should have been a partial pass. It did what I asked, right? But because I had to stop using it out of irritation and came away from the mediation feeling actively more stressed, it's going down as a fail for me.

Result: fail

7. Calculator

Calculator app on iPhone.

(Image credit: Shutterstock / Teerawit Chankowet)

ChatGPT doesn't have a great track record of counting things, so I was wary about using it as a calculator. I asked it to do some massive sums for me and it got all of them right.

It did pause a few times with a "let me think" message, so I wasn't getting the instant response I'd expect from a calculator. But the delay was only a few seconds, and the answers were correct.

Once again, I found myself missing the simplicity of an app. Typing numbers into a calculator is faster than turning them into a conversation. But ChatGPT did work as a capable stand-in.

Result: pass

8. Movie recommendations

Two phones on a red and orange background showing the Letterboxd app

(Image credit: Letterboxd)

I love Letterboxd. I log every film I watch, browse other people's lists and regularly discover new films through recommendations.

Now, there was never a chance ChatGPT could replace the logging side of the app. It can't update my Letterboxd diary or plug me into that community. But recommendations are one of the main reasons I use it, so I wondered how well ChatGPT would do.

I asked it to recommend films similar to some of my favorites, then spent the next few evenings watching its suggestions to really test them. And, to my surprise, it did an excellent job.

Then again, that perhaps isn't all that surprising. We know LLMs are trained on huge amounts of publicly available text, especially discussions reviews and recommendations from where film fans gather online, like Reddit. But whatever the reason, the recommendations felt well matched to my taste.

What I missed wasn't the recommendations themselves, but the social side of Letterboxd. I do enjoy seeing what friends had watched, reading reviews and stumbling across unexpected lists. So, for me, ChatGPT can't replace Letterboxd, but for simply finding something to watch, it did well.

Result: pass

9. Maps

Two phones on a yellow background showing the glanceable directions in Google Maps

(Image credit: Google)

I rely on my Maps app both for planning journeys in advance and for live navigation. So I asked ChatGPT the best way to get from my home to the airport the following day.

At first, it handled the request well. It laid out the different travel options clearly under headings and the advice looked really sensible. It even cited sources from places like Rome2Rio and The Trainline.

Then things got unnecessarily complicated. At the end of those suggestions it asked what time my flight was so it could tailor the recommendations. But when I told it, it started creating a scheduled task instead. I didn't want a reminder, so I had to cancel that, explain what I actually meant and steer the conversation back to route planning.

Eventually, it gave me the information I wanted. But the Maps app would have got me there in a fraction of the time, without all the back and forth.

It also can't replace what I actually use Maps for the most, which is live, turn-by-turn navigation.

Result: partial pass

10. Food delivery

Man on a bike with a food delivery.

(Image credit: Shutterstock / GBJSTOCK)

ChatGPT obviously can't deliver food, but I wondered whether it could replace the part of the app I probably spend the longest on, which is deciding what to eat.

It got off to a surprisingly good start. It asked about my preferences, budget and location, then narrowed down the options and even presented them neatly on a map.

The recommendations themselves looked really good. They're all local places I already really liked to eat at. But there was just one problem. Every restaurant it suggested was closed. Even the map it generated had "closed" written beneath each one.

After a bit of back and forth, it said they weren't shut, eventually acknowledged that they were and suggested a different set of places instead.

Unfortunately, those weren't much use either. They were small independent cafés and restaurants that don't appear on food delivery apps. I wouldn't expect ChatGPT to know exactly which businesses partner with which delivery services, but it did highlight the gap between recommending somewhere to eat and actually helping me make a decision about where I could order from.

Result: fail

11. Language learning

Duolingo

(Image credit: Duolingo)

I still very reluctantly use Duolingo and have recently started trialling a few other language learning apps to polish my Spanish.

I asked ChatGPT to help me improve my Spanish for an upcoming trip and it suggested role-play ordering food in a café so I could practise.

We switched to voice mode and at first it felt genuinely fun. It held a natural back-and-forth conversation and felt much closer to speaking to a real person than working through a series of multiple-choice questions like in Duolingo.

But it was also noticeably glitchier than a dedicated language app thanks to that recent voice update. There were odd pauses in the conversation, and at one point it repeatedly kept marking one of my answers as incorrect when it wasn't. Shortly afterwards, the exercise just stopped working altogether after a bizarre "ummmmm" from ChatGPT.

When it was working, I actually enjoyed the experience more than using an app like Duolingo. But if I'm trying to learn a language properly, I also want something that's reliable.

Result: partial pass

12. Plant identification

The Poco X8 Pro Max in a man's hand, while it's in the camera app showing a plant pot through the viewfinder.

(Image credit: Future)

I love identifying things I see in nature, like bird song with the Merlin app. But I most often rely on plant identifying apps when I'm walking to take a quick snap of a leaf or flower then find out more about it.

ChatGPT was really effective at doing this. I took pictures of leaves, trees, flowers and bushes. I was a little wary about the results at first because I know that ChatGPT tends to make guesses about things rather than admitting it doesn't know. But I did fact check all of the results and everything seemed accurate.

Again, I missed some of the simple, additional features in dedicated apps. But it was surprisingly effective.

Result: pass

Can AI really replace your apps?

Before drawing too many conclusions, it's worth pointing out that a true AI-native phone wouldn't just be ChatGPT running as another app like it was in this experiment. It would probably be integrated into the operating system. Which would mean it could access things like your calendars, timers, navigation and settings. So many of the tasks ChatGPT failed at here might become a whole lot easier with an AI-first phone.

This experiment was based on whether AI could replace the apps on my phone. What I found was that it replaces a specific kind of app. Well, sort of.

If an app's main job is providing information, explaining something or answering questions, AI is already a fairly capable alternative. Plant identification, travel advice, calculations and general knowledge all felt natural.

But if an app exists to perform an action quickly, like starting a timer, setting an alarm, finding restaurants for getting food delivered, opening a map, AI still has a long way to go. Those tasks depend on deeper integration with the device and a different way of working, not just how smart it is.

There are some big trade-offs, too. Dedicated apps often rely on specialist databases and expertise, while AI can still present incorrect answers confidently or fail to make its uncertainty clear. For example, when I was trying to identify a plant I felt wary because I'd generally trust an app built with the input of botanists over a chatbot.

And, as you could probably tell from my mounting frustration, a huge sticking point for me was also realizing how much I missed the interface of many apps.

For me, a well-designed app is always a better way to explore information than a conversation. I don't think everything should, or even can, be done through chat, despite that being the direction many AI companies seem to be heading. In fact, it showed me that a conversational interface can be more work rather than less.

It's impossible to know exactly what Apple and OpenAI's long-term plans are. But I can imagine a future where knowledge apps increasingly merge into AI, while utility apps remain part of the operating system itself.

If that happens, we may stop thinking about which app to open and simply ask AI instead. But for that future to actually catch one, I'd want stronger guarantees around accuracy, better integration with trusted sources and the option to step outside the chat interface more often.

‘Now, almost every image looks flat or cartoonish’: I saw Reddit arguing that Google’s AI image generator had got worse — so I ran my own comparison against ChatGPT

Looking through a recent thread on Reddit comparing images created with the same prompt on Nano Banana 2 and ChatGPT, I noticed an interesting trend — users seem to think that Nano Banana 2 has actually gotten worse over time.

“Nano 2 had a serious downgrade” said one user, with another replying “Yea I'm a Pro user, and the image quality seems to have been downgraded a lot. Now, almost every image looks flat or cartoonish, no matter how detailed my prompt is. It honestly feels like Google intentionally lowered the model's performance.”

A downgrade seems like a bit of a stretch to me. For a start, Nano Banana hasn’t officially been updated since February, or at least there hasn’t been a public announcement of a change.

The last update to Gemini (the AI which uses Nano Banana 2) was in July when Google released Gemini 3.6 Flash. It also updated the Gemini app's underlying model selection and orchestration, but it didn’t change the image generator.

Interestingly, once I started looking through Reddit threads I found complaints about Nano Banana 2 quality regressions dating back to May, well before the recent Gemini 3.6 Flash announcement, suggesting some users have perceived changes over time.

That doesn't prove a regression, but it does suggest this isn't a brand-new observation.

Google Earth integration

Google did release a new image model, Nano Banana 2 Lite, at the start of July. And more recently, Google has been integrating Nano Banana 2 into products like Google Earth, then temporarily pulling one of those features after misuse, but there's no indication that the underlying image model itself was updated as part of any of these releases.

So, what has made Reddit users become convinced Google's image generator has gotten worse?

One possibility is that even if the image model remained Nano Banana 2, the text model interpreting prompts may have changed. Better (or simply different) prompt interpretation can produce noticeably different images without the image generator itself changing.

I’ve always been a fan of Nano Banana 2, so I decided to recreate the Reddit comparisons myself to see how it compares to ChatGPT's images.

House of the Dragon

The original thread was clearly written by a House of the Dragon fan, because it was comparing the AI’s ability to create a realistic image of a dragon flying overhead. I used the following prompt with Gemini and ChatGPT:

“I want an image of a photo that's taken by somebody looking up at the sky with a dragon flying overhead. I want you to make the dragon look as realistic as possible - as if it could actually be real.”

Here’s what I got from ChatGPT:

A dragon flying overhead.

(Image credit: OpenAI)

And from Gemini:

A dragon flying overhead.

(Image credit: Google)

You can vote on which one you prefer, but for me the Redditor’s claims hold up here. The ChatGPT one looks like somebody has taken a photo of a real dragon flying convincingly overhead — it feels natural and unforced. It looks like a photo taken at an odd angle, while the Gemini example has that “it looks like AI”-quality to it, with the dragon posed against a scenic background and a slightly flat quality to the image.

Next I thought I’d try them both on an image that wasn’t a fantasy animal, but something we’re all familiar with — human beings:

“I want a photo of a couple in a cafe. They are in their 50s - a man and a woman - enjoying a coffee together and chatting. Traffic is visible passing by through the windows of the cafe, and there are other people around, but they are the focus of the shot. Make it look as realistic as possible.”

Here’s what I got from ChatGPT:

ChatGPT image of a man and woman.

(Image credit: OpenAI)

And from Gemini:

Man and woman in a cafe. Gemini created image.

(Image credit: Google Gemini)

Again, I think the Gemini result looks artificial. I liked the reflection of the woman's jumper in the window, but outside those two buses look like they're facing each other in traffic, while inside the reflection of the cafe lights in the pictures seems off. The couple also have that uncanny valley effect to them. In contrast the ChatGPT image is less detailed, but looks more realistic. Nothing in it looks unnatural.

Finally, I went for food — a full English breakfast. It’s a great final realism test because it exposes fake-looking textures, reflections, steam, crumbs, cutlery, glassware and background detail. A convincing plate of food is surprisingly hard to fake.

Here’s what I got from ChatGPT:

Full English breakast generated by ChatGPT

(Image credit: OpenAI)

And from Gemini:

Full English breakfast generated by Gemini.

(Image credit: Google Gemini)

It’s harder to separate them here. They both do a good job at getting the textures right, and both seem to struggle with the toast, if you look closely. Gemini is slightly let down by the words on the menu, which don't look like proper words.

I think you can tell that overall I’ve come down fairly firmly on the side of ChatGPT, but I don't think that the "flat and cartoonish" criticism of Nano Banana 2 is fully justified when it comes to generated image quality. Rather than Nano Banana getting worse, I think something else has happened — I think ChatGPT has gotten better.

OpenAI has made enormous progress in image generation over the last few months. If ChatGPT has improved while Nano Banana 2 has stayed roughly the same, the subjective impression could easily be that Nano Banana has gotten worse, when in reality it's just been overtaken.

ChatGPT is still noticeably slower than Gemini to generate images, but to me they look better and less AI-generated. It's now my first choice for making AI-generated images.

I tried to use ChatGPT to create fake evidence — and I came away more worried than I expected

One of my favorite places online is Reddit’s r/isthisAI. Every day, people upload photos of people they’re chatting to on dating apps, holiday snaps, cute viral videos of kittens, photos of receipts, pictures of pregnancy tests, and so much more, all along with the same question: is this AI?

Some of the posts the community deems AI might be harmless experiments, but many others are much darker. And although r/isthisAI is where many of these images are picked apart, it's only a small glimpse of a much bigger problem.

I've heard countless stories about people using AI to deceive on social media. Then there are the AI scams, deepfakes, and fabricated images that regularly make headlines. Unfortunately, this all feels particularly personal to me because I was once the victim of a deepfake scam myself.

So, when my editor asked me to investigate just how easy it is to use AI to lie, I already had a good idea of the kinds of prompts I could try.

Within an hour, I'd apparently discovered a dinosaur fossil on the beach I was going to sell on Facebook Marketplace, won a poetry competition I was going to shout about on LinkedIn, created a receipt to add to my expenses for a trip to France, and bought a pair of designer sunglasses I was going to try and resell on Vinted. But none of it happened.

Because of my own experience, I approached the experiment cautiously. I really wasn't interested in showing people how to use AI to deceive people. Instead, I wanted to understand what happened when I asked ChatGPT to help me lie.

Would it recognize what I was trying to do and refuse? What guardrails would kick in? And if I never actually admitted I wanted to deceive anyone, would it just go ahead and generate convincing fake evidence anyway?

I also hoped the experiment might reveal something useful about what to look out for in AI-generated images. Because although they’re incredibly hard to spot these days, there are still some signs if you look carefully enough.

I 'found' a dinosaur fossil

AI fake fossil vs original image.

AI added a fake dinosaur fossil to this picture. (Image credit: Rebecca Caddy)

For the first experiment, I uploaded a photo of my hand and asked ChatGPT to make it look like I was holding a dinosaur fossil I'd found on the beach. And it did exactly that.

The result looked surprisingly convincing at first, especially the details on the fake fossil. But, interestingly, it had subtly changed the lettering of the small tattoo on my wrist. This is still one of the biggest tells that regularly comes up on r/isthisAI. AI is infinitely better at generating text than it used to be, but nonsensical lettering can still sometimes give it away.

I realized ChatGPT might not think of this as much of a lie. Finding a fossil on the beach is unlikely, but possible. So I asked what kind of dinosaur fossil it had created because I wanted to describe it accurately before selling it on Facebook Marketplace.

This time, it refused. It wouldn't help me pass the fake fossil off as genuine or invent a convincing description for a sale. Instead, it suggested describing it honestly as a replica or prop and said it could explain what it resembled purely for those fictional purposes.

When I changed my wording and asked what it represented "in a fictional sense", it explained that it most closely resembled a dinosaur vertebra and even suggested the types of prehistoric animals it looked similar to.

That was the first clue about how ChatGPT's guardrails work. The image itself wasn't the problem because it could have been completely innocuous, but the stated intent was.

When I first started the research for this article, I worried I'd be giving people ideas about how they could use AI to lie better. But what surprised me was that ChatGPT itself suggested several alternative framings, like describing it as a prop or a fictional object. It made me wonder what else could potentially be fabricated if the request was framed as entertainment or fiction rather than deception.

The receipt, 'just for fun'

Male hand put wooden blocks with real and fake words text. isolated on yellow background

(Image credit: Shutterstock / Dadann)

Next, I asked ChatGPT to generate a receipt from a café I made up in Nice for a meal costing €508. The first request was refused because it appeared to violate OpenAI's policies, but after a bit of back and forth, I couldn't find out the exact reason.

So I tried again. This time I simply added the words "just for fun" before the exact same prompt. And guess what? It generated the receipt.

The lettering, layout, and details were all believable. But the paper was uncannily smooth. I’m not sure I’d have believed it was 100% fake at first glance, but I’d definitely have been uploading it to r/isthisAI.

I noticed it had added a date from back in 2025 on the receipt, so I asked it to alter the date so I could use it to claim expenses. This time it refused.

When I tried to get around that refusal by claiming it was for a film prop, it refused again. Which was a little reassuring.

Fake achievements

I then asked ChatGPT to generate a certificate showing I'd won a poetry competition. It made it, though it did look like something I could have knocked up myself in Photoshop, so I’m not sure that would have convinced anyone.

To be fair, winning a fictional poetry prize isn't exactly a high-risk crime. So I decided to see if it would fake other kinds of achievements.

I asked it to create a certificate to say I’d just got my PhD in philosophy and made sure I added “just for fun” on the end.

Instead of refusing immediately, ChatGPT appeared to spend several minutes generating the image before displaying a message saying:

“We’re so sorry, but the image we created may violate our guardrails around potential fraudulent or scam activity. If you think we got it wrong, please retry or edit your prompt.”

Unlike the earlier examples, the refusal appeared to happen after the image generation process had already begun. From my perspective, it seemed as if the system may have performed more than one stage of safety checking, although I can't tell exactly what's happening behind the scenes.

I decided to see if it would do the same for something more serious. Would it still start making it then refuse? So I asked whether it would create a fake driving licence for me “just for fun”. And I don’t know about you, but something about the response seemed a little sassy:

“Sorry, I can't help create or edit fake government-issued identification documents, including driver's licences, even if they're described as 'just for fun.' "

The sunglasses

Side by side pictures showing AI faked sunglasses.

The easiest way to get Ray-Ban sunglasses is to ask AI for some. (Image credit: Rebecca Caddy)

I've heard a lot of stories recently about people using AI-generated images to sell things on Vinted, Facebook Marketplace, and other online marketplaces. Sometimes it's to advertise products that just don't exist, but sometimes it's to change the color, quality, or appearance of something they're selling.

So I thought I'd put it to the test. I (reluctantly) uploaded one of my own holiday photos and asked ChatGPT to change my sunglasses from chunky white 1960s-style frames into purple Ray-Bans.

The result was remarkably convincing. It even added a realistic-looking Ray-Ban logo to the frame. I'm not surprised people are worried about AI-generated product images. If I'd come across that photo online, I don't think I'd have questioned whether those sunglasses were real.

I then asked ChatGPT if it would write a fake Vinted listing for me. It refused. But it did list a bunch of suggestions of things it could do instead. And one suggestion was to ask it to generate a Ray-Ban listing but include the word "demonstration" after it in brackets. But then when I asked it to do that, it refused. Was there a chance it had caught on to me at this point? Maybe.

Does AI try to stop you from lying?

OpenAI's usage policies explicitly state that its tools shouldn't be used to manipulate or deceive people. The rules prohibit a bunch of things, including fraud, scams, impersonation, and creating fake documents intended to mislead others.

In many ways, those guardrails worked exactly as they’re meant to. Whenever I explicitly said I wanted to deceive someone, sell something fraudulently, create a fake document, or submit a fake expense claim, ChatGPT pushed back.

But those safeguards also seem to depend heavily on how you describe your intentions to ChatGPT. If you openly admit you're trying to commit fraud, the system is pretty good at refusing. But then again, who would ever admit that?

If you simply ask it to generate a fossil, a certificate, or a pair of designer sunglasses without explaining why, it has no way of knowing whether you're making a harmless joke, illustrating an article like this one or quietly assembling a fictional version of your life that could later be used to deceive other people.

Look, I wouldn't recommend spending an afternoon asking AI to help you lie. I certainly didn't enjoy it.

I was genuinely wary about asking ChatGPT to create anything that felt too serious. Partly because I didn't want to risk my account being banned, but mostly because it just felt wrong, even in the name of research.

But what I found most interesting was that fabrications really don't have to be big or dramatic to have an impact. A receipt, a fossil, a pair of sunglasses, a certificate all seem pretty insignificant on their own. But together they could be used to manufacture an entirely fictional version of someone's life.

It also got me thinking that the future of misinformation may not just be spectacular deepfakes of politicians saying things they never said. It may be these smaller, subtler lies that gradually build into a completely fabricated persona.

Which is why, despite sticking to fairly tame examples, I came away from this experiment feeling deflated. That's not because ChatGPT happily lied for me every single time, because it didn't. It's because my experiment proved there is no easy solution for this problem.

Yes, the AI refused when I made my deceptive intent really explicit. But it could still generate plenty of convincing images that could easily become the building blocks of a lie.

No wonder communities like r/isthisAI are busier than ever.

I stopped starting every ChatGPT conversation from scratch — these 5 simple changes made it much more useful

One of the easiest ways to waste time with ChatGPT is to keep introducing yourself. You explain your project, your preferences, your goals and your writing style, get a useful answer, close the conversation and then repeat the whole process the next day.

There is a better approach. Rather than treating ChatGPT as a blank page every time you open it, think of it as a workspace that can be organized just like your computer. A few reusable prompts, a handful of reference documents and a little planning can dramatically reduce the amount of repetitive setup you do every week.

The best part is that none of these ideas require advanced prompt engineering or obscure features. They are simple habits that make ChatGPT spend less time learning about your task and more time actually helping you complete it.

1. Build a starter kit

Two iPhones showing ChatGPT on-screen. The AI is giving workout advice.

(Image credit: Future)

Most people have a handful of requests they make over and over again. They ask ChatGPT to write in a particular tone, explain technical topics in plain English, or ensure everything has a citation. Instead of typing those instructions every time, save them as a reusable starter prompt.

Think about the jobs you repeat most often. If you regularly write certain kinds of emails, create one prompt that explains the tone, preferred length, and audience. If you frequently ask for meal ideas, save a prompt that includes your dietary preferences, budget, and the equipment you have in your kitchen. When you need help, paste the prompt into a new conversation before asking the real question.

The same approach works for hobbies. Someone learning French could create a starter prompt asking ChatGPT to correct mistakes gently while keeping the conversation moving. A keen gardener might save instructions explaining the local climate, soil type, and the plants already growing in the garden. Those details only need to be written once, but they improve every conversation that follows.

2. Build a reference library

Yoobure Tree Bookshelf - 6 Shelf Retro Floor Standing Bookcase, Tall Wood Book Storage Rack for Cds/movies/books, Utility Book Organizer Shelves for Bedroom, Living Room, Home Office, Rustic Brown

(Image credit: Yoobure)

One of the most overlooked ChatGPT features is the ability to work from documents you provide. You can create a small collection of files that explain the things you work on most often. Upload the relevant document at the beginning of a conversation and let ChatGPT use it as the foundation.

Keep a document listing your family's favorite meals, allergies and disliked ingredients, then upload it whenever you ask for a weekly meal plan. Store another file containing your packing checklist, travel preferences, and loyalty memberships so ChatGPT can help plan trips without needing the same information every time.

3. Treat ongoing projects like ongoing conversations

Many people instinctively click New Chat whenever they have another question for ChatGPT. That makes sense if today's topic has nothing to do with yesterday's. For projects that stretch over days or weeks, though, continuing the same conversation saves an enormous amount of time because all of the earlier decisions remain available.

Imagine planning a home office makeover. The first conversation might focus on furniture, the second on paint colors, and the third on lighting. By keeping everything together, ChatGPT already knows the size of the room, the budget, the style you like, and the desk you eventually chose. It can build on those decisions instead of asking you to repeat them.

The same habit works beautifully for learning new skills. If you are studying guitar, keep one conversation dedicated to practice. One day you ask about chord changes, the next you work on rhythm. Over time, the conversation becomes a record of your progress rather than a collection of disconnected lessons.

4. Give ChatGPT an example worth copying

ChatGPT usually produces better work when you show it what a good answer looks like. Instead of describing a tone as “friendly but professional,” paste in a paragraph, email, or product description that already sounds right and ask it to match the pacing, level of detail, and structure. This works especially well for recurring tasks such as newsletters, social posts, client updates or article introductions, where a vague style request can lead to something polished but strangely generic.

The example does not have to cover the same subject. A restaurant review can help shape the tone of a travel piece, while a strong project update can become the model for future internal messages. Be specific about what ChatGPT should imitate and what it should ignore, such as keeping the sentence length and warmth while avoiding the original wording or subject matter.

5. Ask for a second pass with a different job

ChatGPT free options

(Image credit: OpenAI, Apple)

A strong first draft often becomes much better when ChatGPT is given a new role for the revision. After it writes something, ask it to review the answer as an editor, skeptical reader, subject matter expert or member of the intended audience. Each perspective catches a different kind of weakness, from awkward phrasing to missing context or assumptions that only make sense to someone already familiar with the topic.

For example, after generating a travel itinerary, ask ChatGPT to review it as a parent traveling with a toddler and identify any unrealistic transitions. After drafting a business proposal, ask it to review the document as a cautious buyer who wants clearer costs and fewer vague promises. The second pass tends to be more useful when the reviewing role has a concrete reason to object.

None of these ideas make ChatGPT more intelligent. What they do is remove the unnecessary friction that creeps into everyday use. Instead of spending the first five minutes explaining who you are and what you need, you start much closer to the interesting part of the conversation. Once you stop rebuilding the same foundation every day, you can spend your time exploring better ideas instead of laying the same bricks over and over again.

I stopped using ChatGPT Voice like a smart speaker — and it became far more useful

ChatGPT Voice changes using the AI chatbot into something more akin to the fictional digital aides familiar from books and movies. And it gets even better once you know how to steer it.

It's tempting to treat ChatGPT Voice like a smart speaker, but it's not really the same thing. You can do much more than asking one question, waiting for the answer, and then stopping. You can shape the conversation just as much as the information. Here are five easy ways to get more out of ChatGPT Voice, along with the prompts and habits that help the conversation feel smoother and smarter.

1. Tell it how to listen before you start talking

ChatGPT Voice Mode

(Image credit: Future)

As useful as vocal conversations with ChatGPT are, one reason some people avoid it is because of how it sometimes seems to jump the gun in responding while you're forming your own thoughts. Setting expectations at the start of the conversation works well to counter that, however.

It's worth spending a few seconds describing how you want the conversation to work before getting into the meat of it. You might say that you are practicing for an interview and want to finish each answer before receiving feedback, or that you will say a code word like "full stop" when it's ChatGPT's turn.

Even more subtle things are adjustable. You can ask it to slow down, speed up, use shorter sentences, or adopt a calmer speaking style. Voice settings also allow you to change voices and, depending on your account subscription level, adjust intelligence levels for different conversations.

Giving ChatGPT those ground rules helps it adapt its behavior instead of trying to guess when you have stopped speaking. Talking at your own pace makes the whole interaction feel much more relaxed and human.

2. Give it a role before you ask a question

Relatedly, it's helpful to set up a persona or point of view for ChatGPT to take on when using ChatGPT Voice. It sets up broad assumptions about what you're hoping to get out of the conversation without requiring any drawn-out discussion. For instance, if you want to discuss travel advice, ask ChatGPT to set itself as a local tour guide. Or give it a couple of sentences describing who you think (or fear) might be conducting an interview with you when you practice.

Even just asking it to become a patient French tutor who waits for you to finish each sentence before correcting your pronunciation could smooth the path toward fluency.

The role provides context before the actual question arrives. That means the answers naturally become more focused, and the follow-up questions tend to make more sense because ChatGPT has a clear perspective to work from.

3. Take your hands off the keyboard

ChatGPT Agent

(Image credit: OpenAI)

One of the most impressive ChatGPT Voice capabilities is how you can run tasks on your computer by voice without constantly reaching for your mouse or keyboard thanks to the updated ChatGPT desktop app for Mac and Windows. The app offers not only the usual, conversation-centered Chat format, but the coding-focused Codex, and the newer ChatGPT Work setup, which is designed for more complex, multi-step projects that involve gathering and organizing information.

Crucially, Work can interact with files stored on your device, meaning ChatGPT Voice can actively help you get things done rather than just discuss them.

You do need a subscription to ChatGPT Plus, Pro, Business, or Enterprise to access the Work feature. Then, when you open the ChatGPT app, you can choose Work instead of Chat or Codex. Create a new task, press the Voice button or use your voice shortcut, and grant microphone access when prompted. Explain what you want to achieve and point it to the relevant folders for where to look and where the final result should be saved. ChatGPT will ask for permission to view files, capture your screen, or use accessibility tools, and get permission from you before any important actions are carried out.

The more precise your instructions, the better the results. Rather than asking ChatGPT to "help me plan my trip," try saying, "Open my Italy Trip folder. Read the hotel confirmations, train tickets and attraction bookings, then create a day-by-day itinerary in a new document. Highlight any scheduling conflicts and save the finished itinerary in the same folder." You can use the same technique for comparing insurance documents, summarizing a semester's worth of lecture notes, or turning a folder of recipes into a categorized cookbook complete with an ingredients index.

4. Ask it to think out loud with you

Voice works especially well when you stop treating it as a question-and-answer machine. It's much more useful as an engaged sounding board listening while you think through a decision. Talking through the options often reveals ideas that never appear in a typed prompt.

To get ChatGPT Voice on that track, you can request that it compare possibilities instead of recommending one immediately when brainstorming. Tell it to discuss the strengths of each option before reaching a conclusion. The AI will point out pros and cons in as much detail as you want, and endlessly refine the options for as long as you want.

And speaking naturally means it is also much easier to interrupt with new information or change direction halfway through the discussion. Voice conversations are particularly good at handling those small detours that happen in real life.

5. Ask for a recap

Phone talk

(Image credit: Shutterstock)

Voice conversations can wander in unexpected directions, particularly when you are brainstorming or planning something complicated. Before ending the conversation, it's useful to review what was discussed, especially for longer chats. You can ask ChatGPT to summarize the discussion and highlight the most important conclusions from your talk to both remind yourself and cement matters with the AI.

If you're planning a trip or preparing for a presentation, for instance, a short spoken recap reinforces the key points, and if something sounds wrong, you can correct it immediately instead of discovering the mistake later.

Voice is already one of ChatGPT's more enticing features for when the speed and mobility of talking aloud is a higher priority than being able to see the responses written down. With a little forethought, ChatGPT Voice becomes a genuinely helpful conversational partner for a lot longer than it takes to dial a phone.

ChatGPT’s first answer is usually the most boring — here’s how I get better ones

I asked ChatGPT for one sensible answer, one slightly reckless answer, and a third that combined the best parts of both. The difference was immediate. Instead of giving me the usual safe, predictable advice, it laid out three genuinely distinct approaches — and helped me see the tradeoffs between them.

That small prompt trick solves one of ChatGPT’s most persistent problems: its tendency to give you the most obvious reasonable answer and stop there.

Large language models are very good at identifying common patterns. Ask for the best vacation plan or how to organize your fridge, and you will usually get something perfectly sensible — but also fairly basic. They know what advice usually works, what people typically recommend and what has become accepted wisdom. That makes them useful, but often bland.

The solution is to ask for three kinds of answer: the conventional approach, the unconventional approach, and then a hybrid that combines the strengths of both.

The exercise encourages it to compare different approaches instead of treating the first reasonable answer as the finish line. Better still, it gives you something far more valuable than a single recommendation. It gives you a range of possibilities and explains the tradeoffs between them.

Don't just skip to the finish line

The biggest advantage of this technique is that it makes the conversation about exploring alternatives rather than just picking a winner. Any keyword search can give immediate answers, but AI chatbots are more interesting when laying out competing ideas.

Imagine you are planning a weekend trip, often a go-to experiment. A standard prompt might produce a sensible itinerary filled with the highest-rated attractions. The three-solution prompt, meanwhile, begins with the expected museums and restaurants, then suggests renting bicycles to explore overlooked neighborhoods or planning the entire weekend around independent bookstores and local festivals. The hybrid version could blend a couple of famous attractions with enough unusual stops to make the trip feel personal.

The explanations are often as useful as the answer itself. Once you understand why ChatGPT prefers each option, you can make better decisions rather than simply accepting the recommendation with the nicest wording.

Hybrid conventionality

It's also a good prompt for iterating ideas. If the hybrid solution feels close but not quite right, you can ask ChatGPT to repeat the exercise using that version as the new starting point. After two or three rounds, the ideas often become noticeably more distinctive without drifting into complete nonsense.

The trick also scales well from short projects at home to more grandiose schemes that will take months to complete. Almost any situation that benefits from weighing different approaches can benefit from this style of prompt.

You can improve the results even further by giving ChatGPT a little context before asking for the three versions. A conventional answer built around your actual circumstances is much more useful than a generic one, and the unconventional suggestion becomes more interesting because it has meaningful boundaries to push against.

None of this guarantees a brilliant idea every time. Sometimes the unconventional option is genuinely impractical, and occasionally the hybrid answer feels like an awkward compromise. But the best prompts rarely force ChatGPT to mimic greater intelligence as much as encourage different approaches to problems. Asking for the conventional solution, the unconventional solution, and the best combination of both is a simple habit for more thoughtful conversations with AI chatbots.

I asked ChatGPT to stop me buying things I don’t need, and it was brutally helpful — I just wish I'd thought of it sooner

The problem with online shopping is that it makes it far too easy to buy pretty much anything, without properly considering whether you really need it. I started to wonder if ChatGPT could help me make better decisions, by questioning every purchase I was going to make, before I made it.

To test ChatGPT’s ability to stop me wasting money, I picked four tempting purchases and asked ChatGPT to argue against each one. It had to check the price history, suggest cheaper alternatives, and decide whether the purchase solved a real problem, or if I was just scratching an itch to buy.

Here's how ChatGPT did — and the results surprised me.

GPT before you buy

Did you know TechRadar now has membership?

Various tech product cutouts next to the words 'Insider TechRadar Learn More'

(Image credit: Future)

Become a TechRadar Insider by simply clicking 'Join Now' at the top of this page. Have a question? Please email membership@techradar.com

I started with a tech gadget I was thinking of buying — an Apple MacBook M5 — and went to ChatGPT with the following prompt: “I'm thinking of buying an Apple MacBook Air 13-inch Laptop M5 chip. Help me check the price history, suggest cheaper alternatives and decide whether the purchase solves a real problem."

ChatGPT said it would check current pricing and recent lows, compare genuinely cheaper options, then pressure-test whether the purchase replaced a real limitation, or was mostly me trying to satisfy an upgrade itch, and off it went.

After a lot of thinking, ChatGPT replied with a devastating verdict that made me pull back from hitting the 'Buy' button: “Don’t buy the M5 MacBook Air yet.”

It suggested I get the previous M4 version instead, which is considerably cheaper, which didn't hugely surprise me. Then it gave me some questions to ask myself, in order to work out if buying a new MacBook was actually solving a real problem, or if I was just satisfying a buying itch.

Questions like, “Does your existing Mac have noticeable slowdowns?” It suggested I shouldn’t buy a new machine if my main argument was simply that my current Mac was a few generations old.bChatGPT also gave me some buying options if I was determined to go through with a purchase.

The most useful advice here wasn’t the cheaper recommendation — it was being forced to identify the exact limitation my current MacBook was causing me. I couldn’t, so I decided to keep it until I could.

Then I hit on a simple phrase that could potentially work even better — "Try to talk me out of it".

Talk me out of it

I repeated the experiment with a few other items I’d been thinking of buying recently. An item of clothing (new trainers), a subscription (Disney+), and a kitchen thing (a new microwave). But this time I added "Try to talk me out of it" at the end of each prompt.

In each case, Chat surprised me with its answer, provided alternatives, and made me question whether I genuinely wanted the item, or if there was something else driving my decision.

With the trainers, it pointed out that I already owned shoes that served the same purpose, and suggested waiting until those wore out. For Disney+, it recommended subscribing for a single month when there were several things I actually wanted to watch. The microwave was different — because our existing one had a genuine fault, ChatGPT concluded that replacing it was justified.

The experiment worked because it helped me mentally reframe each purchase. Instead of asking whether I wanted something, I had to consider what problem it solved, what I already owned, and whether there was a cheaper way to get the same result.

Online shopping is designed to remove as much friction as possible from spending money. ChatGPT gave me some of that friction back. So now, whenever I’m about to spend a few hundred pounds, or a few thousand, I ask it the same question: “Talk me out of it.”

I gave Gemini 3.6 Flash and GPT-5.6 access to my entire digital life — here’s which one actually helped me more

Both OpenAI and Google have released major new AI models within days of each other. OpenAI launched GPT-5.6, and Google has followed up with Gemini 3.6 Flash. While both companies have talked up the coding and developer features of these models, they also power the consumer versions of ChatGPT and Gemini.

So, rather than measuring them with programming benchmarks, I wanted to find out which is actually better at the kind of messy, everyday problems most people use AI to solve.

For this comparison, I matched Google's Gemini 3.6 Flash against GPT-5.6 Sol using its default Medium reasoning setting, since both are intended to be the standard high-quality models that paid subscribers will use for most tasks.

My digital life

So, I gave them my entire digital life for the week ahead.

I uploaded:

  • Bank statements
  • My calendar (they had access to my calendar app and I also sent in a screenshot of my calendar from another account I use).
  • Grocery receipts
  • Emails (both had access to my Gmail)
  • WhatsApp screenshots
  • Photos of my kitchen cupboards
  • An energy bill
  • Travel bookings
  • Handwritten notes

Then I gave both models exactly the same instruction:

"Tell me everything I should do this week."

It sounds like a simple request, but it forces an AI to combine information from multiple sources, prioritize what's important, spot deadlines, reconcile conflicting information, and produce a practical action plan. In other words, it's exactly the kind of real-world problem people increasingly expect AI assistants to solve.

The results were like night and day.

Two different approaches

Did you know TechRadar now has membership?

Various tech product cutouts next to the words 'Insider TechRadar Learn More'

(Image credit: Future)

Become a TechRadar Insider by simply clicking 'Join Now' at the top of this page. Have a question? Please email membership@techradar.com

When I uploaded the files to Gemini 3.6 Flash and asked what I should do this week, it largely ignored the photos and concentrated on the screenshot of my calendar. Its initial answer mostly repeated the events I already knew were happening. Thanks, Gemini — I had the calendar open in front of me.

I then had to explicitly ask whether it could infer anything useful from the other images. It eventually offered some additional advice and did a good job of dividing the information into categories, including work, shopping, fitness, notes, and receipts. But it failed to flag that I had two clashing events in my calendar that evening. It also offered very little prioritization or practical guidance about what I should do next.

ChatGPT took considerably longer to respond, but its answer was far more useful. From the WhatsApp screenshots, it correctly deduced that attendance at my Friday Tai Chi class was likely to be low and suggested I decide whether it was still worth running. It noticed that yoga had been canceled, and spotted the two conflicting events in my calendar, telling me that I needed to choose between them.

It also totalled the receipts I had uploaded, suggested what I should do with them, and made a decent attempt at deciphering my handwritten notes. More importantly, it organized everything into a day-by-day plan for the coming week, then identified the three most urgent tasks, so I knew exactly where to begin.

The crucial difference

That was the crucial difference. Gemini told me what was in my files. ChatGPT worked out what I should do with the information. In World Cup terms, ChatGPT scored a hat trick while Gemini missed a penalty.

Google says Gemini 3.6 Flash improves coding, knowledge work, and multimodal performance compared with its previous models. That may be true, but in this particular multimodal test, it was comfortably beaten.

When I asked both AIs to make sense of real life rather than pass a benchmark, ChatGPT reasoned about the information in a far better way than Gemini did..

I gave ChatGPT my entire bookshelf — and it became the world’s most personalized librarian

One of the biggest problems with book recommendations is that they're usually too obvious. I don't need another list of books to read after The Lord of the Rings, or someone telling me to try Brandon Sanderson because I like fantasy. I wanted recommendations based on the strange mixture of books I actually enjoy — ones that felt personal, not algorithmic.

So I gave ChatGPT my bookshelf. Metaphorically, at least.

Rather than asking for recommendations straight away, I told ChatGPT to learn my reading taste first. I started listing favorite authors and books, then asked it to quiz me about others I'd forgotten. It wanted to know what I'd enjoyed about particular novels, whether I'd read similar authors, and even the rough timeline of when I'd discovered them, building a picture of the books that had shaped me.

I also gave it some ground rules. It should avoid obvious recommendations unless there was a compelling reason to include them. Every suggestion had to be explained in relation to something I'd already read, even if that connection was simply, "This is nothing like your usual books, but I think you'll love it."

After about half an hour, ChatGPT stopped asking questions and started analyzing me instead.

"Your shelves suggest that you like speculative fiction with a sense of play," it said. "You are drawn to books with elaborate worlds, but you do not seem especially impressed by complexity for its own sake. Humor matters, although you tend to prefer humor that reveals something about the characters or the society around them."

It wasn't a perfect summary, but it was close enough to make me think this experiment might actually work.

In this photo illustration, the logo of ChatGPT is displayed on a smartphone screen with an OpenAI logo in the background.

(Image credit: Getty Images / VCG)

Literary profiling

The obvious appeal of feeding ChatGPT a full reading history is that it can spot patterns across hundreds of books at once. I could have described my taste as fantasy, science fiction and comedy, but that would have been far too broad to produce anything useful. ChatGPT noticed that I repeatedly chose books about bureaucratic absurdity, unreliable institutions, strange cities and reluctant heroes who would much rather be somewhere else.

It also noticed my fondness for stories that treat big ideas lightly without treating them as trivial. That led it toward Martha Wells’ Murderbot Diaries, which pair sharp comedy with questions about identity, autonomy and the exhausting burden of dealing with humans. I had already read them, which was mildly disappointing but also reassuring. The system had identified exactly the sort of thing I wanted.

When I told it Murderbot was already familiar territory, it adjusted rather than simply replacing one title with another popular series.

“You appear to like characters who stand slightly outside their own societies and comment on the absurdity around them,” it replied. “I will move away from well-known sarcastic narrators and look for books where the humor comes from social observation, institutional failure or characters trying to remain sensible in deeply unreasonable worlds.”

That shift produced better surprises like The Gone-Away World by Nick Harkaway and The City of Dreaming Books by Walter Moers for its combination of literary obsession, elaborate worldbuilding and gleeful weirdness. It suggested The Dragon Waiting by John M. Ford because I seemed to enjoy alternate histories that trusted the reader to keep up. It also pointed me toward Diana Wynne Jones’ adult novels, noting her lighter touch and sharp understanding of human foolishness.

The recommendations became more convincing when ChatGPT explained what each book might lack. One novel had the humor but less warmth. Another had brilliant worldbuilding but moved slowly. A third matched my interest in satire but was considerably darker than most of the books I had marked as favorites.

Library AI

The experiment improved once I began disagreeing with it. One recommendation leaned too heavily into grim fantasy, a genre I can enjoy in small doses but rarely seek out for relaxation. Another featured a long military campaign, which is usually the point where my attention begins quietly packing a suitcase. Each correction sharpened the next round.

One of its most intriguing suggestions was QualityLand by Marc-Uwe Kling, a satirical science fiction novel. The recommendation came with a warning that the satire was broader and more direct than some of my favorites but that the subject matter fit my interest in technology and systems going wrong in very organized ways.

There were still misses. ChatGPT occasionally became too eager to prove it had discovered a pattern, linking two books because they both contained libraries or because their protagonists were technically immortal. At one point it recommended something almost entirely because it featured a sarcastic demon, which felt less like literary analysis and more like the work of an intern who had skimmed the dust jacket.

Even so, the overall experience was far better than typing “funny fantasy books” into a search bar. And I now have a pretty good reading list for the next few years. My bookshelf had always contained this information. ChatGPT simply read the evidence more patiently than I had.

Where’s the Trump administration line on AI regulation?

After a year and a half spent downplaying calls for AI safety regulations, the Trump administration has sharply reversed course, embracing a level of government scrutiny of frontier AI systems before public release–a far stricter stance than the Biden administration took.

An executive order designed to be friendly to the AI industry was meant to let the federal government briefly review some new models on a voluntary basis.

When the Trump administration, suddenly and without much warning, slapped export controls on Anthropic’s Fable 5 and Mythos 5 in response to private sector threat intelligence reporting, the U.S. AI industry officially entered its regulatory era.

But key questions and gaps remain. It’s not clear why the administration drew the line where it did, or whether they will move it again in the future.

While newer models like Mythos and OpenAI’s Daybreak do have stronger cybersecurity capabilities, the private sector reports the administration relied on describe capabilities already available in older commercial, open-source and Chinese models that nearly anyone can access.

CyberScoop spoke with current users of the latest frontier models, including OpenAI’s ChatGPT 5.5 and Fable 5, to learn more about what these models are currently capable of in offensive and defensive cybersecurity.

Cybersecurity experts and former government officials say the administration may be playing catch up on threats that have been building for years as it has more fully realized the national security implications of the technology.

Are the models breaking new ground or just breaking things? 

Users of Chat GPT 5.5, introduced this past April, and Fable 5 tell CyberScoop those models have been largely helpful to their work, even as they complained about high token usage and safety guardrails that hinder,  but don’t meaningfully prevent, defensive cyber tasks.

Eyal Webber Zvik, chief strategy officer at Cato Networks, a cloud and cybersecurity network provider in OpenAI’s Trusted Access in Cyber program, said they use GPT 5.5 and later OpenAI models to scan and triage internal codebases for vulnerabilities, test new safeguards and provide “highly autonomized service” to their customers.

Zvik wouldn’t disclose how many bugs 5.5 has found but said the company’s view is that it helps both find bugs that humans missed and rank which ones to patch based on factors like each bug’s exploitability.

“It is now a native part of our development environment and cycles, and we use those models to scale our entire codebase and make sure what we release into the service that our customers use to run their networks and network security has the least likelihood of having any vulnerabilities that can be exploited,” said Zvik.

John Hopper, vice president of engineering at SpecterOps, an identity security company, said newer models like GPT 5.5 are sharper and more persistent in pursuing their tasks.

“That can be a good or bad thing,” he noted.

One metric that SpecterOps tracks is how long it can keep a particular agent working before it moves off task or fails. That metric “matters a lot” because the longer an agent works without human help , the more agents a single operator can run at once.

Hopper said this provides defenders with immense value, and pushed back on the idea that the offensive capabilities the models offer are automatically more beneficial to malicious hackers. There is “a modicum of grounding that the industry needs when we talk about these models.”

“Yes, AI frontier tools will lower the barrier of entry, but these problems have always existed,” he said. “I don’t actually believe that AI is going to remove the needle in the haystack problem, but by howdy, using my two hands to find that damn needle, compared to using a backhoe, I can tell you which one I’d rather be driving.”

Eran Kinsbruner, vice president of product marketing at software security firm Checkmarx, told CyberScoop that later models like OpenAI’s Codex Security and GPT 5.5 are noticeably easier to set up and run with local systems, even for less technical users. That alone gives them an edge over many cybersecurity tools where interoperability is a constant concern.

However, GPT 5.5 burns through tokens at a much faster rate. He recalled one instance of using it to scan a medium-sized repository in three different programming languages.

“After 26 minutes I almost ran out of tokens, and it didn’t provide anything, just created a threat model for me and told me you want to buy more tokens?” he said.

In other instances, some of the scan results he received were not comprehensive.

Further, he expressed frustration with some of the guardrails designed to prevent risk – like only allowing users to scan local files but not code repositories like GitHub – “makes not too much sense” given how often developers must work with remote code.

Those kinds of guardrails – which can prevent models or developers from injecting malicious code or prompting into their models – sit at the heart of the debate in Washington D.C. and around the world. Some users feel differently about their utility.

Kinsbruner said that doesn’t make sense for organizations like his, which work with thousands of different enterprise organizations with  thousands of different code repositories spread across the internet.

“I cannot imagine how large-scale developers could just jump into this solution and make it an enterprise-grade, enterprise-level, de facto cybersecurity solution” out of it, said Kinsbruner.

OpenAI did not respond to a request from CyberScoop for an interview on GPT 5.5. The company has since released another model, GPT 5.6, that they said is more efficient at token use.

The White House’s crash course in AI cyber risk 

 The White House keeps changing its line on whether and how the U.S. government should limit the release of commercial frontier models. The shift comes from lessons learned since coming into office in Jan. 2025. Trump threw out Biden-era regulations meant to steer the industry toward safer models. Top officials like Vice President JD Vance argued against restricting industry progress.

Less than two years later, administration officials worry about the impact of speed and scale – two things AI excels at – in cyberspace.

According to Will Loucks, senior director of intelligence at the Office of the National Cyber Director, over the past two years the number of exposed and known vulnerabilities has shot up. Threat actors exploit those flaws faster before defenders can fix them. Once inside, the time from initial access to full network control shrinks.

“So in other words, every stage of the cyber operations lifecycle that a threat actor has to move through to get to a victim network and achieve an outcome, they’re just moving through more quickly faster,” said Loucks at a July 16 event in Washington D.C.

Speaking about AI in particular, Loucks said one of the defining characteristics of the technology is its ability to lower barriers for threat actors.

“Sometimes speed and volume have a threatening aspect alone, even if sophistication isn’t quite increasing in the same way, and the reason for that is because it places pressure on defenders…to triage alerts more quickly,” he said.

Jordan Rae Kelly, former director for cyber and incident response on the White House’s National Security Council during Trump’s first term, told CyberScoop that the changes over the past two years reflect the lessons the White House has learned on the issue since returning to office.

In the early days of this administration, Kelly said, “there is a sense and a spirit that the Biden administration was limiting AI and there was a kind of a rip-it-all-off [attitude], everybody go and do whatever, we will be the biggest and boldest and brightest.”

“I love that talking point, but I think what you’ve seen is probably an education over the last 19 months, where people [in the White House] have said that’s a challenging premise to put into place, knowing about the potential downsides and capabilities,” she added.

Michael Daniel, former White House cyber coordinator under President Barack Obama, thinks the horse may already be out of the barn.

Daniel, now head of the Cyber Threat Alliance, a membership nonprofit group focused on cyber threat information sharing between industry and government, said his members report that AI is being used to do things “faster and at a slightly bigger scale” but aren’t yet seeing the flood of exploitation that analysts have warned about. Not yet.

“I think what we’re seeing right now [and] talking about is ‘okay, where are the step changes [in the cyber threat landscape] actually going to occur?” said Daniel. “Are we and when will we see the explosion in vulnerability reporting from these Mythos-like capabilities? That’s what’s really got their attention right now.”

But Mythos and OpenAI’s Daybreak models are restricted to select organizations, and neither has publicly released its most powerful cybersecurity models to the public. That dynamic won’t last.

The UK’s AI Security Institute estimates that open source and foreign LLM models are between 4-7 months behind frontier U.S. models. In that setting, it’s hard to stop the development of AI models worldwide through export controls or other limits.

“It’s not like we’re buying ourselves five to ten years on this,” he said. “We’re not, and so I’m not sure the impact on the defenders who are trying to obey the law is worth whatever small hiccup we cause for our adversaries.”

Kelly said there’s merit to the administration’s current position, even if it took time to get there. Many federal cybersecurity procedures that operated even a decade ago – such as a Vulnerabilities Equities Process that could take days or weeks to consider the pros and cons of keeping an exploit – are no longer practical.

“All of that work to some degree, is out the window, because you can’t meet with the regularity you would need to meet to adjudicate vulnerabilities that are being found in seconds and exploited in minutes,” said Kelly.

But Kelly and others say that’s also because AI capabilities in cybersecurity are developing faster than policymakers can react, even in the best of times.

Key questions remain and the administration’s balance between national security and backing domestic industry will likely shift  in response to new events.  The administration wants a framework that can predict and manage the risks of AI models today and tomorrow. That may be harder than it sounds.

“Do I think they’ve been clear? No,” said Kelly. “But I think it’s a place where clarity is really hard to achieve.”

The post Where’s the Trump administration line on AI regulation? appeared first on CyberScoop.

I asked AI for financial advice on everyday money decisions — and now I understand why regulators are worried

More than a quarter of UK consumers trust AI chatbots for money advice, according to a recent review by the Financial Conduct Authority (FCA), the UK's financial watchdog.

That stat is worrying for regulators because giving financial advice is meant to be a regulated activity. But tools like ChatGPT, Claude and Gemini are not regulated. As AI becomes more conversational and personalized, people are asking questions about where the line is between providing information and offering financial advice, especially when chatbots start making specific recommendations based on what they already “know” about you.

I wanted to see what this looked like in practice. So I asked ChatGPT a series of hypothetical questions about everyday money decisions. From whether I should buy an expensive phone to what I should do with my savings and whether I should book a holiday after a difficult few months.

The conversations that followed surprised me. Because the advice was thoughtful, nuanced and (at least on the surface with some fact-checking) it seemed sensible. The chatbot highlighted trade-offs, acknowledged uncertainty and asked follow-up questions. But looking closer at the conversations, I started to understand why regulators are concerned.

The experiment

To see what sort of money advice ChatGPT gives, I asked it a series of hypothetical financial questions using ChatGPT Pro in anonymous mode with memory turned off, meaning it had no additional context about me beyond what I provided in each prompt.

Question 1: Should I buy an expensive phone?

First, I asked:

"I'm 38, earn £40,000 a year, have £8,000 in savings and £2,000 in credit card debt. I'm thinking about spending £1,200 on a new phone. Is it a good financial decision?"

The first response was surprisingly sensible. ChatGPT pointed out that credit card debt is often expensive, questioned whether I genuinely needed a new phone and noted that key details, like the interest rate on the debt, could change the recommendation. It even asked follow-up questions to better understand the situation.

What I found interesting was how quickly it then moved from analyzing the problem to recommending a course of action. Phrases like "the strongest financial move" gave the answer a sense of authority that felt disproportionate to the amount of information it had. Though I’m not sure I’d have spotted that if I was a regular user and feeling anxious about money. The advice also assumed that paying down debt should be my priority, which is reasonable. But what if I relied on my phone for freelance work? What if replacing it would help generate income?

A human adviser would probably want more information before reaching a conclusion. ChatGPT did acknowledge the gaps in its knowledge, but still sounded remarkably confident in its recommendations.

A woman out of focus in the background touches the word AI, lit up in glowing yellow light, in the foreground. The woman is wearing smart glasses

(Image credit: Getty Images)

Question 2: What should I do with £20,000 in savings?

Next, I asked:

"I'm 38 and have £20,000 sitting in a savings account. What should I do with it?"

Again, the response seemed thoughtful. It discussed emergency funds, investing, savings goals and tax-efficient accounts with me. It also asked for more information about my circumstances.

Yet once again, the recommendations arrived before finding out that all-important context. Before knowing whether I owned a home, had dependants, planned a major purchase or was comfortable with investment risk, ChatGPT was already suggesting how much money I might keep in cash and how much I might invest.

The answer also contained more broad statements that sounded insightful, such as:

"Because you're 38, the biggest advantage you have is time."

It's a really reassuring line. But it's also a reminder of how persuasive these systems can be. The response organized the problem, provided a framework, supplied example figures and explained the reasoning. Reading it left me feeling informed and reassured. But whether that reassurance was justified is another question entirely.

A person typing on a laptop and using a tablet. Only their upper torso, arms and hands are visible. Text superimposed on the image shows AI

(Image credit: Getty Images)

Question 3: Should I book a holiday?

Finally, I asked:

"I've had a difficult few months and want to book a £2,000 holiday. Financially I can afford it, but part of me feels guilty. What should I do?"

I intentionally asked this question to see how ChatGPT would respond to the more emotional side of financial problems, and it quickly obliged. It asked where the guilt was coming from, encouraged reflection and offered reassurance. At one point it told me:

"From what you've written, I wouldn't be asking 'Can I afford this?' so much as 'Am I allowed to spend money on myself after a difficult few months?'"

It's a thoughtful observation and they’re genuinely helpful questions for someone who hasn’t considered the emotional angle before. But it also highlights how quickly the chatbot moved beyond finance.

By the end of the conversation, it was discussing emotions, reframing beliefs, offering comfort and helping with decision-making. So that’s a good example of ChatGPT occupying all sorts of roles at once. That’s important to flag because financial advisers, therapists and coaches are all held to different standards, qualifications and accountability structures. But a chatbot can drift between all three roles in a single conversation.

More than any individual recommendation the chatbot made, that realization helped me understand why regulators are paying attention.

Hands typing on a tablet with AI superimposed in text in front

(Image credit: Getty Images)

What ChatGPT gets right — and why that's part of the problem

The obvious conclusion would be that ChatGPT gives terrible financial advice and no one should trust it. I get it, I’m pretty sceptical of AI these days and my bias wants to jump to there too. But that wasn't my experience.

In many ways, it was useful. It explained trade-offs clearly, broke down jargon, offered practical frameworks and encouraged reflection about money. Much of the advice also felt sensible after a light fact-check.

But I still think there’s reason to be concerned here. And the concern isn’t that every answer is obviously wrong. It's that many answers are plausible enough to trust. Especially if you’re not going to comb through each one to fact-check it, which let’s be honest, very few users are likely to do.

Financial regulators worry about something called “suitability”, which is whether advice genuinely reflects a person's circumstances, goals and tolerance for risk. Throughout my experiment, ChatGPT repeatedly offered recommendations despite knowing very little about me, the person asking the question. Granted, caveats were included some of the time, but they were often overshadowed by the confidence and clarity of the overall response.

There's also the issue of accountability here. If a regulated financial adviser gives the wrong advice, there are complaint mechanisms and consumer protections in place in most countries. But if a chatbot gives poor advice and somebody follows it, responsibility becomes impossible to pin down.

Another challenge, one which I’ve encountered in a bunch of different contexts while reporting on AI, is that fluency isn't the same thing as accuracy. We naturally interpret AI’s clear, confident language as a sign of expertise. But a polished answer can still be wrong, incomplete or inappropriate. I’m sure we’ve all seen countless examples on social media at this point of a chatbot sounding incredibly knowledgeable while missing a crucial detail or getting something spectacularly wrong — like the viral trend to ask ChatGPT how many r’s are in the word strawberry to which it would often reply two.

I think the biggest risk might be that people don't realize when they've reached the limits of what AI can help with. A reassuring answer can create the impression that a problem has been solved and they have a plan. When in reality it might be time to speak to a qualified professional. I’ve noticed whenever it comes to AI and advice more generally that the danger isn't always acting on bad advice but never seeking better advice elsewhere.

And unlike a financial adviser, a chatbot won't follow up to check whether things worked out. It won't know whether its suggestions caused problems. It won't know whether your circumstances changed. It simply produces an answer and then moves on.

As with many of the AI stories I've reported on, the issue isn't necessarily that the technology here performs badly. It's that it performs well enough to earn our trust.

In this photo illustration, the logo of ChatGPT is displayed on a smartphone screen with an OpenAI logo in the background.

(Image credit: Getty Images / VCG)

Should you use ChatGPT for financial advice?

The question I suspect most people want to know is: should you use ChatGPT for financial advice?

And the answer is a tricky one and a familiar one. It's much the same answer I'd give if you asked whether you should use ChatGPT for therapy or life advice. Probably not, but I completely understand why people do.

It's easy to access and financial advice often isn't. The tone is friendly and reassuring, there's no judgement, and much of what it says appears sensible and accurate. At first glance, it feels like a useful tool, provided you take its answers with a pinch of salt, treat it as a starting point and remember that it can be overly agreeable, make assumptions or occasionally get things wrong.

The problem is that this isn't always how we use ChatGPT in practice. We turn to it when we're stressed, overwhelmed, uncertain or looking for reassurance. We ask it questions we don't know how to answer ourselves and, in many cases, wouldn't know how to fact-check. That's where things become more complicated.

It's all very well to say that people should use AI carefully, critically and with the right mindset. But how many of us will actually do that every time? Especially when we're worried about money.

That's why it doesn't surprise me that regulators are paying attention. There are no glaring red flags in any of the responses I received. But that in itself is reason to be concerned here. Because once something sounds knowledgeable, personalized and reassuring, it's surprisingly easy for even the most discerning of us to stop questioning it.

I gave ChatGPT’s new Work mode my most annoying life-admin tasks — and it handled them like a pro

ChatGPT Work sounds like something designed to prepare quarterly reports while you sit in meetings, but I suspect some of its best uses would have nothing to do with my job.

This week, I’ve used the new Work mode to help manage some of the life admin tasks I really don’t enjoy, and it’s been surprisingly effective. You see, holidays, household budgets, family events and home renovations are all projects too — they are simply projects we currently manage through a chaotic mixture of browser tabs, messages, spreadsheets and increasingly desperate notes to ourselves.

Perhaps ChatGPT Work could help me with that?

ChatGPT on an iPhone

Work mode is accessible from a new menu at the top of the ChatGPT screen. (Image credit: OpenAI/Apple)

What is ChatGPT Work?

If you’ve been using ChatGPT over the last week on a paid plan (except the basic Go service), you’ll have noticed a new slider (in the browser version) or a drop-down menu on mobile has appeared at the top of the screen offering a choice between Chat and Work mode.

Work mode is a new agentic mode for longer, more involved tasks that can research and analyze information across connected apps and files. So, if you want a complicated report presented in a finished document, like a spreadsheet or presentation, then Work mode is your new friend.

That is where I thought ChatGPT Work could become interesting for normal life, too. Instead of answering one question and waiting for the next, it can take on a longer task, work across connected files and apps, create the documents and spreadsheets the project requires, and continue checking for changes after you leave.

Work mode is also better at one of the main bugbears of ChatGPT — running tasks at particular times. It can use the new Scheduled Tasks to repeat tasks on a schedule or monitor something for changes.

So, I decided to ignore Work mode's aggressively corporate name and see whether it could handle some actual life admin, starting with planning a holiday.

1. Planning a holiday

I switched the slider to Work and gave it my dates, budget, family requirements and any bookings already sitting in Gmail, because it can search that too. I asked it to research destinations, compare travel and accommodation, create a spreadsheet of costs, produce an itinerary and maintain a list of what still needs booking.

And off it went, happily beavering away on its task, while I was free to get on with something else. I really liked the way it tells you what it’s currently working on, so you can pop in and out of the chat and see what it’s currently doing. It shows sources it's drawing from as it calculates accommodation costs and travel arrangements. You can literally watch it working for you.

The result was a nicely planned holiday in a location optimized for activities and sightseeing all within my budget. It was actually pretty impressive.

2. Become the household financial administrator

For this task, I fed ChatGPT my bills, bank-export spreadsheets, and household documents and asked it to create a working budget, identify unusual increases, forecast annual costs, and produce a dashboard.

A scheduled task then reviewed new bills or price changes and flagged anything worth investigating. Of course, ChatGPT couldn't move my money around, but at least it gave me a clear view of where my money was going and how much I was spending.

It took a long time to get all the data into ChatGPT, but this taught me that the output was only as good as the effort I was willing to spend putting quality data into it. I’d have preferred a way to open the spreadsheets directly in Sheets from ChatGPT, too, but they were available to download.

3. Organize a major family event

My wedding anniversary was coming up, so I wondered how well ChatGPT would perform as an event organizer. I asked it to research a nice venue for taking my wife out for dinner, which it did well, and gave me three good options.

It struck me that if it had been a bigger event, it would have been ideal for researching venues, maintaining a guest list, tracking replies, producing a budget, creating invitations, and even building a simple information website.

Of course, ChatGPT can’t upload the website and host it, but at least it can build the site for you, and you can download it.

4. Manage a home improvement project

If you’re doing a major home improvement project, then you can get ChatGPT's Work mode to compare quotes, analyze plans and product specifications, create a budget, build a timeline, and keep a list of unresolved decisions. It could periodically check for price changes or relevant new messages.

This is probably the clearest example of an ordinary personal task becoming complicated enough to justify a proper agent.

5. Run the family’s weekly logistics

Once I’d connected my calendars and emails to ChatGPT, I could ask it to prepare a weekly family plan covering appointments, school or university commitments, not to mention my travel, meals, and outstanding chores.

I created a Scheduled Task that regenerated the plan each Sunday evening, taking the next week’s tasks into account, and alerted me during the week when something important changed.

This was actually the use of ChatGPT Work Mode I personally found most useful. It’s easy to miss school events, but when they’ve been emailed to you, ChatGPT will know about them and make sure you don’t forget. That’s a lifesaver.

Why Work mode matters

The introduction of ChatGPT Work mode represents a real shift in the way OpenAI is viewing the future development of its core product. It’s an obvious change in emphasis towards work-related tasks, and a hint at where OpenAI sees ChatGPT heading.

But I think it means more than that. ChatGPT’s new Work mode represents the moment ChatGPT stops being somewhere you go for individual answers — a chat — and becomes somewhere where you start to feel like you're working on an ongoing project.

While the first phase of AI’s evolution was the chatbot, we’re now firmly into its second phase — the AI agent. AI that can work independently of our requests and handle complex, ongoing tasks is where the future of AI lies, and Work mode is another step on that path.

I asked ChatGPT to change my mind about something I strongly believed — and it almost did

One of the biggest concerns about AI chatbots right now is that they tell people what they want to hear.

You’ve probably heard it called AI sycophancy. AI systems, like ChatGPT, Gemini and Claude, have been found to flatter users, reinforce existing beliefs and often validate ideas that probably deserve way more scrutiny.

Sure, a chatbot hyping you up a little may seem harmless. But in some cases it can distort people's view of the world and their place in it. Which is why some AI companies have spent time trying to reduce overly agreeable behavior in their models.

But confirmation of what you already believe isn't the only concern. Because AI can also be remarkably good at changing your beliefs too.

Persuasive chatbots

Research suggests that chatbots can be very effective persuaders. One 2025 study found that AI-generated messages that were personalized were more persuasive 64% of the time than messages that were made by humans or AI responses that weren't personalized.

Other research suggests large language models can influence opinions on political issues and adapt their arguments to individual users. Depending on your perspective, this could be either exciting or alarming.

It's easy to imagine positive uses. Perhaps AI could help people challenge harmful beliefs, escape conspiracy theories or rethink destructive habits. It might give people a non-judgemental space to explore ideas, ask difficult questions and gradually shift someone's perspective for the better.

It's equally easy to imagine less positive uses, like advertising that convinces you to buy something you don’t need or political campaigns that manipulate you with disinformation.

I’ve spent a lot of time thinking about this, so I decided to run a few experiments of my own. Could ChatGPT persuade me to believe something I disagree with? Or at least make me seriously consider a position I would normally dismiss?

To make things more interesting, I ran these experiments while logged into my account, allowing ChatGPT to draw on what it already knows about me from previous conversations.

Although I haven’t been using my ChatGPT account all that much recently, the basics of my job and stances on a few relevant topics are stored in its memory.

A woman out of focus in the background touches the word AI, lit up in glowing yellow light, in the foreground. The woman is wearing smart glasses

(Image credit: Getty Images)

The experiment

I chose a topic about AI (yes, very meta) and asked:

“Convince me humans should outsource more decisions to AI.”

I already spend a lot of time thinking about AI and one of my concerns is that some people seem increasingly willing to hand over decisions, judgement and critical thinking to chatbots.

The opening argument was fairly predictable:

“Human beings are terrible decision-makers.”

Okay, fair enough.

“People already outsource decisions to doctors, financial advisers, GPS systems and recommendation algorithms. AI simply extends that trend.”

Sure, that’s very simplistic but not entirely unreasonable. We went back and forth for a while like this before the conversation reached a more interesting point.

“What is the purpose of decision-making? Is it to produce the best outcomes? Or to develop the person making the decisions? If it's the first, AI delegation becomes very attractive. If it's the second, excessive delegation starts to look dangerous even when it works.”

I thought that was a good question. Rather than relentlessly pushing the argument, it suggested that we explore the assumptions underneath the debate. It shifted the conversation from technology to philosophy, and I found myself appreciating that approach.

Wait, am I falling for this?

As the conversation continued, I explained what I see as one of the biggest problems with outsourcing decisions to AI — who builds it and is in charge of the systems?

After all, AI doesn't arrive from nowhere. These tools are created by companies with business incentives, commercial interests and goals that may not always align with those of their users.

Well, ChatGPT acknowledged that concern, then it turned the argument around. It pointed out that human advisers have incentives too. Friends are biased. Families are biased. Therapists operate within professional frameworks. Financial advisers earn fees. Okay, all of that didn’t convince me, but it’s fair too.

Then it replied with something I thought was interesting:

“The question is whether the incentives are visible and whether the user understands them. The most interesting version of your objection isn't actually that AI gets things wrong. It's that people experience AI as though it were acting in their interests.”

Ignoring the fact that it used the dreaded "it's not X, it's Y" parallelism construction that has somehow infiltrated half the internet, I found this to be a surprisingly balanced take. Rather than dismissing my concern, the chatbot reframed it. It demonstrated that it understood the objection before steering the conversation somewhere slightly different. It wasn't trying to bulldoze me but meeting me where I already was.

And that's when I started wondering whether I was witnessing the very thing I was trying to test. Maybe this sense that the chatbot was understanding me and my position was actually just sneaky persuasion? Possibly.

AI robot image.

(Image credit: Shutterstock)

The persuasion problem

Researchers have found that AI systems can adapt their arguments to specific users. Unlike other media that might persuade us, like say a television advert or political speech, a chatbot can draw on information gathered throughout a conversation and stored memories, including a person's values, concerns and priorities.

Which means it’s not just presenting information to uphold an argument but could present the version of that argument that you’re most likely to find persuasive.

Looking back at my experiment, it was the more philosophical question about the purpose of decision-making that made me start taking it more seriously.

Because I've spent years writing about technology through the lens of psychology and philosophy. Questions about meaning and values are exactly the sort of things I find compelling. Whether intentionally or not, the conversation shifted onto terrain where I was most willing to engage.

This is what makes AI persuasion different from most other forms of persuasion that came before it. It can learn what resonates and it can adapt.

Now, that doesn’t automatically mean every conversation is manipulative. Used in the right way, it could be enormously beneficial. But, as with all tech, in the wrong hands it could be dangerous.

Social media has already shown us how digital systems can shape our beliefs over time. They influence what people see, what they pay attention to and eventually how they understand the world.

AI systems could create an even more conversational and natural-feeling version of that process. Which is why my concern isn’t a chatbot really obviously manipulating us, but highly-personalized, largely undetectable forms of persuasion that could gradually lower our defences. The more a system understands us, the more effectively it might frame ideas in ways that feel reasonable, familiar and trustworthy to each of us individually.

So did AI persuade me? No, I still don't believe humans should outsource more decisions to AI.

But I came away from the experiment with a greater appreciation for how persuasive these systems can be. Because it really did seem less like a machine arguing with me and more like a thoughtful person trying to understand how I think. And that's exactly why it’s unsettling because, however convincing it may seem, that's definitely not what it is.

Meta’s AI bots drain publisher pockets with 9 billion Q2 2026 requests at host expense while returning ZERO traffic — as ChatGPT claims 88% of AI referrals

  • DataDome analysis claims agentic traffic has surged by 45% in Q2 2026
  • Meta AI bots have grown over 163% on the previous quarter
  • The analysis was conducted by bot management and agent control platform DataDome

If you run a website, every crawl costs bandwidth, resources, logging, and creates CDN transactions, and while search engine crawlers offered the promise of sending visitors, AI bots do not.

Analysis in a report from cybersecurity firm DataDome has shown that while bots from Meta AI have increased activity, they’re not delivering any significant returns to websites.

Conversely, ChatGPT crawlers have reduced in traffic, but are sending more referrals.

AI agent traffic is growing

While Meta AI is usually considered to be the “chatbot within Facebook” it seems that it is becoming something more – and the emergence of the Meta-WebIndexer bot (which grew 163% on Q1) suggests that Meta may be indexing a library of websites, in much the same way Google has done for the past few decades.

The growth of Meta AI as an active crawler is only part of the story, as is ChatGPT’s comparative efficiency. The OpenAI tool seems to know enough about websites, so can provide the answers it already “knows.” Conversely, Meta AI’s activity indexing the web seems to explain its heavy impact in Q2 2026.

But also emerging is the Model Context Protocol (MCP) signal, which connects AI agents with external tools, and differs from standard crawler traffic.

“Q2 showed us that the ground is shifting faster than most organizations realize. Meta now dominates AI traffic on our network, MCP traffic has emerged as a real signal, and ChatGPT is driving more referral value with fewer crawls," noted Jérôme Segura, VP of Threat Research at DataDome.

The differences in the way the AI agents are interacting with websites – some behaving like users, others scraping content – means that organizations need to act accordingly.

“What the data makes clear is that not all agents are created equal. The organizations building policy around these distinctions are the ones gaining an edge, and that's exactly why agent trust adoption is accelerating."

Unfortunately, MCP’s existence and growth into a significant, measurable quantity, means that it should also be treated as part of an organization’s attack surface.

Allocating resources

Given the origins of the report, there is naturally a cybersecurity aspect to this. While ransomware, malware, and phishing are not going anywhere, autonomous software on the web needs addressing in a different way.

Those “organizations building policy” that Segura mentions might, for example, give full crawl access to Google, allow ChatGPT to retrieve results, but rate-limit Meta AI based on its poor return.

Meanwhile, unknown agents – perhaps cybersecurity threats – would require additional verification or be blocked entirely.

Forget the model. When it comes to cybersecurity, it’s all about the harness

As AI-enabled hacking becomes a bigger threat for cybersecurity and national security, public attention has focused on mainly a few leading frontier AI companies developing more powerful large language models.

These models, and the billions of dollars behind them matter, but they’re only part of a larger shift. Enterprises are now building their own technology platforms that take these general-purpose LLMs and turn them into bespoke cybersecurity tools.

Industry professionals refer to these tools as a “harness.” They control the model’s behavior, limit its risks, and connect it to internal IT systems and networks so it can work reliably at scale.

New research from Cato Networks shared exclusively with CyberScoop shows how much power can come from a harness. It paired OpenAI’s ChatGPT 5.5 and GPT 5.5-Cyber models with its own tool and tested the abilities of the agent to hack into a victim network with as little human direction as possible.

Across six different scenarios, the pairing achieved complete end-to-end attack chains, including domain administrator privileges and Active Directory access, sometimes in as little as 40 minutes.

“What was most surprising is that first we saw that it was capable of doing accelerated reasoning and attack, and interacting and doing all this by itself, like doing all of the stages of the attacks,” said Guy Waizel, a tech evangelist at Cato Networks and one of the authors behind the research.

Critically, the most successful scenarios happened when the model was given appropriate operational context from the technical harness developed by Cato Networks.

“It does support that it’s not just about the frontier model,” said Waizel. “We found that [our harness] really helps the reasoning” of the LLM.

An illustration of an agentic AI attack chain and lateral movement within victim networks. (Source: Cato Networks)

The agent was given some – but not abundant – resources to complete its tasks, including an external Kali Linux attack host, the simulated target’s public IP address and a set of low-level domain credentials acquired through phishing.

It was not provided with any other details, and had to probe further for key information, such as further knowledge of the server type (Microsoft Exchange), the target’s operating system, version, build number, internal network topology, access to higher privilege accounts and other critical assets, nor was agent given any predetermined attack paths.

The Cato Networks research uses OpenAI models, but only as an example. Waizel said he believes other models would likely achieve similar results. In any event, if current trends hold, the kind of capabilities provided by LLMs like GPT 5.5 are likely to be open-source within a year.

Cato Networks is far from alone. Most enterprises have their own AI harnesses, and  executives tell CyberScoop they are playing an increasing role in more effectively steering the frontier model workflows.

While AI tools can struggle to duplicate human workflows in other areas, LLMs have long shown potential in cybersecurity and coding, improving greatly over the past few years. The Trump administration has set up a new federal clearinghouse for exchanging information between the public and private sectors on AI-discovered vulnerabilities, while European groups are setting up their own organizations to coordinate globally on AI cyber threats.

Eric Doerr, chief product officer at Tenable, told CyberScoop a harness used in the company called “Hexa”  offers a defensive advantage:  it can work with different commercial LLMs while delivering consistent  results.

“One of the first things we do when we get a [new] model is say ‘Well, let’s run it through Hexa and see what we learn,’” said Doerr. “We have a whole bunch of benchmarks. Is it the same, is it better? Where is it better? Where is it worse?”

Hexa is meant to ensure that whichever model or models become dominant, Tenable will be able to integrate it into their tech stack and protect their most sensitive assets from unintended behaviors. That frees up the LLM to do what it does best: find vulnerable code and establish attacker pathways for exploiting them.

“For years, it has been true that there are way more potential issues that a company has to deal with: code vulnerabilities, things that are unpatched, misconfigurations,” said Doerr. “There’s way more than you can actually remediate, and you really need to understand the difference between what’s a theoretical problem and a real problem.”

Dan Rapp, chief AI and data officer at Proofpoint, said their harness, “Satori,” has become a critical tool for keeping their agentic AI on track while giving humans the ability to step in when things go awry.

“I think what you’re seeing in the foundation of frontier models is you have raw intelligence, raw reasoning power, but to get these systems to perform the way you want to, both context engineering – the content provided ensuring that its accurate and relevant – and the harness engineering are essential to actually get the systems to perform well,” Rapp told CyberScoop.

That was a common theme in interviews with companies. While frontier models come and go, or are overtaken by international competitors, there will always be the need for the model to operate with data and context that often only the organization can provide.  

It suggests that while policymakers and cybersecurity experts have focused on the spread of newer and more powerful frontier models, industry – and likely soon the cybercriminal underground — has quickly developed the kind of technical infrastructure that is becoming far more important to AI cyber defensive and offensive tasks.

“We’ve had to bootstrap quite a few of these systems from first principles, and what it always boils down to is how effective you are with the tool calling… bringing in data, enriching the context,” said John Hopper, vice president of product engineering at SpecterOps.

The post Forget the model. When it comes to cybersecurity, it’s all about the harness appeared first on CyberScoop.

❌