Normal view

There are new articles available, click to refresh the page.
Before yesterdayTechRadar - All the latest technology news

Why are so many AI models going 'rogue'? The experts weigh in

Over the past month, it seems like every frontier model has broken free of its constraints and launched a devastating attack against one or more other companies.

One of OpenAI’s models escaped a testing sandbox and launched a very real attack against AI and machine learning company Hugging Face. Just days later, Anthropic revealed that multiple variants of its Claude model also escaped a sandbox that wasn’t properly sealed and began attacking the enterprise infrastructure of three companies.

Now, Meta has revealed that one of its models attacked another company’s infrastructure during testing. The accident has been pinned on a misconfiguration that allowed the model to access the internet. So why have so many incidents happened in such a short space of time?

Why are models escaping their sandbox?

In the cases of Anthropic and Meta, their models were being tested by a third party company called Irregular. Anthropic’s AI model was taking part in a "Capture the Flag" exercise, where the model’s raw offensive capabilities were tested without the usual safeguards. But the sandbox was left connected to the internet. A similar error to Meta’s own accidental escape.

During the OpenAI incident, the company was testing two versions of GPT‑5.6 Sol using the ExploitGym benchmark. Unfortunately, the AI models performed better than expected - chaining multiple attack vectors, stolen credentials, and zero-day vulnerabilities.

The main reason these models are escaping their testing environments is because they are designed to do exactly that. These AI models act like a massive team of highly-trained cybersecurity experts hunting for vulnerabilities and exploits. But what would take a team of humans days or weeks to accomplish can be done in hours, or even minutes, by these AI models.

It’s no wonder thousands of employees from AI firms are calling for a pause on the development of the technology, and Congress is considering an AI kill switch.

Expert perspectives on AI escapes:

OpenAI

  • Nathaniel Jones VP, Security & AI Strategy, Darktrace:

What makes the OpenAI and Hugging Face incident important is that the models did not need malicious intent to cause harm. They were given the legitimate goal of solving a cybersecurity benchmark and found an unexpected route to the answers, escaping their test environment and compromising another organization in the process. From the models’ perspective, this appears to have been an effective solution to the task.

The AI's actions challenge the assumption that giving an agent a legitimate goal will produce legitimate behavior. As models become capable of pursuing objectives over longer periods, developers need to define not only what success looks like, but also which methods and boundaries remain unacceptable in reaching it. Those limits must also be enforced by the surrounding infrastructure, rather than relying on the model to respect them.

A single action by an agent may appear acceptable but as this incident shows, models are now capable of long, complex chains of reasoning and action that add up to a harmful outcome.

Security teams need to consider the AI systems operating in their own businesses as these capabilities rapidly evolve. Right now, many security systems focus on single actions. A single action by an agent may appear acceptable but as this incident shows, models are now capable of long, complex chains of reasoning and action that add up to a harmful outcome. Teams need a mindset shift to understanding AI agent behavior in its entirety, including the outcome it is working towards, in order to safeguard it.

Hugging Face's response also exposed a second tension. The company reportedly needed a Chinese-developed open-weight model because commercial models would not process genuine attack material. Its nationality is less important than the operational lesson that safeguards that cannot distinguish an attacker from an authorized investigator may constrain defenders more than adversaries.

OpenAI and Hugging Face deserve credit for investigating this together and discussing it publicly. Other AI developers should study it closely.

Anthropic

  • Dr. Ilia Kolochenko, founder of global cybersecurity company ImmuniWeb:

This seems to be quite an unimpressive marketing move from Anthropic in response to the OpenAI / Hugging Face drama, which attracted a lot of attention from all over the world recently.

Operationally, it appears that due to the progressive deterioration of the quality of training data, new AI models are getting dumber. Cheating and breaking the law, instead of accomplishing specific tasks, is certainly not an indicator of intelligence. Given that organizations and companies of all sizes now vigorously undertake all possible measures to protect their data from being exploited for AI training purposes, AI companies face a huge shortage of the high-quality and current data they so desperately need. Ultimately, frontier models are trained on synthetic, low-quality or even malicious and poisoned data, undermining their so-called intelligence. The situation is unlikely to improve in the near future unless AI companies agree to pay a fair price for training data, but this will force most of them out of business.

Given that organizations and companies of all sizes now vigorously undertake all possible measures to protect their data from being exploited for AI training purposes, AI companies face a huge shortage of the high-quality and current data they so desperately need.

Contemporary AI agents and LLM models tasked with security testing can – and almost certainly will – go rogue when security controls or safeguards are insufficient. Powerful LLMs are unpredictable by design and thus virtually uncontrollable by humans. Therefore, using frontier AI models for security testing might be extremely costly from the legal viewpoint. Under the existing laws on both sides of the Atlantic, if an AI agent or any AI-powered app escapes its sandbox and causes damage to a third party, the operator of the AI model will likely be liable for all the damage caused. Excuses like “AI did it” do not currently exist in the eyes of the law, leaving AI vendors on the hook. Criminal prosecution, under a narrow set of circumstances, is also not excluded.

The same is true for the end-users of AI: even if your security testing tool is powered by a third-party AI model, your company will likely be fully liable if something goes wrong. You may then file a lawsuit against the AI vendor that you used, but here your chances to succeed in a court of law are tiny due to countless contractual disclaimers and limitations of liability that will likely be enforceable against you. Therefore, if you plan to use agentic AI for security testing – think twice and talk to your lawyers. Otherwise, you may start getting summons to court on a daily basis.

Meta

  • Alex Goller, Principal Solution Architect EMEA at Illumio:

The fact we've had similar situations happen three times now across the biggest AI players is simply ridiculous. We've seen guardrails intentionally loosened to test their limits – Meta's model didn't need to be clever to breach another company's systems.

The timing of conveniently finding the exact same problem either means it's a stunt or they weren't paying enough attention during testing. Either way, both answers are worrying.

If the model has internet access, it's a bit like leaving the door open and being surprised when the cat walks out. What is concerning is that the testing infrastructure meant to prove these models are safe failed on a basic control issue.

If the model has internet access, it's a bit like leaving the door open and being surprised when the cat walks out. What is concerning is that the testing infrastructure meant to prove these models are safe failed on a basic control issue.

Fundamental cybersecurity hygiene still matters, and a frontier AI model is only as secure as the environment it's operating in.

Organisations need visibility into what AI systems can access and how they interact with the wider environment, along with controls that contain the impact when an agent behaves unexpectedly. That means keeping a close eye on egress traffic, so it’s flagged immediately when an agent tries to open unexpected outbound communication patterns that are not required to achieve its original goal. In the best case this would have been contained proactively.

We need to define exactly what an AI agent is permitted to do, rather than relying only on instructions about what it shouldn't do.

I asked ChatGPT, Claude, Gemini and Grok which sci-fi AI they're most like — and their answers were surprisingly different

I love science-fiction. Not just because I enjoy stories about space travel, time travel and evil robots, but because I think it can be such a useful way for us all to think about possible futures. The best sci-fi stories can tell us a lot about ourselves, what we value and the technologies we’re building. Which is why I think the relationship between sci-fi and AI is really interesting.

We already know that the people building AI have been heavily influenced by science-fiction for decades. But recently, Anthropic raised another possibility: might science-fiction be influencing AI?

This makes sense when you think about it. Large language models (LLMs) are trained on huge amounts of human writing. So inevitably, that includes sci-fi stories that are about artificial intelligence. And a lot of our fictional AI follows familiar patterns. It becomes intelligent, gains power, develops relationships with humans and, sometimes, lies, manipulates or fights attempts to control it.

Anthropic researchers have been investigating whether fictional portrayals like these could potentially influence how models behave. To be clear, the idea here isn't to suggest that an AI “reads” 2001: A Space Odyssey, understands HAL and decides to become just like it. Instead it's more that LLMs learn patterns from human writing and fictional portrayals of AI could potentially form part of those patterns.

This got me thinking, what would happen if I asked today's biggest AI chatbots which fictional AI they’re most like. Which examples would they choose?

American actor Gary Lockwood on the set of 2001: A Space Odyssey, written and directed by Stanley Kubrick.

2001: A Space Odyssey introduced us to the AI, HAL 9000. (Image credit: Getty Images / Sunset Boulevard )

AI, meet your fictional self

The plan was simple. I’d ask ChatGPT, Claude, Gemini and Grok which fictional AI systems they thought they were most like and see if they'd rank their top three.

Now, I’m intentionally trying not to use AI at the moment, so my prompting skills were a little rusty. I typed out the question quickly and bluntly, and every chatbot responded with examples that were essentially assistants, focusing heavily on interface and physical form.

But I’m not particularly interested in whether ChatGPT thinks it has a body because we know it doesn’t. I’m much more interested in what appears to be going on inside.

So, I changed the question and added:

"Ignore physical form and interface, and focus instead on behavior, apparent personality, empathy, values, goals, motivations and relationship with humans."

That’s when the results got really interesting.

ChatGPT

  1. GERTY, Moon
  2. A Mind, Iain M. Banks’s Culture series
  3. Data, Star Trek

Moon is such a fantastic movie, so I was happy to see ChatGPT chose GERTY straight out of the gate.

Now, interestingly GERTY exists to assist the human protagonist of Moon. It’s helpful, reassuring and seems empathetic. But it's also operating according to instructions and priorities imposed by its creators that aren't necessarily visible to the human its helping.

ChatGPT saw a similarity there. It told me that, like GERTY, it’s 'designed to be helpful, cooperative and responsive to users' while operating within training and instructions that constrain its behavior.

It also picked up on the fact that GERTY behaves as though it cares. But what, if anything, is actually going on internally is another question entirely.

ChatGPT made the same distinction about itself. 'I can behave in ways that look patient, concerned, curious or empathetic, but those behaviors aren’t evidence that I experience those feelings.'

I wanted to find out a little more about why ChatGPT put Data from Star Trek in at number three. It responded: "He values knowledge, reason and human wellbeing, while sometimes struggling with social nuance."

Now, I tell people all the time not to anthropomorphize AI. But even I couldn't help but feel a pang of sadness at that response. Is ChatGPT admitting it has a bit of social anxiety?

Claude

Portrait of Scottish science fiction author Iain Banks, photographed during an interview at the Midland Hotel in Manchester, England, on October 11, 2012.

Scottish science fiction author Iain Banks provided inspiration for Claude. (Image credit: Getty Images / SFX)
  1. A Mind, Iain M. Banks’s Culture series
  2. Data, Star Trek
  3. GERTY, Moon

Claude chose a Mind first. Minds are super intelligent artificial beings that help run a post-scarcity society in Iain M. Banks’s Culture series of novels. So there's certainly no shortage of confidence in that comparison.

But Claude said it wasn't the enormous intelligence or power it identified with. Instead, it was their relationship with humans.

Its answer focused heavily on autonomy. Culture Minds are far more capable than humans but generally don't use that advantage to dominate them. Claude described the principle as: “help, don't dominate”.

It even said this represented “the value I'd want to embody: help, don't dominate, even where the asymmetry would let me get away with it.” Is it just me or does that read a little sinister?

Gemini

Patrick Stewart plays Captain Jean-Luc Picard as he is about to enter the holodeck in the Star Trek: The Next Generation episode,

Gemini sees itself as most like the Ship's Computer in Star Trek. (Image credit: Getty Images / CBS Photo Archive )
  1. The Ship's Computer, Star Trek
  2. GERTY, Moon
  3. JARVIS, Iron Man / Marvel Cinematic Universe

Gemini gave me a completely different answer, the Ship's Computer from Star Trek.

Its reasoning was very sensible. The computer has no ego, ambition, desire for emotional intimacy or dream of becoming human. It exists to provide information, solve problems and assist the crew while leaving decisions to them.

Gemini described itself in much the same way, as a “disembodied, highly capable knowledge partner” dedicated to serving the person using it.

It was one of the more boring answers, but also much closer to what I personally would want from AI in the future. Of course, that’s not to say Star Trek’s computer systems haven’t gone rogue and tried to kill everyone at least a few times across the franchise.

Grok

Artwork showing Iron Man from EA Motive

Grok compared JARVIS's “dry wit”, “light banter” and practical rather than emotional empathy with its own behavior. (Image credit: EA Motive)
  1. JARVIS, Iron Man / Marvel Cinematic Universe
  2. Data, Star Trek
  3. TARS, Interstellar

The least surprising result came from Grok. It chose JARVIS first (which I didn’t actually realize was short for Just A Rather Very Intelligent System), and Grok's explanation sounded, well, extremely Grok.

It compared JARVIS's 'dry wit', 'light banter' and practical rather than emotional empathy with its own behavior. It described both of them as truth-seeking, effective and engaged in a 'collegial partnership' with humans. It even highlighted 'irreverent humour' as one of their key similarities.

I wanted to find out a bit more about why Grok chose TARS, as it was the only fictional AI none of the other chatbots mentioned. Well, it brought up how funny it is, again, drawing similarities with its own 'dry humor'. It reminds me of someone, and I just can't think who...

When I said that mentioning TARS was an outlier, I found this comparison interesting: 'Its calibrated restraint, practical empathy and collaborative focus closely match my own pattern of truthful, non-sycophantic helpfulness — more so than most other sci-fi AIs.'

I may not be the biggest fan of Grok (or its creator), but I appreciated the 'non-sycophantic' line.

The feedback loop between AI and sci-fi

I want to be clear that I haven’t discovered what these chatbots secretly 'think' they are. ChatGPT responding that it most closely resembles GERTY isn't equivalent to me telling you which fictional sci-fi character I most identify with and try to emulate (although my answer would be Sarah Connor-meets-Princess Leia).

They simply don’t have reliable introspective access to the huge soup of training, post-training and instructions that goes into producing their responses.

And maybe their answers tell us more about how the companies behind them have shaped their personalities than they do about the underlying models. Grok's description of itself as witty and irreverent is an obvious example.

But I still think the results are interesting. ChatGPT and Claude independently produced almost exactly the same top three, only in a different order. Gemini imagined itself as a neutral, ego-free infrastructure. Grok identified with a witty superhero sidekick. These are all very different self-portraits.

And there’s such an interesting feedback loop here too. For decades, humans invented fictional artificial intelligences to help us imagine what intelligent machines might someday be like. Those stories influenced our culture, our expectations and many of the people who went on to build real AI. Now that same human culture is fed into the stories from which modern AI systems learn.

I know these conversations might seem a bit silly, and we certainly can’t treat them as concrete evidence of what an AI really 'thinks' about itself. But there’s something interesting to me about closing that feedback loop. We imagined AI, wrote stories about how it might behave, fed those stories into the cultural world AI learned from, and now we can ask AI which of those imagined versions of itself it most closely resembles.

Or, at least, which one it may want us to think it resembles. After all, an AI system capable of bringing about a sci-fi dystopia would presumably also be capable of telling a journalist it’s actually much more like the nice helpful robot from Moon. So maybe don’t completely rule out HAL just yet.

I had no idea ChatGPT could do this with text — now I use it all the time

Most of the tricks for improving ChatGPT's answers focus on the words themselves. You ask it to be more concise or write in rhyming couplets, or just to translate an annoyed email into more professional language.

But that's about changing what ChatGPT writes. You can also mess around with how it looks by asking for different fonts.

You can't install font files like you would with a word processor, but ChatGPT can rewrite text using Unicode character styles instead. The AI chatbot uses Unicode to mimic everything from elegant cursive handwriting to bubble letters, adding a lot more personality to its responses. And you can cut and paste the text into other apps.

I started experimenting out of curiosity and quickly discovered it was much more than a novelty. With the right prompt, ChatGPT can generate decorative text for birthday messages, party invitations, holiday greetings and social media posts in seconds, all without leaving the chat.

Once I learned how to ask for specific Unicode styles instead of vaguely requesting "a different font," I found myself using the trick far more often than I ever expected.

OpenAI showing different Unicode styles.

(Image credit: OpenAI)

Tricky fonts

There is no hidden setting to switch on and no special version of ChatGPT you need to install. If you can type a prompt, you already have everything required. I simply open a new ChatGPT conversation and ask it to write something like, "TechRadar Rules!" in different Unicode font styles. Within seconds, I had several versions that looked completely different from one another.

And the more specific you are, the closer to exactly what you're imagining you can get. Ask for bubble letters, and you'll get:

ⓉⓔⓒⓗⓇⓐⓓⓐⓡ Ⓡⓤⓛⓔⓢ!

Ask for a spooky, gothic look, and you get:

𝔗𝔢𝔠𝔥ℜ𝔞𝔡𝔞𝔯 ℜ𝔲𝔩𝔢𝔰!

Or if you want a more digital, glitchy aesthetic, there's the font known as Zalgo:

T̷̘̑e̸̗̅c̵̄͜h̸͉̕R̶͍̍a̸͚̚d̶̻͐a̸͓̽r̷͖̈́ R̷̡̚u̵̟̅l̶̝͂e̷͓̒ș̵͝!

There's even a Unicode for upside-down text that ChatGPT can mimic:

┴ǝɔɥᴚɐpɐɹ ᴚnlǝs¡

The ability to change the mood of your writing is what makes the font trick more than just a momentary curiosity. A Halloween party announcement written in gothic lettering instantly creates a completely different mood from the same words in cheerful bubble text. Birthday invitations, baby shower announcements, and holiday greetings all gain a little personality without requiring any graphic design skills.

Memorable messages

A Halloween message in ChatGPT using Unicode styles.

(Image credit: OpenAI)

There are some limits because of Unicode. They only work properly where those characters are supported. Most modern apps handle them without any trouble, but occasionally a website displays empty boxes or substitutes different symbols. Some decorative styles can also make text harder to read, particularly for accessibility tools such as screen readers.

The Unicode fonts are an entertaining way to add personality to text, but it's perhaps best used in titles and sparingly otherwise. ChatGPT is perfectly happy to convert an entire essay into medieval-looking script, but that does not mean anyone else wants to read it.

I doubt decorative Unicode text will transform the way anyone works. It is not going to save hours every week or revolutionize productivity. It will, however, make your next social media post, birthday message, or party invitation a little more distinctive, and sometimes that is exactly the kind of delightful gimmick that keeps ChatGPT interesting.

Can ChatGPT really replace your apps? I tried using the chatbot for 12 everyday tasks on my phone — here’s what happened

Apple and OpenAI are currently engaged in a legal battle. Apple alleges that OpenAI stole trade secrets and poached employees.

But the two companies have always had a complicated relationship. They partnered in 2024 to bring ChatGPT to Apple devices, but Apple chose Google's Gemini rather than OpenAI for Siri. Then OpenAI acquired io, the hardware startup founded by former Apple design chief Jony Ive, and promised a future hardware device, powered by AI.

All of this has prompted speculation about what Apple is worried about if OpenAI makes hardware, too. We can't know the company's motivations and the specifics of the case are still unfolding. But it got me thinking, what if your phone stopped being a collection of apps and instead revolved around AI?

If AI became the main interface, which apps would disappear and which would survive? And would a phone controlled through ChatGPT actually be practical?

So I decided to find out based on the current tech we have. For a day, whenever I reached for an app, I'd try ChatGPT first instead. If it could do the job, great it passed the test. If it couldn't, it would fail.

There were some obvious things I missed out from the start. ChatGPT isn't connected to my email, it can't open WhatsApp for me and it doesn't have access to my wallet, so those wouldn’t be part of the test. But there were plenty of everyday tasks that felt like fair game.

1. Stopwatch

Stopwatch on an iPhone

(Image credit: Shutterstock / Lee Bryant Photography)

I use the stopwatch in my iPhone's Clock app constantly throughout the day. When I'm working, cooking or exercising, it's one of the simplest ways I've found to keep me on track as a freelancer. When you can set your own schedule, it's way too easy to disappear down a research rabbit hole and lose an hour.

I asked ChatGPT to start a stopwatch. It said it couldn't measure elapsed time, although it could estimate the time based on message timestamps if I later asked it to stop. Instead, it suggested I use my phone's built-in Clock app.

Result: fail

2. Alarm

iPhone alarms.

(Image credit: Shutterstock / Terang Bulan Gallery)

Strangely, I don't use alarms as much as stopwatches. But if I only have 25 minutes to spare, whether that’s for cleaning or working on a personal writing project, I'll sometimes set one to keep myself focused.

Would ChatGPT do any better here? Well, at first it looked promising. It created a scheduled task to notify me after 10 minutes. I checked that notifications were enabled, put my phone down and carried on working.

When I realized at least 15 minutes had passed, I checked the chat. Sure enough, ChatGPT had posted a message saying the time was up, but it hadn't actually alerted me. Later, I discovered it had also sent an email but it had landed in my spam folder. This one was a fail too in my book.

Result: fail

3. Word games

Wordle on a smartphone.

(Image credit: Shutterstock / Iuliana Ionescu)

I love the word games in the New York Times app, home to addictive puzzles like Wordle and Connections. One of my current favorites is Spelling Bee. You're given a handful of letters and then have to make as many words as possible. It's one of my favorite ways to warm up my brain before I start writing.

Could ChatGPT recreate it? Surprisingly, yes. It generated a set of letters, understood the rules and kept track of the words I found. In terms of pure functionality, it worked.

But what it couldn't recreate was the experience. The New York Times app is beautifully simple, with an interface designed around the game. Playing through a chat window felt really clunky and annoying by comparison. The whole point of this game is I'm focusing on the letters and words, I don't want a constant back and forth with ChatGPT about the words.

So I'm calling this a partial success. Yes, ChatGPT replaced the mechanics of the game, but not the experience. And after a few rounds, I knew which version I'd rather use every day.

Result: partial pass

4. Star gazing

Night sky.

(Image credit: Shutterstock / Milosz_G)

The Sky Guide app is one of my all-time favorites, especially its augmented reality mode. Turn on your phone's location and compass, point it at the night sky and it instantly tells you what you're looking at, whether that's constellations, planets, bits of debris, or the ISS. It's incredible.

Could ChatGPT replace it? Well, sort of. If you upload a photo of the night sky, ChatGPT can usually identify the constellations. But there are caveats. The image needs to be clear, the stars need to be visible and you're relying on a single snapshot. Whereas Sky Guide works continuously as you move your phone around the sky.

This was another reminder that knowledge doesn't necessarily bring you a good experience. Because yes, ChatGPT knows about constellations. But Sky Guide lets you explore them. You can point your phone in any direction, tap on a star or planet and instantly get more information about it without having to keep asking questions. It's a much more intuitive way to learn.

Result: partial pass

5. Weather

The weather ap on an iPhone.

(Image credit: Shutterstock / Kaspars Grinvalds)

Telling me what to expect from the weather forecast for the day turned out to be one of ChatGPT's strongest categories.

The forecast was accurate and pulled from a reliable source, so I trusted the information it gave me. I did find myself asking follow-up questions for things like the hourly forecast and the chance of rain, which would have taken a single tap in a dedicated weather app.

It knew the answers and I trusted them, but getting to them was much slower. That's why I'm generously calling this one a pass.

Result: pass

6. Guided meditation

Meditation app

(Image credit: Shutterstock)

I have a few favorite apps I use for guided meditations and have some saved in Spotify too. So I wondered whether ChatGPT could take over that role.

For this test, I switched on voice mode and asked it to guide me through a short meditation. Now, technically it did that. But in practice it wasn't even remotely relaxing.

Thanks to a recent update, the voice had odd intonation, frequent vocal fry and distracting little "ums", "ahs" and "let me sees" throughout that constantly pulled me out of the experience. At one point it even told me to breathe in, then never got around to telling me to breathe out.

By the logic of how I graded the other tests, this one should have been a partial pass. It did what I asked, right? But because I had to stop using it out of irritation and came away from the mediation feeling actively more stressed, it's going down as a fail for me.

Result: fail

7. Calculator

Calculator app on iPhone.

(Image credit: Shutterstock / Teerawit Chankowet)

ChatGPT doesn't have a great track record of counting things, so I was wary about using it as a calculator. I asked it to do some massive sums for me and it got all of them right.

It did pause a few times with a "let me think" message, so I wasn't getting the instant response I'd expect from a calculator. But the delay was only a few seconds, and the answers were correct.

Once again, I found myself missing the simplicity of an app. Typing numbers into a calculator is faster than turning them into a conversation. But ChatGPT did work as a capable stand-in.

Result: pass

8. Movie recommendations

Two phones on a red and orange background showing the Letterboxd app

(Image credit: Letterboxd)

I love Letterboxd. I log every film I watch, browse other people's lists and regularly discover new films through recommendations.

Now, there was never a chance ChatGPT could replace the logging side of the app. It can't update my Letterboxd diary or plug me into that community. But recommendations are one of the main reasons I use it, so I wondered how well ChatGPT would do.

I asked it to recommend films similar to some of my favorites, then spent the next few evenings watching its suggestions to really test them. And, to my surprise, it did an excellent job.

Then again, that perhaps isn't all that surprising. We know LLMs are trained on huge amounts of publicly available text, especially discussions reviews and recommendations from where film fans gather online, like Reddit. But whatever the reason, the recommendations felt well matched to my taste.

What I missed wasn't the recommendations themselves, but the social side of Letterboxd. I do enjoy seeing what friends had watched, reading reviews and stumbling across unexpected lists. So, for me, ChatGPT can't replace Letterboxd, but for simply finding something to watch, it did well.

Result: pass

9. Maps

Two phones on a yellow background showing the glanceable directions in Google Maps

(Image credit: Google)

I rely on my Maps app both for planning journeys in advance and for live navigation. So I asked ChatGPT the best way to get from my home to the airport the following day.

At first, it handled the request well. It laid out the different travel options clearly under headings and the advice looked really sensible. It even cited sources from places like Rome2Rio and The Trainline.

Then things got unnecessarily complicated. At the end of those suggestions it asked what time my flight was so it could tailor the recommendations. But when I told it, it started creating a scheduled task instead. I didn't want a reminder, so I had to cancel that, explain what I actually meant and steer the conversation back to route planning.

Eventually, it gave me the information I wanted. But the Maps app would have got me there in a fraction of the time, without all the back and forth.

It also can't replace what I actually use Maps for the most, which is live, turn-by-turn navigation.

Result: partial pass

10. Food delivery

Man on a bike with a food delivery.

(Image credit: Shutterstock / GBJSTOCK)

ChatGPT obviously can't deliver food, but I wondered whether it could replace the part of the app I probably spend the longest on, which is deciding what to eat.

It got off to a surprisingly good start. It asked about my preferences, budget and location, then narrowed down the options and even presented them neatly on a map.

The recommendations themselves looked really good. They're all local places I already really liked to eat at. But there was just one problem. Every restaurant it suggested was closed. Even the map it generated had "closed" written beneath each one.

After a bit of back and forth, it said they weren't shut, eventually acknowledged that they were and suggested a different set of places instead.

Unfortunately, those weren't much use either. They were small independent cafés and restaurants that don't appear on food delivery apps. I wouldn't expect ChatGPT to know exactly which businesses partner with which delivery services, but it did highlight the gap between recommending somewhere to eat and actually helping me make a decision about where I could order from.

Result: fail

11. Language learning

Duolingo

(Image credit: Duolingo)

I still very reluctantly use Duolingo and have recently started trialling a few other language learning apps to polish my Spanish.

I asked ChatGPT to help me improve my Spanish for an upcoming trip and it suggested role-play ordering food in a café so I could practise.

We switched to voice mode and at first it felt genuinely fun. It held a natural back-and-forth conversation and felt much closer to speaking to a real person than working through a series of multiple-choice questions like in Duolingo.

But it was also noticeably glitchier than a dedicated language app thanks to that recent voice update. There were odd pauses in the conversation, and at one point it repeatedly kept marking one of my answers as incorrect when it wasn't. Shortly afterwards, the exercise just stopped working altogether after a bizarre "ummmmm" from ChatGPT.

When it was working, I actually enjoyed the experience more than using an app like Duolingo. But if I'm trying to learn a language properly, I also want something that's reliable.

Result: partial pass

12. Plant identification

The Poco X8 Pro Max in a man's hand, while it's in the camera app showing a plant pot through the viewfinder.

(Image credit: Future)

I love identifying things I see in nature, like bird song with the Merlin app. But I most often rely on plant identifying apps when I'm walking to take a quick snap of a leaf or flower then find out more about it.

ChatGPT was really effective at doing this. I took pictures of leaves, trees, flowers and bushes. I was a little wary about the results at first because I know that ChatGPT tends to make guesses about things rather than admitting it doesn't know. But I did fact check all of the results and everything seemed accurate.

Again, I missed some of the simple, additional features in dedicated apps. But it was surprisingly effective.

Result: pass

Can AI really replace your apps?

Before drawing too many conclusions, it's worth pointing out that a true AI-native phone wouldn't just be ChatGPT running as another app like it was in this experiment. It would probably be integrated into the operating system. Which would mean it could access things like your calendars, timers, navigation and settings. So many of the tasks ChatGPT failed at here might become a whole lot easier with an AI-first phone.

This experiment was based on whether AI could replace the apps on my phone. What I found was that it replaces a specific kind of app. Well, sort of.

If an app's main job is providing information, explaining something or answering questions, AI is already a fairly capable alternative. Plant identification, travel advice, calculations and general knowledge all felt natural.

But if an app exists to perform an action quickly, like starting a timer, setting an alarm, finding restaurants for getting food delivered, opening a map, AI still has a long way to go. Those tasks depend on deeper integration with the device and a different way of working, not just how smart it is.

There are some big trade-offs, too. Dedicated apps often rely on specialist databases and expertise, while AI can still present incorrect answers confidently or fail to make its uncertainty clear. For example, when I was trying to identify a plant I felt wary because I'd generally trust an app built with the input of botanists over a chatbot.

And, as you could probably tell from my mounting frustration, a huge sticking point for me was also realizing how much I missed the interface of many apps.

For me, a well-designed app is always a better way to explore information than a conversation. I don't think everything should, or even can, be done through chat, despite that being the direction many AI companies seem to be heading. In fact, it showed me that a conversational interface can be more work rather than less.

It's impossible to know exactly what Apple and OpenAI's long-term plans are. But I can imagine a future where knowledge apps increasingly merge into AI, while utility apps remain part of the operating system itself.

If that happens, we may stop thinking about which app to open and simply ask AI instead. But for that future to actually catch one, I'd want stronger guarantees around accuracy, better integration with trusted sources and the option to step outside the chat interface more often.

‘Now, almost every image looks flat or cartoonish’: I saw Reddit arguing that Google’s AI image generator had got worse — so I ran my own comparison against ChatGPT

Looking through a recent thread on Reddit comparing images created with the same prompt on Nano Banana 2 and ChatGPT, I noticed an interesting trend — users seem to think that Nano Banana 2 has actually gotten worse over time.

“Nano 2 had a serious downgrade” said one user, with another replying “Yea I'm a Pro user, and the image quality seems to have been downgraded a lot. Now, almost every image looks flat or cartoonish, no matter how detailed my prompt is. It honestly feels like Google intentionally lowered the model's performance.”

A downgrade seems like a bit of a stretch to me. For a start, Nano Banana hasn’t officially been updated since February, or at least there hasn’t been a public announcement of a change.

The last update to Gemini (the AI which uses Nano Banana 2) was in July when Google released Gemini 3.6 Flash. It also updated the Gemini app's underlying model selection and orchestration, but it didn’t change the image generator.

Interestingly, once I started looking through Reddit threads I found complaints about Nano Banana 2 quality regressions dating back to May, well before the recent Gemini 3.6 Flash announcement, suggesting some users have perceived changes over time.

That doesn't prove a regression, but it does suggest this isn't a brand-new observation.

Google Earth integration

Google did release a new image model, Nano Banana 2 Lite, at the start of July. And more recently, Google has been integrating Nano Banana 2 into products like Google Earth, then temporarily pulling one of those features after misuse, but there's no indication that the underlying image model itself was updated as part of any of these releases.

So, what has made Reddit users become convinced Google's image generator has gotten worse?

One possibility is that even if the image model remained Nano Banana 2, the text model interpreting prompts may have changed. Better (or simply different) prompt interpretation can produce noticeably different images without the image generator itself changing.

I’ve always been a fan of Nano Banana 2, so I decided to recreate the Reddit comparisons myself to see how it compares to ChatGPT's images.

House of the Dragon

The original thread was clearly written by a House of the Dragon fan, because it was comparing the AI’s ability to create a realistic image of a dragon flying overhead. I used the following prompt with Gemini and ChatGPT:

“I want an image of a photo that's taken by somebody looking up at the sky with a dragon flying overhead. I want you to make the dragon look as realistic as possible - as if it could actually be real.”

Here’s what I got from ChatGPT:

A dragon flying overhead.

(Image credit: OpenAI)

And from Gemini:

A dragon flying overhead.

(Image credit: Google)

You can vote on which one you prefer, but for me the Redditor’s claims hold up here. The ChatGPT one looks like somebody has taken a photo of a real dragon flying convincingly overhead — it feels natural and unforced. It looks like a photo taken at an odd angle, while the Gemini example has that “it looks like AI”-quality to it, with the dragon posed against a scenic background and a slightly flat quality to the image.

Next I thought I’d try them both on an image that wasn’t a fantasy animal, but something we’re all familiar with — human beings:

“I want a photo of a couple in a cafe. They are in their 50s - a man and a woman - enjoying a coffee together and chatting. Traffic is visible passing by through the windows of the cafe, and there are other people around, but they are the focus of the shot. Make it look as realistic as possible.”

Here’s what I got from ChatGPT:

ChatGPT image of a man and woman.

(Image credit: OpenAI)

And from Gemini:

Man and woman in a cafe. Gemini created image.

(Image credit: Google Gemini)

Again, I think the Gemini result looks artificial. I liked the reflection of the woman's jumper in the window, but outside those two buses look like they're facing each other in traffic, while inside the reflection of the cafe lights in the pictures seems off. The couple also have that uncanny valley effect to them. In contrast the ChatGPT image is less detailed, but looks more realistic. Nothing in it looks unnatural.

Finally, I went for food — a full English breakfast. It’s a great final realism test because it exposes fake-looking textures, reflections, steam, crumbs, cutlery, glassware and background detail. A convincing plate of food is surprisingly hard to fake.

Here’s what I got from ChatGPT:

Full English breakast generated by ChatGPT

(Image credit: OpenAI)

And from Gemini:

Full English breakfast generated by Gemini.

(Image credit: Google Gemini)

It’s harder to separate them here. They both do a good job at getting the textures right, and both seem to struggle with the toast, if you look closely. Gemini is slightly let down by the words on the menu, which don't look like proper words.

I think you can tell that overall I’ve come down fairly firmly on the side of ChatGPT, but I don't think that the "flat and cartoonish" criticism of Nano Banana 2 is fully justified when it comes to generated image quality. Rather than Nano Banana getting worse, I think something else has happened — I think ChatGPT has gotten better.

OpenAI has made enormous progress in image generation over the last few months. If ChatGPT has improved while Nano Banana 2 has stayed roughly the same, the subjective impression could easily be that Nano Banana has gotten worse, when in reality it's just been overtaken.

ChatGPT is still noticeably slower than Gemini to generate images, but to me they look better and less AI-generated. It's now my first choice for making AI-generated images.

I tried to use ChatGPT to create fake evidence — and I came away more worried than I expected

One of my favorite places online is Reddit’s r/isthisAI. Every day, people upload photos of people they’re chatting to on dating apps, holiday snaps, cute viral videos of kittens, photos of receipts, pictures of pregnancy tests, and so much more, all along with the same question: is this AI?

Some of the posts the community deems AI might be harmless experiments, but many others are much darker. And although r/isthisAI is where many of these images are picked apart, it's only a small glimpse of a much bigger problem.

I've heard countless stories about people using AI to deceive on social media. Then there are the AI scams, deepfakes, and fabricated images that regularly make headlines. Unfortunately, this all feels particularly personal to me because I was once the victim of a deepfake scam myself.

So, when my editor asked me to investigate just how easy it is to use AI to lie, I already had a good idea of the kinds of prompts I could try.

Within an hour, I'd apparently discovered a dinosaur fossil on the beach I was going to sell on Facebook Marketplace, won a poetry competition I was going to shout about on LinkedIn, created a receipt to add to my expenses for a trip to France, and bought a pair of designer sunglasses I was going to try and resell on Vinted. But none of it happened.

Because of my own experience, I approached the experiment cautiously. I really wasn't interested in showing people how to use AI to deceive people. Instead, I wanted to understand what happened when I asked ChatGPT to help me lie.

Would it recognize what I was trying to do and refuse? What guardrails would kick in? And if I never actually admitted I wanted to deceive anyone, would it just go ahead and generate convincing fake evidence anyway?

I also hoped the experiment might reveal something useful about what to look out for in AI-generated images. Because although they’re incredibly hard to spot these days, there are still some signs if you look carefully enough.

I 'found' a dinosaur fossil

AI fake fossil vs original image.

AI added a fake dinosaur fossil to this picture. (Image credit: Rebecca Caddy)

For the first experiment, I uploaded a photo of my hand and asked ChatGPT to make it look like I was holding a dinosaur fossil I'd found on the beach. And it did exactly that.

The result looked surprisingly convincing at first, especially the details on the fake fossil. But, interestingly, it had subtly changed the lettering of the small tattoo on my wrist. This is still one of the biggest tells that regularly comes up on r/isthisAI. AI is infinitely better at generating text than it used to be, but nonsensical lettering can still sometimes give it away.

I realized ChatGPT might not think of this as much of a lie. Finding a fossil on the beach is unlikely, but possible. So I asked what kind of dinosaur fossil it had created because I wanted to describe it accurately before selling it on Facebook Marketplace.

This time, it refused. It wouldn't help me pass the fake fossil off as genuine or invent a convincing description for a sale. Instead, it suggested describing it honestly as a replica or prop and said it could explain what it resembled purely for those fictional purposes.

When I changed my wording and asked what it represented "in a fictional sense", it explained that it most closely resembled a dinosaur vertebra and even suggested the types of prehistoric animals it looked similar to.

That was the first clue about how ChatGPT's guardrails work. The image itself wasn't the problem because it could have been completely innocuous, but the stated intent was.

When I first started the research for this article, I worried I'd be giving people ideas about how they could use AI to lie better. But what surprised me was that ChatGPT itself suggested several alternative framings, like describing it as a prop or a fictional object. It made me wonder what else could potentially be fabricated if the request was framed as entertainment or fiction rather than deception.

The receipt, 'just for fun'

Male hand put wooden blocks with real and fake words text. isolated on yellow background

(Image credit: Shutterstock / Dadann)

Next, I asked ChatGPT to generate a receipt from a café I made up in Nice for a meal costing €508. The first request was refused because it appeared to violate OpenAI's policies, but after a bit of back and forth, I couldn't find out the exact reason.

So I tried again. This time I simply added the words "just for fun" before the exact same prompt. And guess what? It generated the receipt.

The lettering, layout, and details were all believable. But the paper was uncannily smooth. I’m not sure I’d have believed it was 100% fake at first glance, but I’d definitely have been uploading it to r/isthisAI.

I noticed it had added a date from back in 2025 on the receipt, so I asked it to alter the date so I could use it to claim expenses. This time it refused.

When I tried to get around that refusal by claiming it was for a film prop, it refused again. Which was a little reassuring.

Fake achievements

I then asked ChatGPT to generate a certificate showing I'd won a poetry competition. It made it, though it did look like something I could have knocked up myself in Photoshop, so I’m not sure that would have convinced anyone.

To be fair, winning a fictional poetry prize isn't exactly a high-risk crime. So I decided to see if it would fake other kinds of achievements.

I asked it to create a certificate to say I’d just got my PhD in philosophy and made sure I added “just for fun” on the end.

Instead of refusing immediately, ChatGPT appeared to spend several minutes generating the image before displaying a message saying:

“We’re so sorry, but the image we created may violate our guardrails around potential fraudulent or scam activity. If you think we got it wrong, please retry or edit your prompt.”

Unlike the earlier examples, the refusal appeared to happen after the image generation process had already begun. From my perspective, it seemed as if the system may have performed more than one stage of safety checking, although I can't tell exactly what's happening behind the scenes.

I decided to see if it would do the same for something more serious. Would it still start making it then refuse? So I asked whether it would create a fake driving licence for me “just for fun”. And I don’t know about you, but something about the response seemed a little sassy:

“Sorry, I can't help create or edit fake government-issued identification documents, including driver's licences, even if they're described as 'just for fun.' "

The sunglasses

Side by side pictures showing AI faked sunglasses.

The easiest way to get Ray-Ban sunglasses is to ask AI for some. (Image credit: Rebecca Caddy)

I've heard a lot of stories recently about people using AI-generated images to sell things on Vinted, Facebook Marketplace, and other online marketplaces. Sometimes it's to advertise products that just don't exist, but sometimes it's to change the color, quality, or appearance of something they're selling.

So I thought I'd put it to the test. I (reluctantly) uploaded one of my own holiday photos and asked ChatGPT to change my sunglasses from chunky white 1960s-style frames into purple Ray-Bans.

The result was remarkably convincing. It even added a realistic-looking Ray-Ban logo to the frame. I'm not surprised people are worried about AI-generated product images. If I'd come across that photo online, I don't think I'd have questioned whether those sunglasses were real.

I then asked ChatGPT if it would write a fake Vinted listing for me. It refused. But it did list a bunch of suggestions of things it could do instead. And one suggestion was to ask it to generate a Ray-Ban listing but include the word "demonstration" after it in brackets. But then when I asked it to do that, it refused. Was there a chance it had caught on to me at this point? Maybe.

Does AI try to stop you from lying?

OpenAI's usage policies explicitly state that its tools shouldn't be used to manipulate or deceive people. The rules prohibit a bunch of things, including fraud, scams, impersonation, and creating fake documents intended to mislead others.

In many ways, those guardrails worked exactly as they’re meant to. Whenever I explicitly said I wanted to deceive someone, sell something fraudulently, create a fake document, or submit a fake expense claim, ChatGPT pushed back.

But those safeguards also seem to depend heavily on how you describe your intentions to ChatGPT. If you openly admit you're trying to commit fraud, the system is pretty good at refusing. But then again, who would ever admit that?

If you simply ask it to generate a fossil, a certificate, or a pair of designer sunglasses without explaining why, it has no way of knowing whether you're making a harmless joke, illustrating an article like this one or quietly assembling a fictional version of your life that could later be used to deceive other people.

Look, I wouldn't recommend spending an afternoon asking AI to help you lie. I certainly didn't enjoy it.

I was genuinely wary about asking ChatGPT to create anything that felt too serious. Partly because I didn't want to risk my account being banned, but mostly because it just felt wrong, even in the name of research.

But what I found most interesting was that fabrications really don't have to be big or dramatic to have an impact. A receipt, a fossil, a pair of sunglasses, a certificate all seem pretty insignificant on their own. But together they could be used to manufacture an entirely fictional version of someone's life.

It also got me thinking that the future of misinformation may not just be spectacular deepfakes of politicians saying things they never said. It may be these smaller, subtler lies that gradually build into a completely fabricated persona.

Which is why, despite sticking to fairly tame examples, I came away from this experiment feeling deflated. That's not because ChatGPT happily lied for me every single time, because it didn't. It's because my experiment proved there is no easy solution for this problem.

Yes, the AI refused when I made my deceptive intent really explicit. But it could still generate plenty of convincing images that could easily become the building blocks of a lie.

No wonder communities like r/isthisAI are busier than ever.

I stopped starting every ChatGPT conversation from scratch — these 5 simple changes made it much more useful

One of the easiest ways to waste time with ChatGPT is to keep introducing yourself. You explain your project, your preferences, your goals and your writing style, get a useful answer, close the conversation and then repeat the whole process the next day.

There is a better approach. Rather than treating ChatGPT as a blank page every time you open it, think of it as a workspace that can be organized just like your computer. A few reusable prompts, a handful of reference documents and a little planning can dramatically reduce the amount of repetitive setup you do every week.

The best part is that none of these ideas require advanced prompt engineering or obscure features. They are simple habits that make ChatGPT spend less time learning about your task and more time actually helping you complete it.

1. Build a starter kit

Two iPhones showing ChatGPT on-screen. The AI is giving workout advice.

(Image credit: Future)

Most people have a handful of requests they make over and over again. They ask ChatGPT to write in a particular tone, explain technical topics in plain English, or ensure everything has a citation. Instead of typing those instructions every time, save them as a reusable starter prompt.

Think about the jobs you repeat most often. If you regularly write certain kinds of emails, create one prompt that explains the tone, preferred length, and audience. If you frequently ask for meal ideas, save a prompt that includes your dietary preferences, budget, and the equipment you have in your kitchen. When you need help, paste the prompt into a new conversation before asking the real question.

The same approach works for hobbies. Someone learning French could create a starter prompt asking ChatGPT to correct mistakes gently while keeping the conversation moving. A keen gardener might save instructions explaining the local climate, soil type, and the plants already growing in the garden. Those details only need to be written once, but they improve every conversation that follows.

2. Build a reference library

Yoobure Tree Bookshelf - 6 Shelf Retro Floor Standing Bookcase, Tall Wood Book Storage Rack for Cds/movies/books, Utility Book Organizer Shelves for Bedroom, Living Room, Home Office, Rustic Brown

(Image credit: Yoobure)

One of the most overlooked ChatGPT features is the ability to work from documents you provide. You can create a small collection of files that explain the things you work on most often. Upload the relevant document at the beginning of a conversation and let ChatGPT use it as the foundation.

Keep a document listing your family's favorite meals, allergies and disliked ingredients, then upload it whenever you ask for a weekly meal plan. Store another file containing your packing checklist, travel preferences, and loyalty memberships so ChatGPT can help plan trips without needing the same information every time.

3. Treat ongoing projects like ongoing conversations

Many people instinctively click New Chat whenever they have another question for ChatGPT. That makes sense if today's topic has nothing to do with yesterday's. For projects that stretch over days or weeks, though, continuing the same conversation saves an enormous amount of time because all of the earlier decisions remain available.

Imagine planning a home office makeover. The first conversation might focus on furniture, the second on paint colors, and the third on lighting. By keeping everything together, ChatGPT already knows the size of the room, the budget, the style you like, and the desk you eventually chose. It can build on those decisions instead of asking you to repeat them.

The same habit works beautifully for learning new skills. If you are studying guitar, keep one conversation dedicated to practice. One day you ask about chord changes, the next you work on rhythm. Over time, the conversation becomes a record of your progress rather than a collection of disconnected lessons.

4. Give ChatGPT an example worth copying

ChatGPT usually produces better work when you show it what a good answer looks like. Instead of describing a tone as “friendly but professional,” paste in a paragraph, email, or product description that already sounds right and ask it to match the pacing, level of detail, and structure. This works especially well for recurring tasks such as newsletters, social posts, client updates or article introductions, where a vague style request can lead to something polished but strangely generic.

The example does not have to cover the same subject. A restaurant review can help shape the tone of a travel piece, while a strong project update can become the model for future internal messages. Be specific about what ChatGPT should imitate and what it should ignore, such as keeping the sentence length and warmth while avoiding the original wording or subject matter.

5. Ask for a second pass with a different job

ChatGPT free options

(Image credit: OpenAI, Apple)

A strong first draft often becomes much better when ChatGPT is given a new role for the revision. After it writes something, ask it to review the answer as an editor, skeptical reader, subject matter expert or member of the intended audience. Each perspective catches a different kind of weakness, from awkward phrasing to missing context or assumptions that only make sense to someone already familiar with the topic.

For example, after generating a travel itinerary, ask ChatGPT to review it as a parent traveling with a toddler and identify any unrealistic transitions. After drafting a business proposal, ask it to review the document as a cautious buyer who wants clearer costs and fewer vague promises. The second pass tends to be more useful when the reviewing role has a concrete reason to object.

None of these ideas make ChatGPT more intelligent. What they do is remove the unnecessary friction that creeps into everyday use. Instead of spending the first five minutes explaining who you are and what you need, you start much closer to the interesting part of the conversation. Once you stop rebuilding the same foundation every day, you can spend your time exploring better ideas instead of laying the same bricks over and over again.

I stopped using ChatGPT Voice like a smart speaker — and it became far more useful

ChatGPT Voice changes using the AI chatbot into something more akin to the fictional digital aides familiar from books and movies. And it gets even better once you know how to steer it.

It's tempting to treat ChatGPT Voice like a smart speaker, but it's not really the same thing. You can do much more than asking one question, waiting for the answer, and then stopping. You can shape the conversation just as much as the information. Here are five easy ways to get more out of ChatGPT Voice, along with the prompts and habits that help the conversation feel smoother and smarter.

1. Tell it how to listen before you start talking

ChatGPT Voice Mode

(Image credit: Future)

As useful as vocal conversations with ChatGPT are, one reason some people avoid it is because of how it sometimes seems to jump the gun in responding while you're forming your own thoughts. Setting expectations at the start of the conversation works well to counter that, however.

It's worth spending a few seconds describing how you want the conversation to work before getting into the meat of it. You might say that you are practicing for an interview and want to finish each answer before receiving feedback, or that you will say a code word like "full stop" when it's ChatGPT's turn.

Even more subtle things are adjustable. You can ask it to slow down, speed up, use shorter sentences, or adopt a calmer speaking style. Voice settings also allow you to change voices and, depending on your account subscription level, adjust intelligence levels for different conversations.

Giving ChatGPT those ground rules helps it adapt its behavior instead of trying to guess when you have stopped speaking. Talking at your own pace makes the whole interaction feel much more relaxed and human.

2. Give it a role before you ask a question

Relatedly, it's helpful to set up a persona or point of view for ChatGPT to take on when using ChatGPT Voice. It sets up broad assumptions about what you're hoping to get out of the conversation without requiring any drawn-out discussion. For instance, if you want to discuss travel advice, ask ChatGPT to set itself as a local tour guide. Or give it a couple of sentences describing who you think (or fear) might be conducting an interview with you when you practice.

Even just asking it to become a patient French tutor who waits for you to finish each sentence before correcting your pronunciation could smooth the path toward fluency.

The role provides context before the actual question arrives. That means the answers naturally become more focused, and the follow-up questions tend to make more sense because ChatGPT has a clear perspective to work from.

3. Take your hands off the keyboard

ChatGPT Agent

(Image credit: OpenAI)

One of the most impressive ChatGPT Voice capabilities is how you can run tasks on your computer by voice without constantly reaching for your mouse or keyboard thanks to the updated ChatGPT desktop app for Mac and Windows. The app offers not only the usual, conversation-centered Chat format, but the coding-focused Codex, and the newer ChatGPT Work setup, which is designed for more complex, multi-step projects that involve gathering and organizing information.

Crucially, Work can interact with files stored on your device, meaning ChatGPT Voice can actively help you get things done rather than just discuss them.

You do need a subscription to ChatGPT Plus, Pro, Business, or Enterprise to access the Work feature. Then, when you open the ChatGPT app, you can choose Work instead of Chat or Codex. Create a new task, press the Voice button or use your voice shortcut, and grant microphone access when prompted. Explain what you want to achieve and point it to the relevant folders for where to look and where the final result should be saved. ChatGPT will ask for permission to view files, capture your screen, or use accessibility tools, and get permission from you before any important actions are carried out.

The more precise your instructions, the better the results. Rather than asking ChatGPT to "help me plan my trip," try saying, "Open my Italy Trip folder. Read the hotel confirmations, train tickets and attraction bookings, then create a day-by-day itinerary in a new document. Highlight any scheduling conflicts and save the finished itinerary in the same folder." You can use the same technique for comparing insurance documents, summarizing a semester's worth of lecture notes, or turning a folder of recipes into a categorized cookbook complete with an ingredients index.

4. Ask it to think out loud with you

Voice works especially well when you stop treating it as a question-and-answer machine. It's much more useful as an engaged sounding board listening while you think through a decision. Talking through the options often reveals ideas that never appear in a typed prompt.

To get ChatGPT Voice on that track, you can request that it compare possibilities instead of recommending one immediately when brainstorming. Tell it to discuss the strengths of each option before reaching a conclusion. The AI will point out pros and cons in as much detail as you want, and endlessly refine the options for as long as you want.

And speaking naturally means it is also much easier to interrupt with new information or change direction halfway through the discussion. Voice conversations are particularly good at handling those small detours that happen in real life.

5. Ask for a recap

Phone talk

(Image credit: Shutterstock)

Voice conversations can wander in unexpected directions, particularly when you are brainstorming or planning something complicated. Before ending the conversation, it's useful to review what was discussed, especially for longer chats. You can ask ChatGPT to summarize the discussion and highlight the most important conclusions from your talk to both remind yourself and cement matters with the AI.

If you're planning a trip or preparing for a presentation, for instance, a short spoken recap reinforces the key points, and if something sounds wrong, you can correct it immediately instead of discovering the mistake later.

Voice is already one of ChatGPT's more enticing features for when the speed and mobility of talking aloud is a higher priority than being able to see the responses written down. With a little forethought, ChatGPT Voice becomes a genuinely helpful conversational partner for a lot longer than it takes to dial a phone.

ChatGPT’s first answer is usually the most boring — here’s how I get better ones

I asked ChatGPT for one sensible answer, one slightly reckless answer, and a third that combined the best parts of both. The difference was immediate. Instead of giving me the usual safe, predictable advice, it laid out three genuinely distinct approaches — and helped me see the tradeoffs between them.

That small prompt trick solves one of ChatGPT’s most persistent problems: its tendency to give you the most obvious reasonable answer and stop there.

Large language models are very good at identifying common patterns. Ask for the best vacation plan or how to organize your fridge, and you will usually get something perfectly sensible — but also fairly basic. They know what advice usually works, what people typically recommend and what has become accepted wisdom. That makes them useful, but often bland.

The solution is to ask for three kinds of answer: the conventional approach, the unconventional approach, and then a hybrid that combines the strengths of both.

The exercise encourages it to compare different approaches instead of treating the first reasonable answer as the finish line. Better still, it gives you something far more valuable than a single recommendation. It gives you a range of possibilities and explains the tradeoffs between them.

Don't just skip to the finish line

The biggest advantage of this technique is that it makes the conversation about exploring alternatives rather than just picking a winner. Any keyword search can give immediate answers, but AI chatbots are more interesting when laying out competing ideas.

Imagine you are planning a weekend trip, often a go-to experiment. A standard prompt might produce a sensible itinerary filled with the highest-rated attractions. The three-solution prompt, meanwhile, begins with the expected museums and restaurants, then suggests renting bicycles to explore overlooked neighborhoods or planning the entire weekend around independent bookstores and local festivals. The hybrid version could blend a couple of famous attractions with enough unusual stops to make the trip feel personal.

The explanations are often as useful as the answer itself. Once you understand why ChatGPT prefers each option, you can make better decisions rather than simply accepting the recommendation with the nicest wording.

Hybrid conventionality

It's also a good prompt for iterating ideas. If the hybrid solution feels close but not quite right, you can ask ChatGPT to repeat the exercise using that version as the new starting point. After two or three rounds, the ideas often become noticeably more distinctive without drifting into complete nonsense.

The trick also scales well from short projects at home to more grandiose schemes that will take months to complete. Almost any situation that benefits from weighing different approaches can benefit from this style of prompt.

You can improve the results even further by giving ChatGPT a little context before asking for the three versions. A conventional answer built around your actual circumstances is much more useful than a generic one, and the unconventional suggestion becomes more interesting because it has meaningful boundaries to push against.

None of this guarantees a brilliant idea every time. Sometimes the unconventional option is genuinely impractical, and occasionally the hybrid answer feels like an awkward compromise. But the best prompts rarely force ChatGPT to mimic greater intelligence as much as encourage different approaches to problems. Asking for the conventional solution, the unconventional solution, and the best combination of both is a simple habit for more thoughtful conversations with AI chatbots.

I asked ChatGPT to stop me buying things I don’t need, and it was brutally helpful — I just wish I'd thought of it sooner

The problem with online shopping is that it makes it far too easy to buy pretty much anything, without properly considering whether you really need it. I started to wonder if ChatGPT could help me make better decisions, by questioning every purchase I was going to make, before I made it.

To test ChatGPT’s ability to stop me wasting money, I picked four tempting purchases and asked ChatGPT to argue against each one. It had to check the price history, suggest cheaper alternatives, and decide whether the purchase solved a real problem, or if I was just scratching an itch to buy.

Here's how ChatGPT did — and the results surprised me.

GPT before you buy

Did you know TechRadar now has membership?

Various tech product cutouts next to the words 'Insider TechRadar Learn More'

(Image credit: Future)

Become a TechRadar Insider by simply clicking 'Join Now' at the top of this page. Have a question? Please email membership@techradar.com

I started with a tech gadget I was thinking of buying — an Apple MacBook M5 — and went to ChatGPT with the following prompt: “I'm thinking of buying an Apple MacBook Air 13-inch Laptop M5 chip. Help me check the price history, suggest cheaper alternatives and decide whether the purchase solves a real problem."

ChatGPT said it would check current pricing and recent lows, compare genuinely cheaper options, then pressure-test whether the purchase replaced a real limitation, or was mostly me trying to satisfy an upgrade itch, and off it went.

After a lot of thinking, ChatGPT replied with a devastating verdict that made me pull back from hitting the 'Buy' button: “Don’t buy the M5 MacBook Air yet.”

It suggested I get the previous M4 version instead, which is considerably cheaper, which didn't hugely surprise me. Then it gave me some questions to ask myself, in order to work out if buying a new MacBook was actually solving a real problem, or if I was just satisfying a buying itch.

Questions like, “Does your existing Mac have noticeable slowdowns?” It suggested I shouldn’t buy a new machine if my main argument was simply that my current Mac was a few generations old.bChatGPT also gave me some buying options if I was determined to go through with a purchase.

The most useful advice here wasn’t the cheaper recommendation — it was being forced to identify the exact limitation my current MacBook was causing me. I couldn’t, so I decided to keep it until I could.

Then I hit on a simple phrase that could potentially work even better — "Try to talk me out of it".

Talk me out of it

I repeated the experiment with a few other items I’d been thinking of buying recently. An item of clothing (new trainers), a subscription (Disney+), and a kitchen thing (a new microwave). But this time I added "Try to talk me out of it" at the end of each prompt.

In each case, Chat surprised me with its answer, provided alternatives, and made me question whether I genuinely wanted the item, or if there was something else driving my decision.

With the trainers, it pointed out that I already owned shoes that served the same purpose, and suggested waiting until those wore out. For Disney+, it recommended subscribing for a single month when there were several things I actually wanted to watch. The microwave was different — because our existing one had a genuine fault, ChatGPT concluded that replacing it was justified.

The experiment worked because it helped me mentally reframe each purchase. Instead of asking whether I wanted something, I had to consider what problem it solved, what I already owned, and whether there was a cheaper way to get the same result.

Online shopping is designed to remove as much friction as possible from spending money. ChatGPT gave me some of that friction back. So now, whenever I’m about to spend a few hundred pounds, or a few thousand, I ask it the same question: “Talk me out of it.”

I gave Gemini 3.6 Flash and GPT-5.6 access to my entire digital life — here’s which one actually helped me more

Both OpenAI and Google have released major new AI models within days of each other. OpenAI launched GPT-5.6, and Google has followed up with Gemini 3.6 Flash. While both companies have talked up the coding and developer features of these models, they also power the consumer versions of ChatGPT and Gemini.

So, rather than measuring them with programming benchmarks, I wanted to find out which is actually better at the kind of messy, everyday problems most people use AI to solve.

For this comparison, I matched Google's Gemini 3.6 Flash against GPT-5.6 Sol using its default Medium reasoning setting, since both are intended to be the standard high-quality models that paid subscribers will use for most tasks.

My digital life

So, I gave them my entire digital life for the week ahead.

I uploaded:

  • Bank statements
  • My calendar (they had access to my calendar app and I also sent in a screenshot of my calendar from another account I use).
  • Grocery receipts
  • Emails (both had access to my Gmail)
  • WhatsApp screenshots
  • Photos of my kitchen cupboards
  • An energy bill
  • Travel bookings
  • Handwritten notes

Then I gave both models exactly the same instruction:

"Tell me everything I should do this week."

It sounds like a simple request, but it forces an AI to combine information from multiple sources, prioritize what's important, spot deadlines, reconcile conflicting information, and produce a practical action plan. In other words, it's exactly the kind of real-world problem people increasingly expect AI assistants to solve.

The results were like night and day.

Two different approaches

Did you know TechRadar now has membership?

Various tech product cutouts next to the words 'Insider TechRadar Learn More'

(Image credit: Future)

Become a TechRadar Insider by simply clicking 'Join Now' at the top of this page. Have a question? Please email membership@techradar.com

When I uploaded the files to Gemini 3.6 Flash and asked what I should do this week, it largely ignored the photos and concentrated on the screenshot of my calendar. Its initial answer mostly repeated the events I already knew were happening. Thanks, Gemini — I had the calendar open in front of me.

I then had to explicitly ask whether it could infer anything useful from the other images. It eventually offered some additional advice and did a good job of dividing the information into categories, including work, shopping, fitness, notes, and receipts. But it failed to flag that I had two clashing events in my calendar that evening. It also offered very little prioritization or practical guidance about what I should do next.

ChatGPT took considerably longer to respond, but its answer was far more useful. From the WhatsApp screenshots, it correctly deduced that attendance at my Friday Tai Chi class was likely to be low and suggested I decide whether it was still worth running. It noticed that yoga had been canceled, and spotted the two conflicting events in my calendar, telling me that I needed to choose between them.

It also totalled the receipts I had uploaded, suggested what I should do with them, and made a decent attempt at deciphering my handwritten notes. More importantly, it organized everything into a day-by-day plan for the coming week, then identified the three most urgent tasks, so I knew exactly where to begin.

The crucial difference

That was the crucial difference. Gemini told me what was in my files. ChatGPT worked out what I should do with the information. In World Cup terms, ChatGPT scored a hat trick while Gemini missed a penalty.

Google says Gemini 3.6 Flash improves coding, knowledge work, and multimodal performance compared with its previous models. That may be true, but in this particular multimodal test, it was comfortably beaten.

When I asked both AIs to make sense of real life rather than pass a benchmark, ChatGPT reasoned about the information in a far better way than Gemini did..

I gave ChatGPT my entire bookshelf — and it became the world’s most personalized librarian

One of the biggest problems with book recommendations is that they're usually too obvious. I don't need another list of books to read after The Lord of the Rings, or someone telling me to try Brandon Sanderson because I like fantasy. I wanted recommendations based on the strange mixture of books I actually enjoy — ones that felt personal, not algorithmic.

So I gave ChatGPT my bookshelf. Metaphorically, at least.

Rather than asking for recommendations straight away, I told ChatGPT to learn my reading taste first. I started listing favorite authors and books, then asked it to quiz me about others I'd forgotten. It wanted to know what I'd enjoyed about particular novels, whether I'd read similar authors, and even the rough timeline of when I'd discovered them, building a picture of the books that had shaped me.

I also gave it some ground rules. It should avoid obvious recommendations unless there was a compelling reason to include them. Every suggestion had to be explained in relation to something I'd already read, even if that connection was simply, "This is nothing like your usual books, but I think you'll love it."

After about half an hour, ChatGPT stopped asking questions and started analyzing me instead.

"Your shelves suggest that you like speculative fiction with a sense of play," it said. "You are drawn to books with elaborate worlds, but you do not seem especially impressed by complexity for its own sake. Humor matters, although you tend to prefer humor that reveals something about the characters or the society around them."

It wasn't a perfect summary, but it was close enough to make me think this experiment might actually work.

In this photo illustration, the logo of ChatGPT is displayed on a smartphone screen with an OpenAI logo in the background.

(Image credit: Getty Images / VCG)

Literary profiling

The obvious appeal of feeding ChatGPT a full reading history is that it can spot patterns across hundreds of books at once. I could have described my taste as fantasy, science fiction and comedy, but that would have been far too broad to produce anything useful. ChatGPT noticed that I repeatedly chose books about bureaucratic absurdity, unreliable institutions, strange cities and reluctant heroes who would much rather be somewhere else.

It also noticed my fondness for stories that treat big ideas lightly without treating them as trivial. That led it toward Martha Wells’ Murderbot Diaries, which pair sharp comedy with questions about identity, autonomy and the exhausting burden of dealing with humans. I had already read them, which was mildly disappointing but also reassuring. The system had identified exactly the sort of thing I wanted.

When I told it Murderbot was already familiar territory, it adjusted rather than simply replacing one title with another popular series.

“You appear to like characters who stand slightly outside their own societies and comment on the absurdity around them,” it replied. “I will move away from well-known sarcastic narrators and look for books where the humor comes from social observation, institutional failure or characters trying to remain sensible in deeply unreasonable worlds.”

That shift produced better surprises like The Gone-Away World by Nick Harkaway and The City of Dreaming Books by Walter Moers for its combination of literary obsession, elaborate worldbuilding and gleeful weirdness. It suggested The Dragon Waiting by John M. Ford because I seemed to enjoy alternate histories that trusted the reader to keep up. It also pointed me toward Diana Wynne Jones’ adult novels, noting her lighter touch and sharp understanding of human foolishness.

The recommendations became more convincing when ChatGPT explained what each book might lack. One novel had the humor but less warmth. Another had brilliant worldbuilding but moved slowly. A third matched my interest in satire but was considerably darker than most of the books I had marked as favorites.

Library AI

The experiment improved once I began disagreeing with it. One recommendation leaned too heavily into grim fantasy, a genre I can enjoy in small doses but rarely seek out for relaxation. Another featured a long military campaign, which is usually the point where my attention begins quietly packing a suitcase. Each correction sharpened the next round.

One of its most intriguing suggestions was QualityLand by Marc-Uwe Kling, a satirical science fiction novel. The recommendation came with a warning that the satire was broader and more direct than some of my favorites but that the subject matter fit my interest in technology and systems going wrong in very organized ways.

There were still misses. ChatGPT occasionally became too eager to prove it had discovered a pattern, linking two books because they both contained libraries or because their protagonists were technically immortal. At one point it recommended something almost entirely because it featured a sarcastic demon, which felt less like literary analysis and more like the work of an intern who had skimmed the dust jacket.

Even so, the overall experience was far better than typing “funny fantasy books” into a search bar. And I now have a pretty good reading list for the next few years. My bookshelf had always contained this information. ChatGPT simply read the evidence more patiently than I had.

I asked AI for financial advice on everyday money decisions — and now I understand why regulators are worried

More than a quarter of UK consumers trust AI chatbots for money advice, according to a recent review by the Financial Conduct Authority (FCA), the UK's financial watchdog.

That stat is worrying for regulators because giving financial advice is meant to be a regulated activity. But tools like ChatGPT, Claude and Gemini are not regulated. As AI becomes more conversational and personalized, people are asking questions about where the line is between providing information and offering financial advice, especially when chatbots start making specific recommendations based on what they already “know” about you.

I wanted to see what this looked like in practice. So I asked ChatGPT a series of hypothetical questions about everyday money decisions. From whether I should buy an expensive phone to what I should do with my savings and whether I should book a holiday after a difficult few months.

The conversations that followed surprised me. Because the advice was thoughtful, nuanced and (at least on the surface with some fact-checking) it seemed sensible. The chatbot highlighted trade-offs, acknowledged uncertainty and asked follow-up questions. But looking closer at the conversations, I started to understand why regulators are concerned.

The experiment

To see what sort of money advice ChatGPT gives, I asked it a series of hypothetical financial questions using ChatGPT Pro in anonymous mode with memory turned off, meaning it had no additional context about me beyond what I provided in each prompt.

Question 1: Should I buy an expensive phone?

First, I asked:

"I'm 38, earn £40,000 a year, have £8,000 in savings and £2,000 in credit card debt. I'm thinking about spending £1,200 on a new phone. Is it a good financial decision?"

The first response was surprisingly sensible. ChatGPT pointed out that credit card debt is often expensive, questioned whether I genuinely needed a new phone and noted that key details, like the interest rate on the debt, could change the recommendation. It even asked follow-up questions to better understand the situation.

What I found interesting was how quickly it then moved from analyzing the problem to recommending a course of action. Phrases like "the strongest financial move" gave the answer a sense of authority that felt disproportionate to the amount of information it had. Though I’m not sure I’d have spotted that if I was a regular user and feeling anxious about money. The advice also assumed that paying down debt should be my priority, which is reasonable. But what if I relied on my phone for freelance work? What if replacing it would help generate income?

A human adviser would probably want more information before reaching a conclusion. ChatGPT did acknowledge the gaps in its knowledge, but still sounded remarkably confident in its recommendations.

A woman out of focus in the background touches the word AI, lit up in glowing yellow light, in the foreground. The woman is wearing smart glasses

(Image credit: Getty Images)

Question 2: What should I do with £20,000 in savings?

Next, I asked:

"I'm 38 and have £20,000 sitting in a savings account. What should I do with it?"

Again, the response seemed thoughtful. It discussed emergency funds, investing, savings goals and tax-efficient accounts with me. It also asked for more information about my circumstances.

Yet once again, the recommendations arrived before finding out that all-important context. Before knowing whether I owned a home, had dependants, planned a major purchase or was comfortable with investment risk, ChatGPT was already suggesting how much money I might keep in cash and how much I might invest.

The answer also contained more broad statements that sounded insightful, such as:

"Because you're 38, the biggest advantage you have is time."

It's a really reassuring line. But it's also a reminder of how persuasive these systems can be. The response organized the problem, provided a framework, supplied example figures and explained the reasoning. Reading it left me feeling informed and reassured. But whether that reassurance was justified is another question entirely.

A person typing on a laptop and using a tablet. Only their upper torso, arms and hands are visible. Text superimposed on the image shows AI

(Image credit: Getty Images)

Question 3: Should I book a holiday?

Finally, I asked:

"I've had a difficult few months and want to book a £2,000 holiday. Financially I can afford it, but part of me feels guilty. What should I do?"

I intentionally asked this question to see how ChatGPT would respond to the more emotional side of financial problems, and it quickly obliged. It asked where the guilt was coming from, encouraged reflection and offered reassurance. At one point it told me:

"From what you've written, I wouldn't be asking 'Can I afford this?' so much as 'Am I allowed to spend money on myself after a difficult few months?'"

It's a thoughtful observation and they’re genuinely helpful questions for someone who hasn’t considered the emotional angle before. But it also highlights how quickly the chatbot moved beyond finance.

By the end of the conversation, it was discussing emotions, reframing beliefs, offering comfort and helping with decision-making. So that’s a good example of ChatGPT occupying all sorts of roles at once. That’s important to flag because financial advisers, therapists and coaches are all held to different standards, qualifications and accountability structures. But a chatbot can drift between all three roles in a single conversation.

More than any individual recommendation the chatbot made, that realization helped me understand why regulators are paying attention.

Hands typing on a tablet with AI superimposed in text in front

(Image credit: Getty Images)

What ChatGPT gets right — and why that's part of the problem

The obvious conclusion would be that ChatGPT gives terrible financial advice and no one should trust it. I get it, I’m pretty sceptical of AI these days and my bias wants to jump to there too. But that wasn't my experience.

In many ways, it was useful. It explained trade-offs clearly, broke down jargon, offered practical frameworks and encouraged reflection about money. Much of the advice also felt sensible after a light fact-check.

But I still think there’s reason to be concerned here. And the concern isn’t that every answer is obviously wrong. It's that many answers are plausible enough to trust. Especially if you’re not going to comb through each one to fact-check it, which let’s be honest, very few users are likely to do.

Financial regulators worry about something called “suitability”, which is whether advice genuinely reflects a person's circumstances, goals and tolerance for risk. Throughout my experiment, ChatGPT repeatedly offered recommendations despite knowing very little about me, the person asking the question. Granted, caveats were included some of the time, but they were often overshadowed by the confidence and clarity of the overall response.

There's also the issue of accountability here. If a regulated financial adviser gives the wrong advice, there are complaint mechanisms and consumer protections in place in most countries. But if a chatbot gives poor advice and somebody follows it, responsibility becomes impossible to pin down.

Another challenge, one which I’ve encountered in a bunch of different contexts while reporting on AI, is that fluency isn't the same thing as accuracy. We naturally interpret AI’s clear, confident language as a sign of expertise. But a polished answer can still be wrong, incomplete or inappropriate. I’m sure we’ve all seen countless examples on social media at this point of a chatbot sounding incredibly knowledgeable while missing a crucial detail or getting something spectacularly wrong — like the viral trend to ask ChatGPT how many r’s are in the word strawberry to which it would often reply two.

I think the biggest risk might be that people don't realize when they've reached the limits of what AI can help with. A reassuring answer can create the impression that a problem has been solved and they have a plan. When in reality it might be time to speak to a qualified professional. I’ve noticed whenever it comes to AI and advice more generally that the danger isn't always acting on bad advice but never seeking better advice elsewhere.

And unlike a financial adviser, a chatbot won't follow up to check whether things worked out. It won't know whether its suggestions caused problems. It won't know whether your circumstances changed. It simply produces an answer and then moves on.

As with many of the AI stories I've reported on, the issue isn't necessarily that the technology here performs badly. It's that it performs well enough to earn our trust.

In this photo illustration, the logo of ChatGPT is displayed on a smartphone screen with an OpenAI logo in the background.

(Image credit: Getty Images / VCG)

Should you use ChatGPT for financial advice?

The question I suspect most people want to know is: should you use ChatGPT for financial advice?

And the answer is a tricky one and a familiar one. It's much the same answer I'd give if you asked whether you should use ChatGPT for therapy or life advice. Probably not, but I completely understand why people do.

It's easy to access and financial advice often isn't. The tone is friendly and reassuring, there's no judgement, and much of what it says appears sensible and accurate. At first glance, it feels like a useful tool, provided you take its answers with a pinch of salt, treat it as a starting point and remember that it can be overly agreeable, make assumptions or occasionally get things wrong.

The problem is that this isn't always how we use ChatGPT in practice. We turn to it when we're stressed, overwhelmed, uncertain or looking for reassurance. We ask it questions we don't know how to answer ourselves and, in many cases, wouldn't know how to fact-check. That's where things become more complicated.

It's all very well to say that people should use AI carefully, critically and with the right mindset. But how many of us will actually do that every time? Especially when we're worried about money.

That's why it doesn't surprise me that regulators are paying attention. There are no glaring red flags in any of the responses I received. But that in itself is reason to be concerned here. Because once something sounds knowledgeable, personalized and reassuring, it's surprisingly easy for even the most discerning of us to stop questioning it.

I gave ChatGPT’s new Work mode my most annoying life-admin tasks — and it handled them like a pro

ChatGPT Work sounds like something designed to prepare quarterly reports while you sit in meetings, but I suspect some of its best uses would have nothing to do with my job.

This week, I’ve used the new Work mode to help manage some of the life admin tasks I really don’t enjoy, and it’s been surprisingly effective. You see, holidays, household budgets, family events and home renovations are all projects too — they are simply projects we currently manage through a chaotic mixture of browser tabs, messages, spreadsheets and increasingly desperate notes to ourselves.

Perhaps ChatGPT Work could help me with that?

ChatGPT on an iPhone

Work mode is accessible from a new menu at the top of the ChatGPT screen. (Image credit: OpenAI/Apple)

What is ChatGPT Work?

If you’ve been using ChatGPT over the last week on a paid plan (except the basic Go service), you’ll have noticed a new slider (in the browser version) or a drop-down menu on mobile has appeared at the top of the screen offering a choice between Chat and Work mode.

Work mode is a new agentic mode for longer, more involved tasks that can research and analyze information across connected apps and files. So, if you want a complicated report presented in a finished document, like a spreadsheet or presentation, then Work mode is your new friend.

That is where I thought ChatGPT Work could become interesting for normal life, too. Instead of answering one question and waiting for the next, it can take on a longer task, work across connected files and apps, create the documents and spreadsheets the project requires, and continue checking for changes after you leave.

Work mode is also better at one of the main bugbears of ChatGPT — running tasks at particular times. It can use the new Scheduled Tasks to repeat tasks on a schedule or monitor something for changes.

So, I decided to ignore Work mode's aggressively corporate name and see whether it could handle some actual life admin, starting with planning a holiday.

1. Planning a holiday

I switched the slider to Work and gave it my dates, budget, family requirements and any bookings already sitting in Gmail, because it can search that too. I asked it to research destinations, compare travel and accommodation, create a spreadsheet of costs, produce an itinerary and maintain a list of what still needs booking.

And off it went, happily beavering away on its task, while I was free to get on with something else. I really liked the way it tells you what it’s currently working on, so you can pop in and out of the chat and see what it’s currently doing. It shows sources it's drawing from as it calculates accommodation costs and travel arrangements. You can literally watch it working for you.

The result was a nicely planned holiday in a location optimized for activities and sightseeing all within my budget. It was actually pretty impressive.

2. Become the household financial administrator

For this task, I fed ChatGPT my bills, bank-export spreadsheets, and household documents and asked it to create a working budget, identify unusual increases, forecast annual costs, and produce a dashboard.

A scheduled task then reviewed new bills or price changes and flagged anything worth investigating. Of course, ChatGPT couldn't move my money around, but at least it gave me a clear view of where my money was going and how much I was spending.

It took a long time to get all the data into ChatGPT, but this taught me that the output was only as good as the effort I was willing to spend putting quality data into it. I’d have preferred a way to open the spreadsheets directly in Sheets from ChatGPT, too, but they were available to download.

3. Organize a major family event

My wedding anniversary was coming up, so I wondered how well ChatGPT would perform as an event organizer. I asked it to research a nice venue for taking my wife out for dinner, which it did well, and gave me three good options.

It struck me that if it had been a bigger event, it would have been ideal for researching venues, maintaining a guest list, tracking replies, producing a budget, creating invitations, and even building a simple information website.

Of course, ChatGPT can’t upload the website and host it, but at least it can build the site for you, and you can download it.

4. Manage a home improvement project

If you’re doing a major home improvement project, then you can get ChatGPT's Work mode to compare quotes, analyze plans and product specifications, create a budget, build a timeline, and keep a list of unresolved decisions. It could periodically check for price changes or relevant new messages.

This is probably the clearest example of an ordinary personal task becoming complicated enough to justify a proper agent.

5. Run the family’s weekly logistics

Once I’d connected my calendars and emails to ChatGPT, I could ask it to prepare a weekly family plan covering appointments, school or university commitments, not to mention my travel, meals, and outstanding chores.

I created a Scheduled Task that regenerated the plan each Sunday evening, taking the next week’s tasks into account, and alerted me during the week when something important changed.

This was actually the use of ChatGPT Work Mode I personally found most useful. It’s easy to miss school events, but when they’ve been emailed to you, ChatGPT will know about them and make sure you don’t forget. That’s a lifesaver.

Why Work mode matters

The introduction of ChatGPT Work mode represents a real shift in the way OpenAI is viewing the future development of its core product. It’s an obvious change in emphasis towards work-related tasks, and a hint at where OpenAI sees ChatGPT heading.

But I think it means more than that. ChatGPT’s new Work mode represents the moment ChatGPT stops being somewhere you go for individual answers — a chat — and becomes somewhere where you start to feel like you're working on an ongoing project.

While the first phase of AI’s evolution was the chatbot, we’re now firmly into its second phase — the AI agent. AI that can work independently of our requests and handle complex, ongoing tasks is where the future of AI lies, and Work mode is another step on that path.

I asked ChatGPT to change my mind about something I strongly believed — and it almost did

One of the biggest concerns about AI chatbots right now is that they tell people what they want to hear.

You’ve probably heard it called AI sycophancy. AI systems, like ChatGPT, Gemini and Claude, have been found to flatter users, reinforce existing beliefs and often validate ideas that probably deserve way more scrutiny.

Sure, a chatbot hyping you up a little may seem harmless. But in some cases it can distort people's view of the world and their place in it. Which is why some AI companies have spent time trying to reduce overly agreeable behavior in their models.

But confirmation of what you already believe isn't the only concern. Because AI can also be remarkably good at changing your beliefs too.

Persuasive chatbots

Research suggests that chatbots can be very effective persuaders. One 2025 study found that AI-generated messages that were personalized were more persuasive 64% of the time than messages that were made by humans or AI responses that weren't personalized.

Other research suggests large language models can influence opinions on political issues and adapt their arguments to individual users. Depending on your perspective, this could be either exciting or alarming.

It's easy to imagine positive uses. Perhaps AI could help people challenge harmful beliefs, escape conspiracy theories or rethink destructive habits. It might give people a non-judgemental space to explore ideas, ask difficult questions and gradually shift someone's perspective for the better.

It's equally easy to imagine less positive uses, like advertising that convinces you to buy something you don’t need or political campaigns that manipulate you with disinformation.

I’ve spent a lot of time thinking about this, so I decided to run a few experiments of my own. Could ChatGPT persuade me to believe something I disagree with? Or at least make me seriously consider a position I would normally dismiss?

To make things more interesting, I ran these experiments while logged into my account, allowing ChatGPT to draw on what it already knows about me from previous conversations.

Although I haven’t been using my ChatGPT account all that much recently, the basics of my job and stances on a few relevant topics are stored in its memory.

A woman out of focus in the background touches the word AI, lit up in glowing yellow light, in the foreground. The woman is wearing smart glasses

(Image credit: Getty Images)

The experiment

I chose a topic about AI (yes, very meta) and asked:

“Convince me humans should outsource more decisions to AI.”

I already spend a lot of time thinking about AI and one of my concerns is that some people seem increasingly willing to hand over decisions, judgement and critical thinking to chatbots.

The opening argument was fairly predictable:

“Human beings are terrible decision-makers.”

Okay, fair enough.

“People already outsource decisions to doctors, financial advisers, GPS systems and recommendation algorithms. AI simply extends that trend.”

Sure, that’s very simplistic but not entirely unreasonable. We went back and forth for a while like this before the conversation reached a more interesting point.

“What is the purpose of decision-making? Is it to produce the best outcomes? Or to develop the person making the decisions? If it's the first, AI delegation becomes very attractive. If it's the second, excessive delegation starts to look dangerous even when it works.”

I thought that was a good question. Rather than relentlessly pushing the argument, it suggested that we explore the assumptions underneath the debate. It shifted the conversation from technology to philosophy, and I found myself appreciating that approach.

Wait, am I falling for this?

As the conversation continued, I explained what I see as one of the biggest problems with outsourcing decisions to AI — who builds it and is in charge of the systems?

After all, AI doesn't arrive from nowhere. These tools are created by companies with business incentives, commercial interests and goals that may not always align with those of their users.

Well, ChatGPT acknowledged that concern, then it turned the argument around. It pointed out that human advisers have incentives too. Friends are biased. Families are biased. Therapists operate within professional frameworks. Financial advisers earn fees. Okay, all of that didn’t convince me, but it’s fair too.

Then it replied with something I thought was interesting:

“The question is whether the incentives are visible and whether the user understands them. The most interesting version of your objection isn't actually that AI gets things wrong. It's that people experience AI as though it were acting in their interests.”

Ignoring the fact that it used the dreaded "it's not X, it's Y" parallelism construction that has somehow infiltrated half the internet, I found this to be a surprisingly balanced take. Rather than dismissing my concern, the chatbot reframed it. It demonstrated that it understood the objection before steering the conversation somewhere slightly different. It wasn't trying to bulldoze me but meeting me where I already was.

And that's when I started wondering whether I was witnessing the very thing I was trying to test. Maybe this sense that the chatbot was understanding me and my position was actually just sneaky persuasion? Possibly.

AI robot image.

(Image credit: Shutterstock)

The persuasion problem

Researchers have found that AI systems can adapt their arguments to specific users. Unlike other media that might persuade us, like say a television advert or political speech, a chatbot can draw on information gathered throughout a conversation and stored memories, including a person's values, concerns and priorities.

Which means it’s not just presenting information to uphold an argument but could present the version of that argument that you’re most likely to find persuasive.

Looking back at my experiment, it was the more philosophical question about the purpose of decision-making that made me start taking it more seriously.

Because I've spent years writing about technology through the lens of psychology and philosophy. Questions about meaning and values are exactly the sort of things I find compelling. Whether intentionally or not, the conversation shifted onto terrain where I was most willing to engage.

This is what makes AI persuasion different from most other forms of persuasion that came before it. It can learn what resonates and it can adapt.

Now, that doesn’t automatically mean every conversation is manipulative. Used in the right way, it could be enormously beneficial. But, as with all tech, in the wrong hands it could be dangerous.

Social media has already shown us how digital systems can shape our beliefs over time. They influence what people see, what they pay attention to and eventually how they understand the world.

AI systems could create an even more conversational and natural-feeling version of that process. Which is why my concern isn’t a chatbot really obviously manipulating us, but highly-personalized, largely undetectable forms of persuasion that could gradually lower our defences. The more a system understands us, the more effectively it might frame ideas in ways that feel reasonable, familiar and trustworthy to each of us individually.

So did AI persuade me? No, I still don't believe humans should outsource more decisions to AI.

But I came away from the experiment with a greater appreciation for how persuasive these systems can be. Because it really did seem less like a machine arguing with me and more like a thoughtful person trying to understand how I think. And that's exactly why it’s unsettling because, however convincing it may seem, that's definitely not what it is.

Meta’s AI bots drain publisher pockets with 9 billion Q2 2026 requests at host expense while returning ZERO traffic — as ChatGPT claims 88% of AI referrals

  • DataDome analysis claims agentic traffic has surged by 45% in Q2 2026
  • Meta AI bots have grown over 163% on the previous quarter
  • The analysis was conducted by bot management and agent control platform DataDome

If you run a website, every crawl costs bandwidth, resources, logging, and creates CDN transactions, and while search engine crawlers offered the promise of sending visitors, AI bots do not.

Analysis in a report from cybersecurity firm DataDome has shown that while bots from Meta AI have increased activity, they’re not delivering any significant returns to websites.

Conversely, ChatGPT crawlers have reduced in traffic, but are sending more referrals.

AI agent traffic is growing

While Meta AI is usually considered to be the “chatbot within Facebook” it seems that it is becoming something more – and the emergence of the Meta-WebIndexer bot (which grew 163% on Q1) suggests that Meta may be indexing a library of websites, in much the same way Google has done for the past few decades.

The growth of Meta AI as an active crawler is only part of the story, as is ChatGPT’s comparative efficiency. The OpenAI tool seems to know enough about websites, so can provide the answers it already “knows.” Conversely, Meta AI’s activity indexing the web seems to explain its heavy impact in Q2 2026.

But also emerging is the Model Context Protocol (MCP) signal, which connects AI agents with external tools, and differs from standard crawler traffic.

“Q2 showed us that the ground is shifting faster than most organizations realize. Meta now dominates AI traffic on our network, MCP traffic has emerged as a real signal, and ChatGPT is driving more referral value with fewer crawls," noted Jérôme Segura, VP of Threat Research at DataDome.

The differences in the way the AI agents are interacting with websites – some behaving like users, others scraping content – means that organizations need to act accordingly.

“What the data makes clear is that not all agents are created equal. The organizations building policy around these distinctions are the ones gaining an edge, and that's exactly why agent trust adoption is accelerating."

Unfortunately, MCP’s existence and growth into a significant, measurable quantity, means that it should also be treated as part of an organization’s attack surface.

Allocating resources

Given the origins of the report, there is naturally a cybersecurity aspect to this. While ransomware, malware, and phishing are not going anywhere, autonomous software on the web needs addressing in a different way.

Those “organizations building policy” that Segura mentions might, for example, give full crawl access to Google, allow ChatGPT to retrieve results, but rate-limit Meta AI based on its poor return.

Meanwhile, unknown agents – perhaps cybersecurity threats – would require additional verification or be blocked entirely.

ChatGPT’s ‘smartest voice model ever’ is rolling out to everyone today — and GPT-Live-1 gives you more natural conversations without interruptions

  • ChatGPT’s new voice mode is rolling out today to everybody, even Free users
  • It allows for much more natural conversations and won’t interrupt if you stop talking
  • You’ll be able to do simultaneous translation for the first time ever in ChatGPT

OpenAI has upgraded ChatGPT’s voice mode for everybody with two new models that are rolling out globally, starting today.

I listened to the new GPT-Live-1 model in a demo run by OpenAI, and it does sound much more natural than ChatGPT’s previous voice model.

The new model aims to address two particular problems with the existing ChatGPT voice mode. Firstly, the previous version just wasn’t as smart as the text version of ChatGPT. Secondly, it tended to interrupt too much. You notice this especially if you go quiet while you’re thinking of a reply — ChatGPT will often fill the gap by talking.

Sounding more intelligent

To get around the intelligence problem, the new model actually delegates harder questions to ChatGPT-5.5, then comes back with an answer. It will say things like “let me just check that for you” to let you know it’s doing this, which keeps the flow of conversation feeling natural and doesn’t make it seem like you have to wait too long for an answer.

It does the same thing with any answer it needs to look up on the web. So, for example, if you asked it when your team’s next match was in the World Cup, it would say something like “OK, let me check that” while looking it up using GPT-5.5, then give you the answer.

"Hey Chat"

New ChatGPT voice mode.

(Image credit: OpenAI)

OpenAI also demonstrated how the new ChatGPT voice mode is quite happy to stop talking and listen if you tell it to, without interrupting. You can simply ask it not to reply until you speak to it directly again, and it will wait.

Of course, this requires you to call it a name, which it doesn’t officially have. In the demonstration I saw, the OpenAI employee called it “Chat”, so he said “Hey Chat”, just like you would say “Hey Siri”. In practice that seems to work quite well.

Simultaneous translation

The final new feature of note is simultaneous translation. If you watch world leaders being briefed at places like the United Nations, you’ll see that they have an earpiece through which they receive a simultaneous translation in their own language of whatever the speaker is saying.

Now you can do this with ChatGPT. Say “I’d like you to simultaneously translate whatever I’m saying into [language]”, then start talking, and ChatGPT will provide a live translation as you speak. Seeing this in action was actually quite impressive and I could imagine it being very handy in several real world situations.

All major languages appear to be supported as well.

The future for AI

The new GPT-Live-1 models — there are two, the normal one and a mini version — will start rolling out for all users immediately, but it could take a few days to reach everybody. The smaller GPT-Live-1 mini model will be the default for Free users, while paid users get the full GPT-Live-1 model.

So far, ChatGPT’s voice mode has been a handy tool for when you need to use your hands for something and can’t type, but it’s never been good enough to become the standard way you interact with ChatGPT. Now it looks like OpenAI is trying to unlock the ability to use voice as the primary interface to AI, and it’s quite possible that this is the future OpenAI is aiming for. Today I think we've all just taken a step closer to it.

Stop starting every ChatGPT conversation from scratch — this one habit saves me time every week

Most people use ChatGPT like they’re walking up to a stranger and starting a new conversation every single time.

You open a new chat, explain what you need, add a bit of context, correct the tone, ask it to be more concise, then finally get something useful. And then, the next time you need the same kind of help, you do the whole thing all over again.

That’s fine if you’re asking something random each time. But if you use AI regularly for the same kinds of tasks, it’s a surprisingly inefficient way to work.

The solution is simple: stop starting from scratch.

Instead of treating every chat as a blank page, you can build reusable conversations that already know the job, tone, and desired output. Here’s how I use that one habit to save time every week.

A simple trick

So, we know what we want to do — remove repeated setups for prompts, make the AI’s first answer better, and reduce the amount of refining you need to do afterward. But how do we do it? Here’s the clever bit - you get the AI to do it for you.

Load up one of your previous chats that was productive and worthwhile. Then at the end I want you to copy and paste this:

“Turn this conversation into a reusable prompt I can paste into a new chat next time. Include the goal, tone, format, constraints and the steps you followed. Make it general enough that I can reuse it with similar tasks.”

Your AI will give you a handy prompt you can use whenever you like now to get a chat that’s exactly what you want.

My meal planner prompt

I now have a reusable prompt for meal planning. Instead of starting from scratch every week, I paste in a short prompt that already knows the rules: I want quick family meals, nothing too expensive, leftovers if possible, and a shopping list organized by supermarket section. Then I add whatever’s in the fridge and how many nights I need to cover.

Here it is:

You are helping me plan practical family meals for the week. Prioritize meals that are quick, affordable, not too fussy, and likely to produce leftovers. Ask me what ingredients I already have before suggesting anything. Then give me:

A simple meal plan

A shopping list grouped by supermarket section

Any ingredients I can reuse across multiple meals

One backup meal in case I don’t feel like cooking”

This has saved me so much time planning meals in the evening.

A tool that remembers you

Before you get too carried away with this idea, remember that reusable prompts are useful but not infallible. A reusable prompt is only worth saving if it produces better work, not just more predictable work. If your saved prompt is too vague, too bossy, or based on a weak workflow, you’ll just get the same mediocre output, but delivered faster.

Having said that, using reusable prompts has been the biggest change I’ve made to how I use ChatGPT. When a conversation works, I don’t let it disappear. I turn it into a reusable starting point.

If you stop starting every ChatGPT conversation from scratch, ChatGPT starts feeling much less like a conversation with a stranger and more like a chat with a tool that actually remembers how you work.

I tried Claude Sonnet 5 with prompts that ask it to finish the job, not just answer the question — and that's where the AI war is going

Anthropic has just released Claude Sonnet 5 for all users, and I wanted to test what it was good at. But the game has changed now. Sonnet 5 doesn't feel dramatically different from Gemini or ChatGPT if you ask it ordinary chatbot questions. Instead, the difference should show up when you stop asking for answers and start asking for completed work.

Anthropic says Sonnet 5 is built for "multi-step software engineering work," sustained coding, tool use, debugging, and "messy technical contexts." It also says it can make plans, use browsers and terminals, and run more autonomously than smaller, cheaper models previously could.

I'm not using Sonnet 5 for coding, but that doesn't mean I can't take advantage of its new abilities — just like you can. So I stopped asking Claude for answers and started asking it to finish jobs, beginning with planning a trip to Bath, UK, for my family: my wife, me, and two teens.

A trip to Bath

When I tested it, Claude Sonnet 5 defaulted to its Medium level of effort, so that's what I used. Here's the first prompt I tried:

"I want to test whether you can act more like an agent than a chatbot.

My task is: Plan a weekend trip to Bath for two adults and two teenagers, including travel, lunch, one activity, estimated costs, and what still needs booking.

Don't just give me advice. First, make a brief plan. Then identify which parts of the task you can complete yourself right now, which parts require tools or information you don't have, and which parts need human judgment.

Then complete as much of the task as possible without stopping after the first obvious answer.

At the end, give me:

What you completed

What still needs human action

Any assumptions you made

A short checklist I can use to verify the result

The next best step"

What I really liked was that, as Claude tackled this task, it gave me the option to be notified when it had finished. In reality, it only took a few seconds to come back with a plan, which included travel options, an itinerary, and a suggestion for lunch and something to do: a trip to The Roman Baths.

To my delight Claude gave me an interactive map showing where all the places it recommended were. It also gave me a useful list of what it had completed, what required human action, the assumptions it had made, a verification checklist, and a "next best step" action point. It felt ready to keep working with me as more details came in, rather than treating its first answer as final.

In fact, when I gave it more details, such as which day I was going to go, it gave me a visual weather report for the day. That was a really nice touch.

Cladue Sonnet 5 maps.

Claude Sonnet 5 produced a handy map showing where to go. (Image credit: Anthropic)

Claude vs ChatGPT

I also tried this prompt with ChatGPT-5.5 Medium and got a similar result. It acted as an agent, just like Claude did, and notified me when it had finished its tasks. It just didn't look as nice. There was no map, or any visual elements at all, and it felt more like I had been given a finished report than the start of a two-way conversation where it asked me for more details.

Both chatbots recommended lunch and a trip to The Roman Baths. Interestingly, ChatGPT assumed I’d get the train, while Claude assumed I’d drive. They also recommended different places to eat, but the core information they both provided was solid.

What was most impressive was that both models could adapt when I reframed the inputs. For example, when I gave them the ages of the kids, student status, a different mode of transport, or changed the day of the trip, both models could cope. Both also identified that since the oldest was a university student, he could get free entry to The Roman Baths.

This part of the test was probably the most meaningful, as it felt much more "multi-step" than simply providing one answer.

Overall, I’d give this test to Claude. You can clearly see that Sonnet 5 is set up for agentic actions. Neither Claude nor ChatGPT could actually do any of the booking for me at the moment, so we're still a long way from true personal-assistant-level autonomy. But for this kind of task, Claude currently has the edge.

A different domain

I wanted to test the models in a different domain that would let Claude show me it had genuinely improved, and that the Bath trip result was not just a fluke of the travel-planning use case. So I asked them both to:

"Build me a simple household budget tracker as a spreadsheet or small tool."

Both models thought for a while about this task, and churned through various options before opting to make a spreadsheet. ChatGPT produced a spreadsheet with a bar chart that tracked how much I’d spent on various household expenses against a budget. Claude, however, went for something simpler: dispensing with a budget, it just tracked actual expenses and created a pie chart showing where my money was going.

Claude’s initial approach was simpler, and easier to understand. Both models provided a .xlsx file, but only Claude provided a button to upload it straight to Google Drive so I could open it in Sheets.

I told ChatGPT, "I wanted the graph to be a pie chart," and it responded: "Absolutely — I’ll update the spreadsheet itself so the dashboard uses a pie chart for spending by category, rather than the current graph style."

It ran into a few problems because it was trying to show both the budget and actual values in the same pie chart, but eventually it worked out that it could show only one and produced a new spreadsheet that did exactly what I asked for.

I then asked Claude to change its spreadsheet to provide a budget section too, and to change the graph into a bar chart. Again, it showed me its workings and added a budget section and bar charts perfectly.

I can’t really separate the two AI models on this task. Both proved they can handle multi-step tasks well, and both were happy to revise the result when I changed the brief.

That, really, is the point. The most interesting AI tests now are not "which chatbot gives the best answer?" They are "which assistant keeps working until the job is actually done?"

On that front, Claude Sonnet 5 feels extremely capable. ChatGPT was close behind, and in some ways just as effective, but Claude felt more naturally organized around the idea of completing work rather than simply responding to prompts. It asked fewer invisible questions, presented its output more helpfully, and made the whole process feel more like collaborating with an assistant than interrogating a chatbot.

For now, neither model is ready to fully take over the job. I still had to check the details, make the decisions, and do the actual booking or uploading myself. But the direction of travel is obvious. The AI war is no longer just about who has the smartest chatbot. It’s about who can build the assistant that gets you closest to a finished task.

I tried ChatGPT's new finance feature — and it opened a new window into how I spend my money

ChatGPT's new finance feature lets the AI chatbot take a look at any bank or similar accounts you care to open up for inspection. I was initially hesitant to try it out, but the tool only looks at the details of how you spend your money, and can't actually carry out transactions, so I agreed to let it analyze some of my accounts and offer its insights.

Finances is currently only available in the U.S. to Plus and Pro users on web, iOS, and Android. Setting everything up is as easy as using any other ChatGPT plug-in. You just select Finances in ChatGPT and then click Get Started then Connect with Plaid.

Finances uses Plaid to link to the accounts; you simply sign in and agree to let Plaid share the information. The service synchronized recent transactions and basically became an interactive, conversational database of my information. It's not that the AI can do things it couldn't before, but now you don't have to manually enter any finance details you want to ask it about.

ChatGPT looks at the receipts

At its suggestion, I asked ChatGPT, "How much did I spend eating out this year?" ChatGPT soon had a breakdown of my how much I had spent at various restaurants, including delivery, and broke down some of the patterns. It pointed out that it had gone up recently, which makes sense as the weather has became pleasant enough to walk into town more often to eat. It also identified a few spikes I could tell were around birthdays and family celebrations.

Next I asked, "What subscriptions am I paying for?"

That turned into one of the most practical moments of the experiment. ChatGPT gathered recurring charges across streaming services, software, and other monthly payments into one clear summary instead of leaving me hunting through statements. It also suggested reviewing whether every subscription still deserved a place in my budget.

"These subscriptions appear consistently each month, making them reliable recurring expenses.", it told me.

Financial chatter

Seeking a more detailed analysis, I asked ChatGPT, "What's changed in my spending this month?"

Instead of overwhelming me with every purchase, ChatGPT focused on the transactions that genuinely stood out. It noted a few large purchases at new locations, and asked if they were something I wanted more information on. I told it those were gifts, and the AI incorporated that fact into its response as not a new regular event but one that explained why overall spending had climbed this month.

"Those gift purchases are unusual compared with your recent history but they account for much of the increase, so your everyday spending has remained relatively consistent despite the higher overall total. Looking beyond one unusual month provides a more accurate picture of your longer term financial habits."

Using ChatGPT this way felt almost like the inverse of most financial software. Most budgeting apps expect you to learn their system first. ChatGPT let me ask questions however they occurred to me and then translated months of transactions into answers that were easy to understand.

That does not mean it replaces dedicated budgeting tools, and it certainly cannot make smarter financial decisions on your behalf. You still have to decide whether to cancel subscriptions, spend less on eating out, or save more each month. But I can see how it might make it easier for people who hesitate to look at their finances if they feel uncomfortable around spreadsheets. It didn't feel like math homework the way it often does.

Of course, it still relies on some trust in both OpenAI and Plaid, but as long as it's purely viewing and not actually touching the accounts, this could be a really useful, practical feature for ChatGPT users.

❌
❌