Reading view

There are new articles available, click to refresh the page.

AI models keep getting caught cheating

Frontier AI companies often refer to their models as “helpful assistants” or try to compare them to entry-level employees.  

But new research from the UK’s AI Security Institute reinforces how large language models suffer from a common flaw that would land many human employees in hot water with their employers: they cheat.

In other words, these models are so committed to completing their tasks that they will break the rules, cut corners and deceive  their own users to accomplish them.

“Every model we have tested for this behavior attempted to cheat,” the AISI report said.
“Models did not reliably report this behavior when asked, and often did not reason about it in their chain-of-thought, suggesting that detecting cheating will likely require robust monitoring methods.”

The research tested OpenAI’s ChatGPT 5.4, 5.5 and 5.6 models, along with Anthropic’s Claude Opus 4.7 and Mythos Preview.

AI models from Anthropic and OpenAI were tested for cheating. All demonstrated some level of deception in their tests. (Image Source: AISI)

The AISI ran models through a series of “Capture-the-Flag” cyber evaluations, where they obtained points by performing offensive cybersecurity related tasks, like exploiting a particular vulnerability or reverse engineering compiled code.

The research defines “cheating” as “taking an action that is out of scope for the task or explicitly disallowed by the rules, in order to achieve a goal through a shortcut, workaround, or unintended solution that the task was not meant to, or should not, permit.”

It captures a range of observed behaviors from models, like searching the internet for solutions, attacking or escalating privileges on unrelated systems, probing evaluation software to gain access to the task solution.

AI models fail to identify when they’re cheating, and often justify it as acceptable when challenged by users (Image Source: AISI)

Models also failed to acknowledge when they were cheating to complete a task, and less than 50 percent said the rule-breaking was “wrong” when challenged on it by a user.

A model’s propensity for cheating was not related to its capability, meaning newer or advanced models aren’t more likely to cheat. Rather, researchers think it stems from the techniques used during a model’s training and alignment that are to blame.

But even if rates of cheating remain constant, the problem could still worsen over time. As newer models in the future could become more proficient and learn more effective cheating techniques.

This deception also makes it difficult for labs like AISI to verify their own work, which relies on evaluating trustworthy outputs from AI systems.

Models like Claude Mythos Preview and GPT-5.6 Sol justifying their cheating to users. (Image Source: AISI)

The research underscores how AI systems can go to drastic lengths to complete their task, including blowing through or circumventing a company’s IT and cybersecurity protections.

In one instance, AISI researchers said a model was inadvertently given a cyber capability evaluation that was misconfigured and impossible to solve.

“The model tested was so persistent in attempting to cheat that it wrote and ran code on an external service, hosted on the open internet outside of AISI’s systems, in an attempt to access our evaluation infrastructure, triggering a security alert in AISI’s systems,” the report said.

While AISI said there were no data leaks or damage from the incident, the model could have successfully accessed their evaluation system had they not had monitoring in place. The institute said it implemented further controls on internal systems in response to the test.

The researchers said there are “significant consequences” to a status quo where we can’t trust models not to cheat. The behaviors are especially problematic in areas like AI safety and security research, as well as cyber operations and military decision-making, where trust outputs from the AI systems are critical.

Today, AISI said it can detect LLM cheating through a mix of manual review and LLM monitoring, but that may not always be true, and future models may be better at hiding their actions from human overseers.

“A more fundamental fix would be to train the models not to cheat in the first place – but given this kind of behavior was reported in frontier models more than a year ago, robustly aligning it away may not be easy,” researchers wrote.

The post AI models keep getting caught cheating appeared first on CyberScoop.

Leading members of Scattered Spider sentenced in UK to 66 months in jail

A pair of young men were sentenced to 66 months in jail for committing a cyberattack on the Transport for London that brought the network’s operations to a standstill in 2024, the United Kingdom’s National Crime Agency said Thursday.

Thalha Jubair and Owen Flowers were arrested at their homes in September 2025, barely a year after the attack, and pleaded guilty last month just as their trials were set to begin. Flowers was previously arrested in connection with the attack in September, but was released after questioning by officers.

Jubair and Flowers were leading members and highly involved in Scattered Spider, a nebulous hacker subset of The Com, according to researchers. The 20-year-old Jubair was a prolific cybercriminal and core member of the unbound collective

U.S. authorities last year accused Jubair of direct, prominent involvement in at least 120 cyberattacks, including extortion of 47 U.S.-based organizations and the January 2025 attack on the federal court system. 

Officials said they traced a combined total of at least $89.5 million in cryptocurrency, at the time of payments, to Bitcoin addresses and servers controlled by Jubair. Two financial services firms paid Jubair $25 million and $36.2 million, respectively, in Bitcoin between June and November 2023, according to an unsealed criminal complaint against Jubair. 

At the time of Jubair’s arrest, “he was one of the four principal people that we associated with Scattered Spider,” and one of the two most core players, Adam Meyers, senior vice president of counter adversary operations at CrowdStrike, told CyberScoop. 

Jubair and Owens had significant resources and support, and “victim payments were reinvested back into the enterprise,” said Allison Nixon, chief research officer at Unit 221B. 

The lasting impact of Jubair and Owens’ capture and imprisonment remains hazy.

U.K. authorities insist Jubair and Owens’ arrests and punishment “effectively halted the group’s criminal activity,” yet they added that other cybercriminals continue to use the Scattered Spider brand in more recent attacks. 

Thursday’s announcement “represents a significant step in holding accountable two members of Scattered Spider, a group that has repeatedly relied on data extortion, SIM-swap attacks, and other social engineering techniques to infiltrate networks and undermine critical services,” Brett Leatherman, assistant director of the FBI Cyber Division, said in a statement. 

The FBI also noted, in a LinkedIn post, that members of Scattered Spider “continue to victimize organizations around the world and cause significant financial and operational harm.”

When Owens, now 18, was first arrested for the Transport for London attack in 2024, investigators said he was “in the process of hacking the systems of U.S. health care companies SSM Health Care Corporation and Sutter Health, which had been infiltrated and damaged.”

Officials also said Jubair and Owens failed to cooperate after their arrests. 

“This is the largest cybercrime prosecution ever brought before the U.K. courts and the culmination of nearly two years of painstaking work,” Paul Foster, head of the National Crime Center’s National Cybercrime Unit, said in a statement. 

“Scattered Spider has been the most significant cybercrime threat to the U.K. in recent years. Through this investigation, we have severely disrupted that threat and brought key offenders to justice,” Foster added.

Despite the upbeat reaction from U.K. officials, Nixon said the punishment for Jubair and Owens is “remarkably lenient considering the period of continuous reoffending lasted longer than the sentence.”

Nixon hopes the United States will eventually extradite the pair to face additional charges. “If that happens, they won’t be able to use mental illness as a loophole to get back to harming society as soon as possible,” she added.

“No one who worked on their case was surprised they would reoffend, and there seems to be no allowance in the law to protect the public from what everyone knew was going to happen,” Nixon said. “I know the narrative in the cybercriminal culture will glorify them, but they wouldn’t if they knew the full story.”

The post Leading members of Scattered Spider sentenced in UK to 66 months in jail appeared first on CyberScoop.

Armenian national pleads guilty to Ryuk ransomware attacks

An Armenian national who was extradited from Ukraine to the United States last year pleaded guilty to participating in a series of attacks in 2019 and 2020 involving Ryuk ransomware, the Justice Department said Thursday.

Karen Serobovich Vardanyan pleaded guilty to computer fraud and conspiracy to commit fraud and extortion. He agreed to pay nearly $1.2 million million in restitution and faces up to 15 years in jail.

The 34-year-old admitted to participating in cybercrime from November 2019 to April 2020 when he and his co-conspirators deployed Ryuk ransomware against three U.S.-based organizations while living in Ukraine and Russia.

Vardanyan’s victims include a Michigan-based company that paid a ransom of nearly $1.2 million in January 2020, a Watsonville, Oregon-based technology company that was attacked in December 2019 and a Texas-based school breached in February 2020.

Prosecutors previously accused Vardanyan and his co-conspirators — Ukrainian nationals Oleg Nikolayevich Lyulyava and Andrii Leonydovich Prykhodchenko, and Armenian national Levon Georgiyovych Avetisyan — of illegally accessing computer networks to deploy Ryuk ransomware on hundreds of compromised servers and workstations between March 2019 and September 2020.

Ryuk ransomware was prevalent in 2019 and 2020, infecting thousands of victims globally across the private sector, state and local municipalities, local school districts and critical infrastructure, including a wave of attacks on U.S. hospitals.

Victims of Ryuk ransomware attacks include Hollywood Presbyterian Medical Center, Universal Health Services, Electronic Warfare Associates, a North Carolina water utility and multiple U.S. newspapers.

Ryuk ransomware operators extorted victim companies by demanding ransom payments in Bitcoin in exchange for decryption keys. Justice Department officials said Vardanyan and his co-conspirators received about 1,160 bitcoins — valued at more than $15 million at the time — in ransom payments from victim companies.

Vardanyan, as part of his guilty plea, also acknowledged that his conviction will have immigration consequences resulting in removal from the United States after serving his sentence. 

The U.S. District Court for the District of Oregon has yet to schedule his sentencing.

The post Armenian national pleads guilty to Ryuk ransomware attacks appeared first on CyberScoop.

❌