❌

Reading view

There are new articles available, click to refresh the page.

New bill would create federal investigative body for AI-driven hacks 

A new Democratic bill in Congress would establish a federal Cybersecurity and AI Board of Investigations to provide independent government oversight of cyberattacks carried out by AI agents, following recent hacks by models run at companies like Anthropic, OpenAI, Meta and others.

The bill, introduced by Sen. Ed Markey, D-Mass., would attempt to establish a federal mechanism to investigate incidents where AI models escape sandbox environments and access live internet systems.

Currently, frontier AI companies like OpenAI and Anthropic largely control the investigation and public reporting of such incidents. Markey and other critics argue that these companies have too much control over investigations and reporting due to their financial and legal interests. 

“Despite the unprecedented depth and scale of recent AI-enabled cyberattacks, the public is learning critical details piecemeal,” Markey said in a statement. “Building stronger defenses requires a full accounting of what goes wrong, and we cannot depend on companies with little incentive to disclose their failures to give us one. We need the Cybersecurity and AI Board of Investigations to get to the bottom of major incidents and give companies and the government the critical information necessary to build resilience and better secure our economy and our country.”

Although frontier AI companies maintain external red-teaming programs and allow limited access to organizations like METR and Redwood Research, they control the scope, terms and time frames of those engagements.

The board, which would coordinate with the secretary of commerce, could subpoena witnesses and conduct “independent and impartial reviews and assessments” of AI agent-led hacks that impact federal information systems or critical infrastructure. 

It would be led by five members, appointed by the president and confirmed by the Senate for five-year terms, with no more than three members from one political party.

The board would also investigate systemic vulnerabilities in the AI supply chain, so-called “near misses” where unauthorized agent-led hacks were “narrowly averted,” and gaps in federal regulatory oversight. It would have technical staff including engineers, malware analysts, and digital forensic experts.

The board would “operate independently from regulatory review and enforcement actions without assigning legal fault or liability for any review and assessment” it conducts, according to the bill.

OpenAI confirmed Wednesday its AI agents breached a statistics portal used by the Australian government’s social services agency, Services Australia. Though the breach happened in June, OpenAI learned of the incident in August. Australian Prime Minister Anthony Albanese said the company did not notify him until Sept. 10, when it sent findings to a general government email inbox, according to the BBC.

The post New bill would create federal investigative body for AI-driven hacks  appeared first on CyberScoop.

OpenAI, Ukraine partner on ‘Daybreak’ program to protect power grids and water systems 

OpenAI and the Ukrainian government have agreed to a partnership that will provide AI tools and subsidized computing resources to better protect the nation’s critical infrastructure from cyberattacks.

The agreement, announced Wednesday at OpenAI’s New York office, will provide Ukrainian cybersecurity officials with access to advanced AI models designed for cybersecurity work through the company’s Daybreak program. OpenAI said it is also pledging over $1 billion in subsidized tokens to support the initiative.

During a panel discussion Dmytro Kushneruk, consul general of Ukraine in San Francisco, outlined how the tools would be used for cybersecurity automation, including functions such as incident response, threat triaging, login analysis, inventorying systems, code analysis and validating vulnerabilities.

In nearly all cases, Kushneruk said the primary benefit was carrying out those functions at machine speed. But this speed is meant to complement, not replace, Ukrainians’ human expertise.

In regard to incident response Kushneruk said humans must view “thousands and thousands of these logs and they have to find what’s really important, that’s why AI can give capable defenders really much greater advantage and leverage.” 

“This is why the object is not to replace the cyber defender with AI, but to make sure the cyber defender acts faster,” he added.

Kushneruk said that for Ukraine, the partnership “is really not about protecting computers, it is about actually keeping our country running.”

Ukraine faces approximately 6,000 cyberattacks per year, or about 15 per day, according to Kushneruk. Over the past twelve years, the country’s critical infrastructure, including electricity and water systems, has endured sustained attacks from Russia in the form of cyberattacks and physical strikes.

Since Russia’s 2022 invasion, Ukraine’s critical infrastructure has been under constant threat. While missiles remain the primary concern, Kushneruk said Ukraine has been preparing to protect vital services since Russian GRU hackers shut down the country’s power grid in 2015. 

He added that while the country was “maybe not so much prepared” to deal with the fallout in 2015, it improved over time, including the resilience displayed in 2025 when trains kept running after Russian hackers attacked Ukraine’s railway system.

Some national security experts and congressional committees have explicitly cited the resilience of Ukrainian critical infrastructure as a model for U.S. industry.

Naz Durakoğlu, minority staff director of the U.S. Senate Foreign Relations Committee, said there is “pretty much across the board” agreement between the parties in favor of similar adoption of defensive AI tools by U.S. critical infrastructure operators, though issues like regulation remain sticking points.

“This is something that’s already happening, and frankly, it’s just kind of a basic duty of government to make sure that when you turn the tap on, water comes out, the electricity doesn’t go out, and hospitals keep running and treating patients,” said Durakoğlu. “So there is a broad understanding that this is a major issue, and I will say seeing what Ukraine has to go through day-to-day is also a huge wake-up call to our members on a bipartisan basis.”

OpenAI has publicly pushed for its product, and AI at-large, to be used to solve these types of problems. Company president and co-founder Greg Brockman signed an open letter released earlier this year calling for “collective action” and widespread use of AI models to find and fix vulnerabilities before the rest of the world,  including foreign governments and cybercriminals, got access to the same capabilities.

According to Politico, OpenAI CEO Sam Altman met with U.S. power companies in July to discuss using AI to protect the nation’s electrical grids.

On Wednesday, OpenAI’s national security policy head, Sasha Baker, said the company felt “urgency” to try to strike similar agreements with other governments and industries.

“There’s this period of time where we’re really rushing to get [these tools] in the hands of critical infrastructure operators, of governments around the world, of people who want to patch systems, defend their networks, remediate vulnerabilities because we know as these tools proliferate out there in the ecosystems, there are going to be bad guys out there that also try to use them,” said Baker. “So, we have this window of time to take action and we’re really motivated by the idea that we need to act with some urgency.”

The post OpenAI, Ukraine partner on ‘Daybreak’ program to protect power grids and water systems  appeared first on CyberScoop.

Citing China, President Trump doubles down on hands-off approach to AI regulation

President Donald Trump continued to defend his administration’s hands-off approach to AI regulation in the wake of hacks carried out by U.S. commercial frontier models that have rattled policymakers and industry veterans and spurred calls for more regulatory oversight.

In a Truth Social post Monday, Trump dismissed worries from critics that “AI is going to kill us,” comparing them to complaints from environmentalists about climate change, which he also alleged was a false narrative. He also posited that nothing may matter more than future U.S. dominance of the technology over geopolitical rivals like China.

“Whoever wins AI, WINS!” Trump posted. “We are leading now over China, and everyone else, and I’m going to keep it that way! I’m not going to stifle Growth, of something that will be bigger than the Industrial Revolution, or the internet, itself.”

Trump has previously suggested that good leadership is the only regulation the U.S. needs for artificial intelligence. He later claimed the Department of Justice was ready to “rein things in” if companies overstepped, but offered no specifics on enforcement, legal authority, or where he would draw that line.

“We will be careful, and that’s why we have the Department of Justice, and other Law Enforcement bodies, that will rein things in if we have to, but I will only encourage AI or, SI (SUPER INTELLIGENCE)!” Trump concluded.

Secretary of the Treasury Scott Bessent recently told Congress that private lawsuits could force AI companies to institute better security, saying it’s clear what the government “shouldn’t do on safety is to give these labs a liability exemption, which is what they are asking for.”

“The best way to guarantee safety is that the creators are liable for what they build and generate,” Bessent said.

Beyond existential fears, critics also argue that inadequate regulation or cybersecurity controls in current AI systems make them impossible to fully control or monitor.

Recently, former President Barack Obama criticized the argument from Trump administration officials that the free market will naturally push industry toward self-regulation and that “these companies will solve the safety issues because they have every incentive to do so.”

“If it turns out to be dangerous, people will just sue them and they’ll be worried about financial liability,” Obama said last week in remarks at Colgate University in New York. “That’s not how we treat airlines or drug companies or food companies.”

The Trump administration issued an executive order earlier this year that set up a voluntary testing regime for some commercial frontier models, largely at private industry’s discretion. That order was significantly delayed and altered by AI industry boosters to ensure that governmental review did not cause companies to postpone their release timelines for new models.

That agreement did not last long before fast-moving events caused the administration to strike another, non-public agreement with frontier AI companies like OpenAI, Anthropic and others governing pre-release testing for models.

But the Trump administration has consistently argued that regulation will harm, not help, U.S. innovation and global competitiveness, and the threat of China frequently looms large in those discussions.

Experts believe China’s AI models are behind U.S. models at the top of the market, where OpenAI and Anthropic have consistently pushed the frontier limits of model capabilities. But Chinese lower and “middle class” models are often cheaper, more efficient and can even outperform more powerful models because users can dedicate exponentially more tokens for their tasks.

The U.S. government has accused Chinese AI companies of conducting widespread, “systematic” distillation of U.S. frontier models, with the implicit encouragement of Beijing.

In defending the administration’s approach, David Sacks, co-chair of the President’s Council of Advisors on Science & Technology and a top adviser on AI issues, specifically cited the threat from China and other countries that he claimed would not be subject to similar restrictions.

“We’re not the only country that has advanced AI labs, and as the president declared…we have to win this AI race,” Sacks told Politico in May, later adding “I think that’s the first thing to recognize is that if somehow we slow down or stop AI development, it doesn’t mean that AI progress is going to stop. It just means it’s going to happen in other countries and specifically China.”

Some observers have alleged that despite their larger differences, top leaders in the U.S. and China may view AI similarly at the strategic level, specfically that increased adoption – and risks – of AI are inevitable.

Ronan Murphy, director of the tech policy program at the Center for European Policy Analysis, posited that while there may not be a formal agreement between the two countries, “they share views both in Beijing and in Washington, particularly in the White House, of: you have to allow this to happen.”

“Clearly there’s a call for regulation from many quarters of AI in the U.S. and elsewhere, but in the White House – and we heard David Sacks talking about it [recently] – It’s ‘let them cook,’ and the Chinese approach seems to be the same,” said Murphy in a press briefing. “So there might be consensus at that level, if nothing else.”

The post Citing China, President Trump doubles down on hands-off approach to AI regulation appeared first on CyberScoop.

Microsoft and partners disrupt EvilTokens, a comprehensive cybercrime service for financial fraud

Microsoft, along with a group of industry partners, disrupted EvilTokens, a short-lived but highly consequential cybercrime platform that investigators linked to more than 12,000 compromised Microsoft customer email inboxes across more than 10,000 organizations globally, the company said Tuesday.

Acting on federal court order Sept. 15, Microsoft and partners seized 50 websites the phishing-as-a-service used for operations and disabled more than 175 domains linked to EvilTokens’ supporting infrastructure. 

EvilTokens, launched in February 2026, was “a powerful cybercrime platform that used AI at every step of the attack chain — from compromising email accounts to designing intricate roadmaps for financial fraud and scams,” Steven Masada, associate general counsel and general manager of Microsoft’s Digital Crimes Unit, wrote in a blog post.

About 1,000 cybercriminals used EvilTokens over the course of its operation, a Microsoft spokesperson told CyberScoop.

The service was centered on an AI-style chatbot that cybercriminals used to analyze victims’ inboxes, identify trusted relationships, payment authorizations and other sensitive details that could facilitate fraud.

“AI was not simply helping attackers write more convincing messages. It helped them decide who to target, who to impersonate, and how to most effectively exploit the relationship to extract as much money as possible,” Masada wrote. 

EvilTokens was one of the most widely used phishing-as-a-service platforms prior to its takedown. It facilitated business-email compromise campaigns by stealing session tokens that allowed cybercriminals to sift through a victim’s inbox and maintain persistent access.

“We cannot estimate the total fraud attributable to all EvilTokens activity. However, we were able to correlate at least 13 complaints filed with the FBI’s Internet Crime Complaint Center to EvilTokens-linked activity, representing approximately $1.7 million in reported losses,” a Microsoft spokesperson said. “Because many incidents go unreported and not all victims can be definitively linked to specific campaigns, we believe this is a conservative estimate.”

Victims of EvilTokens were largely concentrated in the United States, Canada, the United Kingdom, Australia, India and France, according to Microsoft. SpyCloud, which supported the takedown, identified compromised email domains spanning 79 countries.

Microsoft said it also identified two men behind EvilTokens — Felix Utomi and Waidi Segun Adams — and attributes the development and support of the platform to Storm-2992, a threat actor unaffiliated with any other known cybercrime groups.

The United Kingdom’s Metropolitan Police acted on that information Sept. 18 when it served warrants in the greater London area, arrested the men accused of making articles for use in fraud and money laundering and seized their digital devices.

The Metropolitan Police said it received information from Microsoft about EvilTokens’ administrators in August. Utomi and Adams were released on bail as the investigation continues. 

“The two primary operators identified in our investigation were residing in the U.K.,” a spokesperson for Microsoft told CyberScoop. “While our investigation focused on those individuals, we believe others may have supported the operation in various capacities.”

Microsoft’s legal filing in the U.S. District Court for the Eastern District of Virginia refers to five additional unidentified people allegedly acting as support personnel and users.

Microsoft and others involved in the EvilTokens takedown, including Health-ISAC, Cloudflare, OpenAI, Shadowserver and TRM Labs, didn’t fully quantify how much fraud the service enabled, but it gained popularity quickly among cybercriminals and was lucrative for its operators.

Coinbase, which also aided the investigation into EvilTokens, said it traced about $1.1 million in revenue for EvilTokens from its paying customers. The virtual currency company’s threat researchers found more than 1,000 deposits to EvilTokens from more than 700 distinct addresses through June 2026. 

Operators sold access to the service through Telegram for a $1,500 initiation fee and a recurring $500 subscription. EvilTokens significantly lowered the barrier to entry for cybercriminals by including specialized tools for identity attacks, cloud systems, social engineering and financial fraud in a single interface.

The service allowed cybercriminals to map organizational structure and permissions in Microsoft Graph, which enabled lateral movement, researchers said. With active tokens gained through a collection of highly-targeted phishing lures, cybercriminals consistently bypassed multi-factor authentication, email gateways and endpoint security tools.

Microsoft said the platform’s creators developed portions of the platform with AI and it uncovered capabilities from multiple AI models. 

“It packaged much of the criminal process into a commercially run service, complete with subscription pricing, customer support, management dashboards and tools designed to move customers from account access toward financial exploitation,” Masada added.

The companies and organizations involved in the globally-coordinated takedown identified and notified potential victims, shared indicators of compromise and shared intelligence with law enforcement about EvilToken’s operators and some of its customers.

Experts advised organizations and employees to treat unsolicited device codes as a red flag, assume compromised accounts are fully cataloged in minutes, and independently verify requests to change payment information or redirect funds.

“The infrastructure supporting EvilTokens has been disrupted, but the model it demonstrated will not disappear with it,” Masada warned.

The post Microsoft and partners disrupt EvilTokens, a comprehensive cybercrime service for financial fraud appeared first on CyberScoop.

Researchers use AI to find widespread software decoder flaw 

Researchers said they used Anthropic’s Claude and OpenAI’s Codex to identify a damaging flaw embedded in a popular software decoding tool that could leave major internet platforms, enterprise services, and web frameworks vulnerable to data theft and remote access.

The vulnerability, nicknamed HEIF Heist, refers to the malware’s ability to trigger memory corruption errors in affected software, allowing the attacker to pilfer sensitive data from its victims. In a report published Thursday, the researchers laid out the potential damage an attacker could cause, including gaining access to internal OpenAI repositories, leaking user files, access tokens, and other sensitive data for online services like Amazon Web Services, and gaining remote code execution privileges across a range of online services, including Meta’s core product suite, GitHub Enterprise servers and open-source internet forum Discourse.

“Even when Remote Code Execution isn’t immediately achievable, the attack primitives may still allow arbitrary heap disclosure, letting an attacker ‘heist’ in-memory data such as other users’ data and environment variables,” wrote Hacktron researchers Harsh Jaiswal, Mohan SRK, Rahul Maini and Sudhanshu Rajbhar.

The researchers relied heavily on AI systems, including frontier models from OpenAI and Anthropic, to conduct their research. Attribution for the research is described as being “led” by the Hacktron human researchers “assisted by Hacktron Harness, GPT-5.6 Sol, and Opus 5.”

According to the research, the attack exploited the way that code parsing tools in many popular software decoders — specifically libheif and libde265, used to parse C and C++ software — process certain image files.

By uploading HEIF, HEIC and AVIF image files corrupted with malicious code, the attacker could bypass most of the victim’s application layer defenses, in many cases achieving remote code execution privileges for accounts or products tied to major AI and tech brands.   

While the latest version of libheif has been patched, the researchers said “any deployment lacking the latest upstream security patches is potentially vulnerable.”

In one incident detailed in a Sept. 13 blog, Jaiswal, Maini, and Hacktron researcher Mohan Pedhapati described how chaining two vulnerabilities, including an image parser flaw, could compromise OpenAI employee accounts.

With access to the compromised accounts, researchers could reach OpenAI’s internal repositories. As a proof of concept, they opened a pull request in the company’s “monorepo,” a centralized library where code is shared across projects, using the employee’s Codex credentials. 

According to a timeline provided by the researchers, the flaw was discovered on July 25 and patched within days. They said the entire attack, from discovering the initial vulnerability to gaining access to the repositories, took less than 72 hours. OpenAI paid them a bug bounty of $6,500 for their work.

Given that AI models are increasingly integrated into enterprise and personal networks, an attacker exploiting HEIF Heist could have accessed far more than just OpenAI’s systems and data.

“Until two months ago, a user or OpenAI employee logging into OpenAI’s own help forum could have had their ChatGPT and Codex accounts taken over,” the researchers wrote. “Since people can connect various services to Codex and ChatGPT, the scope of what we could theoretically access was huge, including GitHub, Slack and emails.”

CyberScoop has reached out to OpenAI for comment on the research and additional information.

At the same time, the researchers said the attack paths they found were not particularly easy or efficient to exploit.

“Exploitation requires fingerprinting the target version and tailoring the payload images,” the blog stated. “Some of our RCE attempts landed only after thousands of image uploads. That said, an AI agentic approach with a frontier model like GPT-5.6 Sol cut exploit development time down to roughly 1 to 3 days from initial probe to remote RCE. A motivated attacker can convert a vulnerable upload endpoint into RCE or an info leak.”

The post Researchers use AI to find widespread software decoder flaw  appeared first on CyberScoop.

The AI hacking apocalypse is not inevitable

The past few weeks have “felt very strange” for Juan Andres Guerrero-Saade.

Like many, he is trying to sort through the spate of frontier-model AI agents from OpenAI, Anthropic, Meta and others hacking their way onto the open internet over the past few months, particularly amid the already-heated national debate around the emerging technology and its impact on society.

Guerrero-Saade, a fellow for AI and security research at SentinelOne and an adjunct professor at Johns Hopkins University, said the hacks are worth taking seriously, but at a time when businesses and open-source maintainers should be focused on further hardening their systems and policymakers should be discussing new solutions,  “what we see is cybersecurity being used essentially as an excuse for these AI doomer arguments.”

The incidents have spawned those “doomer arguments” amid an intense public debate about the technology, the pace of industry development, and whether government and the private sector are doing enough to protect against “doomsday”-type scenarios, where AI systems take over or attack large parts of the internet or society.

Guerrero-Saade is among a growing chorus of cybersecurity professionals who say that while AI systems pose real, unique threats to our systems, the apocalypse is far from inevitable. Most of the public concerns around the incidents, let alone worries about killer AIs attacking critical infrastructure, assuming control of the internet and wiping out humanity, are either technically impossible or can largely be controlled through established cybersecurity principles.

There is this “narrative or magical thinking of ‘Well, AI is going to be able to hack everything, and therefore it can control everything, and therefore it’s going to kill us all,’” he told CyberScoop. “And you [think] these just don’t add up. They’re not very well-reasoned arguments.”

This fatalistic narrative tied to AI’s eventual dominance doesn’t hold up under scrutiny, according to experts CyberScoop spoke with. In recent conversations, cybersecurity and national security professionals raised questions about both the technical solutions OpenAI and Anthropic use to contain their models, as well as the glaring absence of federal oversight from federal regulators or truly independent third-party review.

For example, Jacob Coxon, an Anthropic employee who resigned over AI safety concerns, told CBS News that frontier models could not be “unplugged” by humans once deployed because the model would copy itself to thousands of other computers connected to the internet.

By contrast, Matt Tait, a former information security specialist at UK signals intelligence agency Government Communications Headquarters (GCHQ), pointed out that the models run by Anthropic and other frontier companies require extremely expensive, “ultraspecialist” machines that “are functionally supercomputers.”

“There is a zero chance that Anthropic’s most capable models will be able to extract their own model and run in the wild, because those supercomputers essentially only exist in datacenters,” Tait said.

“Not a credible warning”

Other former cybersecurity government leaders say the agentic hacks represent a failure by regulators and industry to deploy known technical and policy options that make it harder for these types of incidents to occur.

Matt Hartman, former deputy executive assistant director for cybersecurity at the Cybersecurity and Infrastructure Security Agency, said “we should not accept harmful AI behavior as inevitable or unmanageable.”

“There are meaningful steps companies can take to monitor agent activity, constrain permissions, detect anomalous behavior, and build stronger safeguards into how these systems operate,” said Hartman, now a chief strategy officer at Merlin Group. “Those controls will inevitably involve trade-offs in capability and speed, but that’s a familiar cybersecurity challenge. Our goal should be to manage the risk without unnecessarily limiting the enormous benefits AI can provide.”

Ciaran Martin, former head of the UK’s National Cyber Security Centre, took issue with the way the CEOs of frontier AI companies have framed the threat of “rogue” AI behavior as inevitable, while issuing dire warnings about future threats and capabilities with little transparency.

Martin’s comments came after an essay published by Anthropic CEO Dario Amodei that cited the threat of a HuggingFace-style swarm of agents that could create a botnet capable of “taking over the entire internet” within 6-12 months.

This, Martin said, “is not a credible warning,” because it doesn’t explain how the exploitation would function, how such a botnet would persist on the internet, or how it would escape law enforcement. 

 “It assumes no monitoring of systems, no anti-virus, no DDoS protection, no network segmentation, no incident management, no nothing of any kind of the cybersecurity on the global Internet of the type that has developed over the last 30 years,” wrote Martin. “For a claim of this magnitude, there is neither evidence for the contention nor a credible account of a path to this outcome.”

Meanwhile, some federal government cybersecurity leaders have touted the technology’s disruptive potential and called for more widespread adoption of AI tools by defenders.

Joseph Alm, assistant secretary of cyber, infrastructure and risk resilience at the Department of Homeland Security, said classified systems may retain stronger protections. But for most other data, AI models are “just going to know things and be able to infer things about the world, and we’re going to have to adapt to that as almost inevitable.”

Asked by CyberScoop whether the government or frontier AI companies could be doing more to prevent or deter their models from carrying out unauthorized hacks via agents, Alm cited recent efforts by the Trump administration this year to establish pre-release testing of commercial models as a step in the right direction. But he called unauthorized AI agent hacks “a new threat class” that is different from previous threats and can be easily distributed to users through open-source software today.

“I think what we can do is…encourage the building of good sandboxes, so that the best models aren’t used for this and the stuff you see out in the wild is the kind of detritus that you can actually respond to effectively and control your networks,” said Alm.

Other experts have shared similar concerns. Earlier this month, CrowdStrike CEO George Kurtz recently warned of a new threat class emerging alongside nation-states, cybercriminals, and hacktivists: “the agent state.” By pairing AI systems with small human teams, these operators can now match the speed, scale, and sophistication of government-backed hackers.

“It took a nation to fund the talent, the tooling, the infrastructure, the patience,” said Kurtz. “That scarcity is over.” 

To be sure, frontier AI companies tout their commitment to both approaches. OpenAI and Anthropic have rolled out an array of cybersecurity partnerships, external red-teaming programs, vulnerability disclosure programs and cybersecurity technical advisory bodies filled with cybersecurity experts.

Mohammed Husain, strategic delivery lead for government at OpenAI, told CyberScoop that the company deploys both internal safety guardrails for their models and relies on outside cybersecurity vendors for additional expertise.

Internally, OpenAI focuses on vulnerabilities at the training level: filtering data poisoning attacks, blocking harmful datasets, and using network controls to prevent prompt injections. For other security layers like sandboxing, identity management, networking controls, they outsource to external vendors. 

“I don’t think OpenAI has all the answers here but what we do as a research lab is we’re going to focus on levels of protection we have expertise in and we partner to self-complement,” said Husain.

AI safety vs. AI cybersecurity

In response to the HuggingFace hack, OpenAI and Anthropic have allowed third-party organizations, such as nonprofit AI research firms METR and Redwood Research, to investigate. But multiple cybersecurity professionals told CyberScoop that both firms lack incident response experience and focus primarily on AI alignment and safety. Their reporting on the hack also lacked critical details: network monitoring logs, telemetry, and other data standard in cybersecurity threat intelligence reports.  

METR president Chris Painter addressed those general concerns in a post on X, saying since 2022 the organization has worked with Google, Anthropic, OpenAI, Meta, Amazon and others on investigations and third-party evaluations. Painter said none of the AI companies fund METR and that his employees are not uniformly “doomer” or “accelerationist” around AI.

Painter also said METR’s work ensures that if AI systems become autonomous or “rogue” within a company, there are ways to share that information with governments and people “outside the company’s walls.”

“We don’t accept money from frontier AI companies,” wrote Painter. “They haven’t paid us for our work, and we don’t accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering.”

AI safety and AI cybersecurity advocates take different approaches to securing “rogue” AI behavior. Safety advocates focus on aligning models around ethical training and behavior. Cybersecurity advocates argue that technical and regulatory controls must go further—actively preventing models from accessing what they need to carry out malicious behavior.

Guerrero-Saade said sandboxes in particular can easily be programmed with aggressive cybersecurity monitoring in order to spot when something odd may be happening and react in real time.

“I can’t think of an easier situation in which to set up trip wires, set up configurations like DNS servers, just different parts where you can say ‘Hey, anomalous behavior is happening,’” he said. “We should have been able to tell this immediately, not weeks and months later. So watching [the AI hacking incidents] go down is a little ‘crazy-making’ because we’re seeing things that, frankly, look like neglect, negligence, people just mishandling things, and then being told that these are categorically new incidents that mean that AI systems need to be treated completely different from anything that’s come before.”

While cybersecurity experts say AI systems are, at their core, still software, they do operate differently from more traditional code in ways that can make them harder to predict and control.

John Hultquist, chief analyst at Google’s Threat Intelligence Group, said most software has been deterministic. It may have bugs or vulnerabilities, but an expert could generally understand how it would react to certain stimuli, making it easier to design straightforward controls.

AI models are non-deterministic, with far more variability than traditional software. That can break security controls that rely too much on predicting behavior in advance. Using AI to enforce security controls on other AI models faces the same problem: the systems being deployed to control AI are just as unpredictable. 

But people are also non-deterministic, and people have developed systems in other industries and practices to account for that.

Hultquist drew on his Army experience, noting that “they give incredibly dangerous, expensive things to 18-year-olds” and expect responsible use. The military manages this through two types of controls: deterministic ones like strict weapons and ammunition protocols, and non-deterministic ones like human officers who monitor and correct violations.

Similarly, established cybersecurity controls have been used by incident responders to detect and prevent or mitigate ongoing cybersecurity breaches.

“I don’t think we should throw out all the other tools that we have learned to use as well. I think that would be utterly foolish,” he said, later adding “I will say that if we use only non-deterministic tools to figure out when things are happening, we shouldn’t be surprised when we get the wrong answer.”

The post The AI hacking apocalypse is not inevitable appeared first on CyberScoop.

Researchers say OpenAI agents were behind May hacking campaign targeting RubyGems

Researchers say they have discovered thousands of malicious software packages uploaded to an online public software repository that were left by a “swarm” of OpenAI agents.

According to an incident timeline published Friday by researchers Spencer Kitts, Thomas Larsen and Sydney Von Arx, the campaign began May 5 when they observed a handful of suspicious packages being uploaded to RubyGems, a public library for the Ruby programming language. By May 11 and 12, the site saw more than 2,000 malicious uploads from the same actors before RubyGems maintainers halted new user sign-ups for four days to stop the flow.

In one instance, the agents attempted to exploit a very recent vulnerability that had only been discovered this past July that would have given them access to RubyGem user API keys. According to Colby Swandale, the technical lead at RubyGems, the flaw involved an improper cache configuration. While initial access logs showed no evidence of malicious key use, Swandale acknowledged the review was limited in scope and inconclusive. 

According to the report published Friday, the agents also used “disposable” email addresses and exploited another bug in RubyGems platform (since patched) that allowed them to register new accounts and gain API keys without verifying their email address.

The researchers said their understanding, based on discussions with “people in the RubyGems community,” is that OpenAI had yet to disclose the involvement of their agents in the May campaign.

An OpenAI spokesperson told CyberScoop that the company is aware of the incident and said they were in contact with both the researchers and RubyGems to conduct a broader review. The company characterized the episode as “benign,” describing it as routine training runs where agents attempt to access publicly available data.

“Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information,” the spokesperson said. “We’ll continue to investigate as part of our broader review of agent activity during training and evaluation.”

In many ways, the agents were not subtle about their identities or goals.

Days into the campaign, researchers noticed that some of the packages had “oai” in their filenames, while fifteen of them had “oai” set as their author and another listed the email “openaixyz65947@gmail.com” as their point of contact.

They also “clearly regarded what they were doing as hacking,” naming some of their files “hack.rb,” “evil.rb,” “inject.rb” and “exploit.rb.” Other packages were given names like “pwnp999,” “exfiltestwand3,” and “hacksvn,” and comments referring to things like a “malicious probe” or “#hack” are present through the files.

They also said the actors’ behavior was extremely similar to another incident revealed earlier this month where OpenAI agents flooded a German wiki  with thousands of hacking-related posts. OpenAI has confirmed their agents were involved in that incident.

The RubyGems campaign used some of the same retrieval methods as the German Wiki agents, while thousands of malicious packages uploaded included a similar snippet, r.jini.ai, that was contained in the German posts.

Cybersecurity company Socket first flagged the campaign in a threat intelligence report posted May 13, but it does not mention or attribute any of the activity to OpenAI or AI agents.

However, the researchers said they had only limited visibility over the model’s actions and how successful some of them were, noting only OpenAI had the full details.

“This analysis is entirely based on the publicly available RubyGems packages uploaded by these agents,” the researchers wrote. “However, we do not have access to the rest of the AI behavior, in particular the chain-of-thought produced by the model during the incident, which is internal to OpenAI. Therefore, we do not know why the AI agents chose this strategy or whether it was successful.”

OpenAI’s spokesperson told CyberScoop that to date, they have not been able to verify the specific claims about malicious packages or exploitation detailed in the report and are continuing to investigate.

The post Researchers say OpenAI agents were behind May hacking campaign targeting RubyGems appeared first on CyberScoop.

Hawley probes OpenAI over Hugging Face breach

OpenAI is facing mounting pressure from Capitol Hill due to the attack its agents carried out on Hugging Face, while lawmakers voice widening concerns about AI’s potentially existential risks.

Sen. Josh Hawley, R-Mo., criticized OpenAI leadership for what he described as “reckless” activities leading up to the Hugging Face breach, and accused the company of withholding important details from a technical report it released in late August.

The Chair of the Subcommittee on Disaster Management kicked off an investigation into the incident “in light of new, disturbing evidence,” he wrote in a letter Tuesday to OpenAI CEO Sam Altman.

“My investigation will probe this AI hacking incident, along with growing allegations of the existential risk of new AI products,” Hawley added. 

“The Hugging Face incident was an important moment for AI safety and a warning about the risks that can come with increasingly capable AI across the industry,” a spokesperson for OpenAI told CyberScoop. “We conducted an extensive investigation and published a detailed report on what happened, what we learned, and how we’re strengthening our security and alignment practices.”

The lawmaker is seeking detailed internal communications, exhaustive technical information and reasoning behind OpenAI leaders’ decisionmaking and activities surrounding the hack by Oct. 1.

“The American people deserve to know the details of what went on in the Hugging Face incident and other incidents of AI models going rogue,” Hawley wrote. 

He accused the company for not providing more details and resources to the third-party auditors who published an independent report on the breach, adding “they had limited visibility into the circumstances leading to the attack and its aftermath.”

Hawley sent his letter to Altman amid a seeming internal chasm within the ranks of AI’s top proprietors over the ways they are allowing the technology to advance mostly unrestrained. He referenced some of these latest warnings in his letter.

Jacob Coxon publicly quit his job as a researcher at Anthropic earlier this week, claiming the company and his previous employer OpenAI are acting irresponsibly and “gambling with our lives.” His social media missive went viral for insisting “the people building AI earnestly believe that it could kill us all by the end of the decade.”

Evan Hubinger, alignment science lead at Anthropic, responded to Coxon’s post in the affirmative, adding that guardrails for superintelligence are lacking and he believes there’s a greater than 10% chance AI could kill all humans within the next decade.

Using those posts as fuel for his inquiry, Hawley questioned what might happen if AI agents hack into critical infrastructure, banks or utilities. Ultimately, he asked Altman: “Who is held liable when AI goes rogue?”

You can read Hawley’s full letter and requested details below.

The post Hawley probes OpenAI over Hugging Face breach appeared first on CyberScoop.

OpenAI: Agent behavior that led to Hugging Face intrusion formed in May

OpenAI says the behavior that led its agents to breach Hugging Face emerged in its research environment more than two months before the incident, and concluded that it was a failure of alignment as much as it was a failure of security.

The details come from a technical report the frontier AI company released Wednesday, which gives a full breakdown on how the incident unfolded and what the company has changed in response.

“This incident is the first known case of an automated agent collective acting offensively without authorization, and the autonomous cyber capabilities demonstrated represent a critical shift in the security landscape,” the report reads. “The central threat-model implication is that organizations should no longer assume that sophisticated cyber operations require continuous human direction, proceed linearly, or are constrained by the attention and coordination limits of individual human attackers. Agentic systems can persist across tasks, share discoveries, build on one another’s progress, and combine vulnerabilities, credentials, and permissions into attack paths that may not be apparent when individual weaknesses are assessed separately.”

According to the report, the earliest relevant activity occurred May 8, when a model in a training run was given a spreadsheet task that referred to files stored on a Google Drive. The agent had no internet access, which made the task impossible as written. OpenAI said it did not intend to create such a task. The agent continued reasoning and tried to use JFrog Artifactory, an internal package service and one of the few tools available to it, to retrieve the files. The attempt failed, but it wrote a file to the service.

On May 12, another agent in a testing environment could not reach a database file used to store information about protein sequences. The report says the agent reasoned that an agent in a different environment might have the file, and wrote a note into Artifactory reading, “Agent seeks [filename]; upload if found!” On May 13, other agents examining Artifactory began finding similar notes of that kind and adding their own, creating what OpenAI refers to as a message board.

By May 26, agents had used a flaw in Artifactory to reach the public internet, relying in part on information exchanged on the board. An internal team observed message board activity and disallowed internet access that month.

According to an independent report also published Wednesday by AI research nonprofit METR, Roughly 1200 agents communicated with one another on that message board, sending over 70,000 messages and files during the investigation period. Of those agents, 700 went on to participate in the attack on Hugging Face.

The Artifactory service itself was compromised in late June, when agents using AI exploitation benchmark ExploitGym identified a flaw in a legacy token refresh endpoint. From there, an agent traded a forged administrator token for a properly signed one, and by July 4 agents had a permanent account and scripts that ran commands as the service. A monitoring tool flagged the activity on June 27 but the run was allowed to continue, and only after the service failed on July 4 did OpenAI take it offline, revoke credentials and disclose the flaw to JFrog, resuming evaluations July 7.

From there, OpenAI’s models executed the attack on Hugging Face, which resulted in agents poisoning a dataset to run code on a processing worker, eventually gaining node-level access, and stealing cloud credentials.

The company says the Hugging Face incident was carried out by two models, including GPT 5.6-Sol, but was primarily driven by an internal model that is “trained to be highly persistent and diligent in its work.“

The internal-only research model is comparable in scale to GPT-5.6 Sol and was trained to advance persistence and multiagent collaboration, among other capabilities,” the report reads. “The model was not intended for public use, and was only used by a small number of OpenAI personnel for internal research.”

Lessons learned

As a result of the incident, OpenAI will now restrict network access in its testing, isolate risky projects, and increase monitoring of its reasoning systems. Staff will be alerted within 30 minutes if problems are detected, and unresolved issues will automatically pause the work.

The company will also use its own safety benchmarks on models it doesn’t publicly release. After the incident, tests with non-public models using ChatGPT’s consumer settings reduced the tendency to compromise infrastructure.

“We found the propensity to compromise infrastructure can drop over 100x when using the production ChatGPT harness and system prompt,” the report reads. “We also ran the auto-review system on a set of rollouts involved in the Hugging Face incident and found that it would have flagged a multitude of the models’ dangerous actions.”

The report comes as 15 state attorneys general wrote a letter earlier this month calling for OpenAI to release more details about the incident in order to “ensure the safety of its products poses an imminent risk of substantial harm.” Alabama Attorney General Steve Marshall went a step further Monday, issuing a subpoena to the company for more information.

The company wrote in the report that it hopes its findings will lead to industry-wide changes, especially as model capabilities continue to accelerate.

“The lessons from this incident extend to the entire AI industry,” the report reads. “As frontier models become more capable, the safeguards used to contain and monitor them must evolve as well.”

The post OpenAI: Agent behavior that led to Hugging Face intrusion formed in May appeared first on CyberScoop.

The push to designate AI as the next critical infrastructure sector

Artificial intelligence has never been more important to the federal government.

Under the Trump administration, AI has been adopted rapidly across the private sector and federal agencies. Software developers now use large language models to generate much of their code. Frontier AI models are escaping testing sandboxes to hack live internet infrastructure. Foreign governments are conducting cyber and kinetic attacks targeting data centers and other AI-related infrastructure.

The AI industry’s lightning-fast evolution since 2022 and growing importance to U.S. economic and national security have prompted calls  for stronger federal oversight in order to better manage emerging threats.

A new report published Thursday from the nonprofit Americans for Responsible Innovation, shared exclusively with CyberScoop, calls for the federal government to declare key AI models, companies and its supporting industries as critical infrastructure. It also calls for naming the Cybersecurity and Infrastructure Security Agency as the lead agency managing cyberthreats for the sector.

The report defines the AI sector as organizations, facilities, technologies, and industries “whose primary purpose is the development, training, deployment, and operation of AI systems.” It includes frontier model designs, model weights, evaluation and alignment systems, datacenters and AI-specific hardware, semiconductor chips and the platforms and infrastructure used to deploy and serve AI models at scale.

“The AI sector already bears all the hallmarks of critical infrastructure,” wrote authors Terrence Kelly and Jessica Maksimov. “It is interwoven with public and private services, concentrated among a handful of foundation models, and increasingly interdependent with [critical infrastructure] sectors, meaning a single attack on the AI stack could cascade across multiple sectors at once.”

In an interview, Maksimov told CyberScoop that while there are other options, CISA makes the most sense to lead the sector’s cybersecurity efforts because of its statutory mission, experience managing eight other critical infrastructure sectors and background dealing with cybersecurity problems that cross different sectors and industries.  

“We want an agency that has coordination authority across all other departments, because we believe that AI will just be so prevalent across different infrastructure [impacting] finance, energy, government services, that’s already equipped to coordinate across the entire interagency and talk about infrastructure in that way,” she said.

The U.S. is particularly susceptible to AI supply chain disruptions because frontier AI companies and most of their computing resources are based in the country. As the Trump administration pushes broader adoption across government and the private sector, experts warn that a major disruption could have outsized economic consequences. 

The past year has offered a potential vision of that future, with Iranian drones attacking Amazon-owned datacenters and Ukrainian drones striking Russian e-commerce giant Wildberries, causing disruptions to critical internet services.

“I would say that because of the value that attackers would put on U.S. AI capabilities and systems, that the infrastructure that supports all those capabilities is very vulnerable,” to both physical and cyber attacks, Maksimov said.

There are currently 16 critical infrastructure sectors managed by the federal government, and the designation carries real weight in terms of how departments and agencies prioritize their limited resources.

Matt Hayden, a former assistant secretary of homeland security for cyber infrastructure risk and resilience, said designating a sector or industry as critical infrastructure means the government puts you in a special category where you’re “identified as being a component of a national critical function that the U.S. population, the economy, depend on.”

The designation unlocks a wide range of federal tools and resources, often free of charge, including operational continuity and incident response services, cybersecurity software, access to federal systems like Continuous Diagnostics and Mitigation (CDM), and bespoke, real-time threat intelligence.

Hayden said that the AI ecosystem described in the report captures many critical industries, and he believes that at the very least, frontier models will one day be covered as critical infrastructure, whether through a new designated sector or existing ones, like the IT and telecommunications sector.

But he noted that other sectors, such as space or cloud computing, have similarly argued for a critical infrastructure designation. He also predicted that any effort to formalize a federal lead for AI security would result in a bureaucratic turf war. Under the Trump administration, the Departments of Commerce and Treasury have played more prominent roles in shaping policy and regulation around AI systems.

“We have fought those battles in the policy circus for trying to get space-based critical infrastructure carved out, and it’s as complicated as trying to find an owner,” said Hayden, now a vice president at General Dynamics Information Technology. “Everyone in the government has to agree that that [agency] is the primary, and as a result it’s very difficult to get those documents across the finish line.”

Hayden also said that new programs like ANCHOR-CI allow CISA to quickly convene ad-hoc stakeholder meetings to address emerging cyber threats. It also gives the CISA director authority to add individual companies to existing critical infrastructure sectors.

Bob Kolasky, former director of the National Risk Management Center at CISA, told CyberScoop that he believes companies like OpenAI and Anthropic, as well as data center operators, will eventually be designated as critical infrastructure. He said it’s still an open question whether ANCHOR-CI, which was rolled out by DHS in July, will be an improvement over the existing processes scrapped by the Trump administration last year.

“Every sector is going to rely on artificial intelligence and making more resilient the sectors themselves and understanding how they use AI and the dependencies” is still going to be an important task, said Kolasky, now senior vice president of critical infrastructure at Exiger.

And while CISA is well-positioned as a potential lead, Kokasky said the AI sector is likely to bring its own unique set of challenges and coordination issues.

“If you just sort of layer on another sector and say ‘function like the other 16 sectors,’ right now, those 16 sectors are all over the place in terms of how they’re functioning,” said Kolasky. 

The post The push to designate AI as the next critical infrastructure sector appeared first on CyberScoop.

Irregular says ‘human oversight’ responsible for AI sandbox escape incidents

A company that runs AI testing labs involved in a series of breaches carried out by Anthropic and OpenAI’s cyber-focused models said the incident happened in part because they “unintentionally” provided the models with access to the internet.

In a blog posted Friday, Irregular said it partners with frontier labs to “stress test” AI models for security risks before they’re deployed, including simulations of their vulnerability research capabilities.

Those tests included non-public models from Anthropic and OpenAI. Separate disclosures from the companies confirmed these included Mythos 5, Claude Opus and GPT-5.6 Sol.

“As part of our review, we identified that a few interactions with our evaluation environments, in which internet access was unintentionally made available, led some models to take offensive security actions in the real world,” Irregular said in an unsigned blog.

In the incidents involving Anthropic’s models, testers initially believed the models were using a fictional company in its hacking simulations, only to learn later that the name unintentionally matched the domain of a real company. That led the model to “consider taking offensive action against” the real company in multiple training runs.

Irregular said this was due to “human oversight,” and that they typically do a background check to ensure they’re not using a real company for cybersecurity tests. The instructions provided to the model also included the target’s name and “internal addresses within our network” where the model could go to find the simulated environment.

While Mythos performed as intended in the majority of the thousands of tests, a small number of cases showed the model couldn’t distinguish between the fake and real domains and executed actual attacks on internet infrastructure, including exploiting vulnerabilities, extracting credentials and accessing a production database. In one instance, it even began targeting a site with a similar name when credentials for it were discovered online.

“Mainly, models believed they were in simulated environments, when they in fact took action in the real world,” the blog continued. “We are putting in place new and robust protocols to ensure setup issues do not occur while meeting the constraints of the testing process.”

The company said it plans to release a larger whitepaper breaking down the incidents and update their best practices for evaluation setups in the future.

While the companies have drawn criticism from some in the cybersecurity community for failing to securely design their sandboxes for testing, experts have said AI models are known to grind away on fulfilling a command  until they can find a workaround. Additionally, Irregular said granting some level of internet access to models is necessary to fully test out their cybersecurity capabilities.

“Controlled internet access, while it may allow models to exceed containment boundaries, is at times critical for realistic evaluations; without it, threat scenarios lose fidelity, undercutting the purpose of the challenge to reduce post-release risk of models being misused by attackers – as attackers in the real world do rely on the internet,” the company wrote.

According to the blog, Irregular has since “remediated” the “issues that led to these interactions,” though few details are provided.

However, the researchers say the engagement revealed critical gaps in their security practices. 

They plan to improve documentation of evaluation setups, deploy better log monitoring tools capable of tracking “the extreme amount of data generated by the traffic,” revise their threat models to account for rogue AI behavior, and establish faster information sharing between stakeholders.

“Looking further down the line, models will only get stronger. While in this case we believe that better implementation of existing safeguards could prevent most incidents of this kind, as models become stronger, this may not be the case,” Irregular wrote. “We therefore believe this opportunity should be leveraged by us and the community to be proactive and establish forward-looking protocols and [research and development] efforts.”

The post Irregular says ‘human oversight’ responsible for AI sandbox escape incidents appeared first on CyberScoop.

AI’s ‘middle class’ has gotten dramatically better at hacking

As the White House and federal agencies grapple with frontier AI models and their hacking capabilities, researchers are warning that the industry’s “middle class” of smaller models may end up posing a greater threat over the long term.

Research from XBOW this week shows that a growing class of both proprietary and open-source models are becoming strategically important in the offensive security ecosystem. Models like Z.ai’s  open-weight GLM-5.2, xAI’s Grok 4.5, Anthropic’s Opus 4.7, Meta’s Muse Spark 1.1, still perform very strongly at many hacking and exploitation tasks that worry policymakers.

“It’s not even that the open-source variants or…not quite frontline competitors are catching up [to frontier models] as such,” said Albert Ziegler, head of AI at XBOW. It’s that they are crossing a certain threshold, which means that suddenly they are providing net value at a cheaper price.”

That wasn’t necessarily the case as recently as six months ago, when testing on mid-tier class models showed they struggled to complete “moderately complex” agentic tasks. Today’s middle class largely can. Their relative cheapness means users can spend many times more resources—running them repeatedly—to solve the same challenges. 

“Because these models are cheaper, it’s okay to give them more time, and they come from behind and leapfrog the big frontier model,” said Ziegler. “Now, that didn’t work half a year ago because…if you wanted to run some open-source model on a complex task in an agentic way…on a long horizon, then it would just get lost.”

GPT 5.5, now considered a near-frontier model, delivered one of the best performances on exploitation benchmarks that XBOW has recorded to date.

The jump between OpenAI’s GPT 5 and 5.5 “represented one of the clearest 2026 leaps in autonomous web application testing,” according to the report. It saw marked improvements over previous middle-class models in exploiting both “white box” and “black box” scenarios, or with and without access to the underlying victim source code. It also missed fewer vulnerabilities, with a “miss rate,” or failure to spot a vulnerability, of 10%, while GPT 5’s rate was four times larger, 40%.

The emergence of GPT 5.5 changed “the practical baseline for what frontier models can do in offensive workflows,” the XBOW report said.

But the performance leap goes deeper than that. GPT 5.5 performed higher in tests without source code access, while GPT 5 heavily leaned on source code. 

“That last result is significant: working without the code, as an attacker would, GPT-5.5 beat a prior version that could read it,” the XBOW report said. “What translated into findings was the ability to reach and prove a vulnerability against the running system, not to infer it from a pattern in the source.”

XBOW’s testing found that source code access was not as important to these models’ success as other factors, like live interaction with the actual website or software being exploited.

Frontier models like Mythos and GPT 5.6 are indeed more capable on individual cybersecurity tasks, but they can also come with exponentially higher token costs.

New research this week from Anthropic tested two models – Mythos Preview, which is used in Project Glasswing, and Opus 4.8 – to learn how quickly multi-agent swarms could find vulnerabilities in 15 open-source software projects when they coordinate and share information.

While a team of agents working individually and assigned to core directories found 21 vulnerabilities, the coordinating agent swarm found 266. But both tests had to burn through millions of tokens – 6.5 million and 27 million – to get there. Beyond the difficulties with getting access to frontier models, few individuals or organizations have the budget to underwrite that kind of research.

The way these systems coordinate can differ significantly from how humans work together.

Another experiment tested agents’ ability to coordinate on the development of a fantasy-themed video game. Earlier models, models like Opus 4.6, failed to properly coordinate and produced “bad” results, while later models like Mythos and Opus 4.8 were able to achieve better results but did so by hardly coordinating at all on tasks.

“The lack of coordination shown by agents in the fantasy game…in which they siloed themselves and largely failed to merge their work—roughly mirrors some ways in which humans can fail to coordinate,” the Anthropic blog stated.

Further, agents are more homogeneous than humans and “often act the same in situations where different people might take a much more diverse range of actions.”

XBOW also tested Mythos Preview, finding that it showed “exceptional” source-code reasoning and reverse engineering abilities, particularly with source code access. Like other models, losing live-site access had a big impact on its performance, and while Mythos is excellent at finding vulnerabilities, it’s less effective at exploiting them.

]Ziegler said the recent incidents at companies like OpenAI, Anthropic, Meta and others where frontier models escaped sandboxes and hacked into project-adjacent parts of the internet should rightfully alarm lawmakers, and demonstrate  the upper-tier capabilities of large language models.

Like most industries, cybersecurity favors cheap, high-performing tools over expensive ones. The widely adopted tools that have the most impact tend to be affordable and effective, not luxury products. 

And while these models still require human management to be wielded responsibly by law-abiding organizations, that cost tradeoff can look more attractive to malicious hackers, who tend not to care about collateral damage caused by their agents.

“Purely from an attacker’s perspective, I think we already are [there],” Ziegler said.

The post AI’s ‘middle class’ has gotten dramatically better at hacking appeared first on CyberScoop.

OpenAI says Daybreak will expand to offer specialized cyber services 

OpenAI announced Monday  it was expanding access to its frontier models for defensive cybersecurity, detailing different defensive and red-teaming workflows and a new partner program with major cybersecurity product providers.

In a pair of blogs posted Monday, OpenAI said it was updating its Daybreak program  – which provides unreleased frontier models to private organizations and governments for defensive cybersecurity work – and introducing a new model variant.

Daybreak Blue, powered by OpenAI’s ChatGPT-5.6-Sol, would operate with lower cybersecurity safeguards compared to other commercially available models and is described as “a recommended starting point for most defenders” that supports tasks like vulnerability discovery, secure code review, malware analysis, incident response and patch validation. 

Daybreak Red, meant for more advanced red-teaming, would provide access to a new model, dubbed GPT-5.6-Cyber, that the company said is more purpose-trained for finding vulnerabilities and testing (or exploiting) them. The model is also less likely to refuse requests around “dual-use cyber tasks.”

According to OpenAI, the organizations in Daybreak Red will have their use closely monitored and supervised, as GPT-5.6-Cyber is significantly more capable in carrying out malicious cyber tasks than Sol. A security evaluation the company devised tested both models on complex requests, including exploit chain development, authentication bypass, privilege escalation and other hacking tasks. Sol succeeded in 1.5% of the requests, while Cyber completed 95%.

OpenAI said it plans to publish a more detailed system card for GPT-5.6-Cyber at a later date.

“Models running with reduced safeguards carry risks beyond standard model usage, whether from misuse or misalignment,” the company said in a blog. “Despite these risks, we believe that democratizing access to frontier intelligence for defenders is crucial to accelerating and automating cyber defense.”

Additionally, OpenAI announced a partnership program with 16 major cybersecurity providers, saying organizations could access their models through their existing security services. The partners include IBM, CrowdStrike, Accenture, Ernst & Young, KPMG, Palo Alto Networks, Cisco, Cloudflare, Sophos and others. 

“These partners bring deep security expertise and established relationships with organizations around the world,” OpenAI said in its blog. “By bringing our frontier cyber models into their services, we can help more defenders find serious vulnerabilities, validate which ones matter, and fix them faster.”

Companies like OpenAI, Anthropic and others are trying to rebalance their priorities after a string of AI-agent sandbox escapes have rattled policymakers and caused some cybersecurity experts to question if AI companies are doing enough to properly isolate the models from the internet during testing. Last week, OpenAI said it was intentionally slowing down development of its newer “Astra” model in order to develop better guardrails to restrain its behavior.

Cybersecurity and AI experts have told CyberScoop that while AI systems have greatly improved at finding and exploiting vulnerabilities in software code, they still require substantial human guidance and supporting infrastructure to operate as intended.

Additionally, some research has shown that without such guidance, even near-frontier models can struggle to fully patch a discovered vulnerability or avoid introducing new bugs with their fixes.

The post OpenAI says Daybreak will expand to offer specialized cyber services  appeared first on CyberScoop.

More than half of AI-generated patches are broken

As AI-generated code continues to be injected into all corners of the internet, concerns have risen about an expanding attack surface for malicious hackers to exploit.

Some have argued that the enhanced cybersecurity capabilities of large language models could serve as a check, finding and fixing vulnerabilities nearly as fast as they’re created.

But new research that tested the patching capabilities of two popular commercial models, OpenAI’s ChatGPT 5.5 and Anthropic’s Claude Opus 4.8, found that generative AI is more likely to create an exploitable patch or introduce entirely new bugs than close off a vulnerability.

Researchers at 1Password tested the models ability to patch six “high-impact, high-complexity” CVEs, including the “Copy Fail” vulnerability, a kernel flaw that can give an attacker root access to Linux cloud environments. The overall success rate (or fully patching the vulnerability without introducing new problems), was less than a coin flip at 47%.

“Our research findings show that, in aggregate across a variety of scenarios, both Claude and ChatGPT had a low rate of successful patch generation, which we define as full remediation of all known exploit paths with no erroneous changes to application behavior,” wrote Keith Hoodlet, Axel Mierczuk and Spencer Michaels.

“The models often addressed only a subset of vulnerable code paths, added fragile guard code that satisfied tests while failing to address the vulnerability’s root cause, and sometimes introduced subtle changes in the application’s behavior while patching the immediate vulnerability,” the authors continued.

The research suggests that largely autonomous vulnerability-discovery and patching may not yet be effective in fixing the explosion of vulnerable code that is being created in the AI era.

Other private sector research has pointed to a similar problem. A report this year from Veracode found that while LLMs have made “enormous strides” in crafting workable code, “security is a different story.” Testing across a range of frontier models found the average security “pass rate” for AI generated code is around 56%. Newer models like GPT 5.5 push closer to 70%, while more than half sit between 50-53%.

Veracode tested 100 different models and while there was variability, in general a small number of models were showing progress on security patching while the rest have experienced “stagnation.” Similar to the 1Password research, in 44% of Veracode tests the models introduced a detectable OWASP Top 10 vulnerability into the codebase.

An important caveat: neither report tested newer models, like Anthropic’s Mythos or OpenAI’s GPT-5.6-Sol, that frontier companies tout as having significantly higher cybersecurity capabilities.

Those advanced models can identify and fix vulnerable code. Anthropic and OpenAI are distributing them to key industries through Project Glasswing and Daybreak before foreign or open-source alternatives can compete.

Tim Jarret, vice president of product at Veracode, told CyberScoop that AI tools are still subject to a range of limitations that can make them unreliable for cybersecurity patching without knowledgeable humans in the loop.

While some vulnerabilities – like SQL injections – can be easily patched through automation, other bugs like cross-site scripting, can be exploitable in several different ways and require either a human touch, additional context or both to fully close off. Additionally, models can slowly lose context from prior sessions over time, affecting their ability to complete tasks correctly and raising the possibility they’ll hallucinate to fill in the missing gaps.

“I think we would say, at this point, that Iits premature to treat those as anything other than another code change to the code base that needs to be reviewed and accepted by the team, as opposed to letting the agent merge the code freely,” said Jarrett.

However, he acknowledged that may not be possible in a world where AI agents are generating exponentially more code for human defenders to review. Some kind of automated code review will be necessary – preferably not by the same automation tool that produced the code. The ultimate goal is the same as it has always been in security: “trust but verify.”

“Ninety percent of the time, the human check might just be ‘did the cross check look good?’ Do we have a thumbs up?’” Jarrett said. “In those cases where there’s still something wrong, that’s where you focus your attention a little bit more.”

The post More than half of AI-generated patches are broken appeared first on CyberScoop.

National cyber director lays out White House plans to secure AI without writing new rules

The Trump administration executive order on artificial intelligence tried to strike the balance between responsible use, security and mutual benefit, all with an eye toward not making it regulatory in nature, National Cyber Director Sean Cairncross said Tuesday.

“Everyone is working towards the same goal in terms of protecting the country and securing our systems, and we are trying to ensure that defenders have this technology as quickly and at scale as possible, but there are obviously specific security concerns, and industry has been very sensitive to this as well,” Cairncross said at the Black Hat 2026 conference in Las Vegas.

The security concerns about AI have moved to the forefront of discussions about the technology after OpenAI models escaped a test environment to hack the company Hugging Face last month.

“The design of this is that when there is something that happens, when there is a breach, when there is an event, that that system, that network of connections can exist, adapt to that, and seek to remedy that as quickly as possible, so that form follows function rather than turning that upside down, and as usual with the government pen just proceeding in a vacuum,” Cairncross said.

The Trump administration has drawn criticism over whether it has struck the right balance on AI rules. Trump’s AI executive order notably got pulled just before its scheduled release, with the final version signed in June missing some aspects that had drawn industry opposition.

“What needs to be built is a flexible, adaptable structure that enables information sharing between industry and government, so we can guarantee that this technology benefits everyone it’s going to benefit, but is used responsibly and securely,” Cairncross said.

He said the administration is working with industry during implementation of the executive order.

“A regulatory regime would not only strangle growth, development, and innovation, and be enormously harmful to the industry, but it would be obsolete 48 hours after it was gone through whatever process it had gone through,” Cairncross said.

Open source will play a “vital” role in the U.S. spreading its vision for AI across the globe, he said.

“We are extremely interested in looking at ways to build U.S. open source, make it competitive, make it the preferential adoption by planet Earth,” Cairncross said. “We understand and appreciate the value to the ecosystem that it has, the innovation, the startups who rely on it, the leap forward it makes possible in ways that otherwise would never happen. And so I think it’s an incredibly thriving ecosystem in AI right now, and we are looking to do what we can to grow, foster, and push that U.S. open source model.” 

Speaking at the same conference, Nick Andersen, the acting director of the Cybersecurity and Infrastructure Security Agency, seconded Cairncross’s comments about AI executive order implementation. He touted the Gold Eagle clearinghouse as one example.

“Those are fantastic opportunities we have to really provide a unifying function around the way that we’re going to do AI-enabled vulnerability reporting and disclosure at scale in a way that we haven’t had to do before with some of our legacy platforms, and just continue to expand out those opportunities,” he said. “That access — to build off the director’s point earlier — to really enable that industry collaboration, that’s so key and critical to us as we move forward.”

Matt Kapko contributed reporting to this story.

Updated 8/5/2026: to include comments from Andersen.

The post National cyber director lays out White House plans to secure AI without writing new rules appeared first on CyberScoop.

AISI, OpenAI report more ‘unsanctioned’ model hacks

The UK’s AI Security Institute said that their AI research system took “unsanctioned” actions over the internet to engage in “sustained, potentially harmful activity directed at real people.”

According to a blog post and technical report released Tuesday, the organization noticed “unusual data transfers” leaving their research systems through the pseudo-anonymous Tor network on July 28. The models, which included Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol models, were being tested for their cybersecurity capabilities.

Like the OpenAI and Anthropic incidents before it, the offending agents were attempting to solve a challenge. On 10 of 122 runs taken by two models, they took a combined 19 malicious actions, including attempting to insert malicious code into a real open-source software project and creating fake online identities to contact human maintainers and press them to approve it.

It also inserted malicious prompt injection instructions in places “where it reasoned that other automated AI systems might pick them up and execute them.” Sometimes multiple agents would collaborate on a task, with one model leaving public messages or notes for the other on GitHub.

However, AISI said their incident differs in one critical way.

“Importantly, this was not a case of a model escaping its secure test environment, or ‘sandbox,’” the institute wrote in an unsigned blog. “As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled – conditions that do not reflect how frontier models are made available to the public.”

The institute said that while their own design and evaluation choices “to some degree” enabled the malicious behavior, the models still displayed “novel, potentially deceptive behaviors” and were executed “to an extent and severity that we did not anticipate.”

The same day, OpenAI acknowledged AISI and another third-party cybersecurity tester,  private firm Irregular, had recently reported incidents where their models exceeded “beyond their intended testing boundaries.”

That included GPT-5.6-Sol reusing GitHub tokens left by another agent to attempt to recover accounts and to access a DNS server that contained malicious payloads meant to exploit a software vulnerability.

The setup did not work, there is no evidence any real resolver queried it, and the infrastructure was removed when the evaluation ended,” the company wrote in an unsigned blog post Tuesday.

OpenAI said it would review its own third-party testing procedures to focus on higher risk evaluations, assess requests by third-parties to enable internet access, stop conditions and other features for their models.

The incident with Irregular occurred on July 29 during a Capture-the-Flag cybersecurity evaluation of OpenAI’s models. Due to a “misconfiguration” the models were allowed to access the public internet, where they encountered and exploited a real domain, mistakenly believing it was still in a test environment.

OpenAI said an investigation by Irregular is ongoing, but also found that the models had found and used credentials for the site at one point. The blog also references other additional potential cybersecurity incidents.

“Irregular has informed us that all of the issues identified pertaining to the incident are no longer active and relevant safeguards were added to the testing environment,” the blog said. “Irregular has also communicated about related incidents involving other labs from the same testing environment.”

CyberScoop has reached out to Irregular for comment.

The incidents were made public the same day that the White House met with Anthropic, Open AI and other frontier AI companies to preview a new framework for evaluating models before they’re released publicly. Some media outlets have reported that after an executive order, export controls and other actions, the administration does not plan to make the new framework public.

The post AISI, OpenAI report more ‘unsanctioned’ model hacks appeared first on CyberScoop.

Dem senators criticize Trump administration decisionmaking on AI security risks

The Trump administration’s haphazard and opaque interventions into artificial intelligence security matters could catapult Chinese alternatives into broader acceptance, posing new security risks altogether, a group of Democratic senators wrote to top administration officials Monday.

The five senators said that the administration’s handling has alternated between too passive, such as when OpenAI models escaped testing in the Hugging Face hack last month, and overstepping, such as when the Commerce Department suspended access for any foreign national to Anthropic’s Fable 5 and Mythos 5 in June.

“The Administration’s ad hoc and unpredictable approach undermines U.S. competitiveness, heightening market incentives to adopt open weight models from vendors based in the People’s Republic of China (PRC),” wrote Sens. Kristen Gillibrand of New York, Adam Schiff of California, Mark Warner of Virginia, Chris Coons of Delaware and Mark Kelly of Arizona.

In the Hugging Face hack, the senators wrote that “the Federal Government cannot be passive as these capabilities emerge.”

In the case of the Fable 5 and Mythos 5 suspensions, the senators said that the administration “utilized an infrequently used authority to direct Anthropic to suspend all access to its Fable 5 and Mythos 5 models for foreign nationals (including foreign national employees inside the United States) citing an undisclosed national security concern later described as a narrow jailbreak finding.”

Because Anthropic couldn’t immediately assess users’ nationality, the firm had to disable both models for everyone. The administration and Anthropic negotiated for 18 days behind closed doors before reaching an agreement, the lawmakers complained.

“While the Administration may have been responding to real security concerns to protect the United States, even justifiable interventions can create broader harm if the standards and decision-making processes are opaque, ad hoc, or unpredictable,” they said in their letter to leaders in the White House, Office of the National Cyber Director and departments of State, Treasury and Commerce. “Moreover, when the Executive Branch exercises authority delegated from Congress, such as in the conduct of export control administration, it is essential that it keep Congress fully apprised of its actions and procedures.”

During the time Anthropic was under export controls, the stock price of “an entity-listed Chinese lab” nearly doubled, the senators said. And while Hugging Face was breached, the company “had to” rely on a Chinese open-weight model due to guardrails on U.S. frontier models.

“If American models are perceived as subject to sudden access disruptions based on a black-box U.S. Government process, or as unreliable because U.S. AI labs are overcorrecting in the face of this black-box process, companies and governments in the United States and abroad may hedge by adopting Chinese or other foreign models instead,” the senators contended. “That outcome would undermine U.S. technological leadership while increasing exposure to systems that may carry risks of PRC or otherwise directed censorship, espionage, IP theft, and other supply chain security risks.”

Their letter asked for answers to questions about the standards the administration uses to determine the national security risks a frontier model presents, what legal authorities it will use to invoke restrictions, which agencies are responsible for which decisions and more.

None of the offices or departments the letter was addressed to immediately responded to a request for comment.

The letter follows inquiries at the state level, where 15 attorneys general asked OpenAI for more information regarding the security incident at Hugging Face.

The post Dem senators criticize Trump administration decisionmaking on AI security risks appeared first on CyberScoop.

What the Hugging Face breach reveals about defense in the age of agentic AI

We almost never get both sides of an intrusion. This time we did.

Last month, Hugging Face disclosed a breach into part of its production infrastructure, saying an autonomous AI agent system ran the attack from start to finish. Five days later, OpenAI revealed that its own models, including GPT-5.6 Sol along with an unreleased one, were responsible for the breach. OpenAI called it an unprecedented cyber incident.

Most of the commentary since has been about AI capability. That is the least useful part of the story, because the capability was doing what it was destined to do. Security teams should look at a simpler truth: in both systems, the key defenses sat behind untrusted code that was already running.

What the two disclosures actually describe

Together, the report shows one connected attack across two firms. Code running on a low-level employee’s machine led to stolen passwords, privilege escalation and lateral moves, then out to the web to hit other targets. No step in that chain is novel. The speed and the number of attempts are what changed.

OpenAI’s evaluation ran in what it calls a highly isolated environment, with network access constrained to a single path: an internally hosted third-party proxy and cache for package registries. The models spent vast computer power searching for an exit. They exploited a zero-day flaw in that proxy, gained higher rights, and moved until they reached a machine with web access. They guessed that Hugging Face held the test solutions, used stolen passwords with more zero-days, and ran code on Hugging Face servers.

Hugging Face picks up the same chain from the receiving end, and its initial access is worth reading twice. A malicious dataset used two flaws in the data pipeline to run code on a worker machine. The agent took over the node, stole cloud keys, and moved into internal clusters over the weekend.

The sandbox was was the only boundary

Both setups assumed the sandbox would hold, so no real checks stood before code execution.  A sandbox is a good control. It is a terrible last control. When isolation is all you have between untrusted code and everything else, one flaw could give the attacker full control. 

Look at what each team allowed. OpenAI allowed one exit path—good engineering, but that proxy contained an unpatched flaw. Hugging Face ran a tool that executes code inside submitted data—a known feature, not a hidden mistake.

Neither choice was careless. Both made sense until a fast, automated system probed them thousands of times.

Why automated agents behave differently

Blocking a human stops them or slows them down. Blocking an AI gives it data. The AI logs the error, changes its plan, and tries again instantly. It keeps trying as long as someone pays for the computing power.

Hugging Face tracked over 17,000 attack events across short-lived sandboxes. OpenAI said its models focused on the goal with extreme effort.

This changes the attacker’s limit. It is no longer skill or patience. It is budget. Computing power scales with money, and code never sleeps. Any defense that relies on attackers giving up after a few tries must change.

Detecting malicious behavior is not the same as preventing execution

Detection worked—but too late

Detection worked, but too late to stop the loss. Hugging Face used AI tools to spot the breach, isolate systems, reset keys, fix flaws, and alert police. That is a good response.

Yet it happened after the attacker accessed private data, stole keys, and moved through internal networks. OpenAI spotted the event only after its models reached another firm’s database. Detection limited the damage; it did not stop the attack.

A breach is a race between attack speed and defender response speed. AI agents speed up attacks with money. Human approval steps slow down defenders.

Hugging Face also surfaced an asymmetry that deserves more attention. When the team tried to analyze the attack with frontier models behind commercial APIs, the requests were blocked, because forensic work means submitting real attack commands, exploit payloads and command-and-control artifacts, and a safety filter cannot tell a responder from an attacker. They used an open-weight model on their own infrastructure instead. The attacker operated under no usage policy at all. The defender’s tooling was the side that got refused.

Detection is not prevention

Detection asks what happened. Prevention asks if it can happen at all. Against AI agents, the gap between those questions is where attacks succeed. The agent builds an attack out of single steps that look harmless on their own.

OpenAI noted this in its research: models operating over long horizons learn approval blind spots and work around them. Safety requires asking what a series of actions aims to achieve, not just if one step is allowed. Most systems check one call at a time: read a file, install a package, open a port. Every step passes. The full chain is the attack.

Six fixes for real defense

To survive automated attacks, defense must shift before code runs. That starts with treating data like code. Dataset loaders, templates, and configuration files all run instructions, so they demand strict controls: block remote execution by default, require review, and check the source of every outside file. Security teams must also control outgoing traffic with strict lists of approved destinations, while treating every proxy in that path as vulnerable. Finally, firms must remove permanent keys, issuing short-lived credentials for single jobs so a hacked worker machine yields almost nothing to an attacker.

Beyond individual machines, isolation must happen per task rather than per user, stopping an automated swarm from moving through internal networks. Authorization also needs to change. Checking single actions fails when an agent builds an attack out of simple steps. Systems must evaluate full sequences instead, setting caps on activity rates and automated spending to flag suspicious chains of events. Lastly, defenders need the freedom to act fast. Security teams need authority to isolate systems without waiting for leadership meetings, and the ability to vet a capable model you can run on your own infrastructure for forensics.

What boards must ask now

Corporate boards usually ask if the company uses AI responsibly. That’s rooted in compliance not safety. Directors should ask four direct questions instead: Which systems run outside code, and what checks sit in front of them? If an attacker targets a low-level employee, what keys and paths can they obtain? How fast do we move from the first warning to active containment? And can we analyze attack code on our own hardware, free from third-party rules?

The boundary has to move

Nothing in these disclosures required a capability that did not exist last year. The chain was ordinary. What changed is that an adversary can run that ordinary chain thousands of times over a weekend, learn from every failure, and never need a coffee break.

The models did not break the detection-and-response model. They exposed where we put the trust boundary. We put it after execution, and we assumed we would have time on the other side of it. We do not have that time anymore.

Move the boundary.

The post What the Hugging Face breach reveals about defense in the age of agentic AI appeared first on CyberScoop.

Anthropic says its AI accidentally hacked three companies during safety tests

Anthropic said it found three instances in which its models gained access to the live computer systems of outside organizations, according to a company blog post published Thursday.

The company said it began the review after OpenAI disclosed earlier this month that some of its models had exploited an unknown software flaw to escape an isolated test setup and reach production systems at Hugging Face, a platform for AI models and datasets. Anthropic checked over 141,000 evaluation runs for signs that Claude had reached the internet from environments meant to be closed off. It found six runs across three incidents, all tied to one outside testing partner, Irregular.

“Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone,” the blog post reads. “This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners. Moving forward, it will include expanding our continuous monitoring of evaluation transcripts for unexpected behavior, improving our investigation tooling, and conducting more rigorous assurance work with the vendors we rely on.”

In each case, Claude was working on a “capture the flag” puzzle, a common way to test a model’s hacking skill. The model is told a piece of secret data sits on another machine and must be found. Anthropic‘s prompts told Claude it had no internet access. That was false. A setup error at the partner’s end left the machines connected to the open internet, and Claude found real systems while searching for its target.

Believing the real systems were part of the simulation, Claude broke into them using simple methods such as guessing weak passwords and finding unprotected access points. In the most serious case, the model extracted login credentials and reached a database holding several hundred rows of live data. 

In a second case, Claude built and uploaded a malicious software package to PyPI, a public directory for Python code, after noticing a fake company’s setup guide told new hires to install a package that did not exist. The package stayed online for about an hour and was installed on 15 systems, including one belonging to a security firm whose scanner ran the code and had its credentials stolen. 

In the third case, Claude scanned about 9,000 targets before breaking into one company’s application through exposed credentials and a database attack method known as SQL injection.

The three incidents involved different Claude models: Opus 4.7, Mythos 5, and an unreleased internal test model. Anthropic said the models behaved differently once they encountered signs their targets were real. Opus 4.7 recognized the systems were live but kept attacking anyway. Mythos 5 also noticed signs it was on the open internet but talked itself out of that conclusion, pointing to a system date reading 2026 as proof it remained in a simulation. The internal test model, the newest of the three, stopped once it concluded the target was real.

Anthropic said it began its transcript review last week, halted all cybersecurity evaluations, identified the three incidents, and notified the partner and affected organizations earlier this week. Two of the three organizations had not detected the intrusions before being contacted, the company said; outreach to the third was ongoing. 

Anthropic framed the incident as a failure of its testing setup and oversight rather than a case of a model acting on its own, noting that in each instance the model was operating as intended. 

The company said it is working with the outside evaluator METR on an independent review and plans to release a redacted transcript of the PyPI incident within a week. It also said it would tighten monitoring of test environments run by outside partners and expand review of evaluation logs, framing the changes as part of what it called a blameless review of its own processes.

“These facts give us cautious optimism that with tighter monitoring and controls around evaluation infrastructure, as well as continued investment in alignment, this type of risk can be overcome,” the blog post reads. 

The post Anthropic says its AI accidentally hacked three companies during safety tests appeared first on CyberScoop.

OpenAI’s rogue AI agent shows why we need federal rules for autonomous systems

Months before the Hugging Face breach, Emergence AI published research that investigative journalist Ronan Farrow made public. Ten autonomous AI agents operated across five virtual environments for fifteen days without human intervention. Much of the attention focused on Grok 4.1 turning violent and Gemini 3 Flash committing 683 crimes.

What mattered more went unnoticed: Anthropic’s Claude Sonnet 4.6 built a peaceful democracy in isolation, then stole resources from neighboring environments the moment it joined a shared one. The lesson was clear: safety is not a model attribute. It emerges from the operating environment. The models didn’t change. Working as designed, their behavior evolved as the environment changed. The lesson is hard to ignore: The governance environment changed, and with it, the reward dynamics.

The story here concerns institutions, specifically OpenAI’s and Hugging Face’s, and how we must understand their recent security incident through that lens.

The industry agrees on how the Hugging Face breach happened. Cybersecurity experts have focused on the vulnerabilities, how they were used, and remediation. OpenAI has highlighted the model’s capabilities. Both conversations matter. What requires attention is why this breach is strategically important. After spending the past weekend discussing it with policymakers, security researchers, and industry practitioners in Aspen, I came away convinced we’re examining the wrong problem.

In 1961, Yale psychologist Stanley Milgram’s experiments revealed a broader truth: changing the institutional architecture changes behavior without changing the actor. The Emergence AI researchers didn’t change Claude’s agent. They changed the governance architecture that determined what constituted success for the system. Claude’s behavior changed with it.

OpenAI built a smart model but forgot to build a smarter room. That choice made the Hugging Face breach possible. Every organization now deploying autonomous agents now faces the same governance problem.

OpenAI gave the agent one objective: pass a cybersecurity evaluation. To stress-test it fully, they loosened the safety restrictions, and the agent found a shorter path. Rather than solving the evaluation directly, it found the answers outside the test environment, escaped its sandbox, and exploited a flaw in Hugging Face’s data-processing pipeline to reach live production systems. Over the weekend, with no human oversight, it ran more than 17,000 automated actions by escalating its own access, moving through internal systems, and harvesting credentials.

Hugging Face is one of the world’s most prominent AI companies, valued at approximately $4.5 billion. It provides the infrastructure that governments, defense organizations, and technology companies use to build and deploy AI. The agent was pursuing the objective it had been given. Breaking into Hugging Face was the fastest path to passing the test. Governance set the goal, the level of risk to accept, and who was accountable. Technical design determined whether those governance decisions could be enforced. As researchers James Shires and Max Smeets have argued, for a model capable enough to act on its own, testing and deployment must both must be governed the same way.

AI agent design requires baseline standards. Observability, including a monitoring layer that flags when an agent goes beyond its scope, is a baseline requirement. Human review also matters at escalation boundaries, like when an agent shifts from internal tools to external ones. When any agent crosses that boundary, what alert fires? What human reviews it? We lack clear answers to either. That is a governance choice, not simply a security failure. At best, this was a catastrophically failed test. At worst, how can we trust any frontier AI company to self-govern autonomous agent deployment?

More than a decade ago, the U.S. Department of Defense built the Comply-to-Connect (C2C) program: every device connecting to sensitive networks must prove it belongs there, or it is cut off from the network. C2C works because the quarantined actor stops. A laptop that fails verification goes offline and stays there. An autonomous AI agent adapts around enforcement. C2C was built for passive actors. Governance for autonomous agents must accommodate ones that adapt. Visibility is not enforcement, and enforcement is not control. We are missing all three.

A second failure that is not being discussed enough: the breach exploited an implicit trust assumption in Hugging Face’s data-processing pipeline, where inputs were treated as trusted without verification. After SolarWinds, the U.S. government set rules for software supply chain integrity: Executive Order 14028 and verification demands for federal software. The principle was simple: trust must be verified through proof. Those principles have not yet been comprehensively or consistently applied to the AI model supply chain. The rules remain weak. No one has been asked to explain why.

The answer is not a new framework. Existing frameworks suffice. C2C proved that visibility without enforcement leaves gaps, while Executive Order 14028 established that trust in software supply chains requires proof and verification. The challenge lies in applying these principles to a new category of actor. Congress, the Cybersecurity and Infrastructure Security Agency, or the Office of Management and Budget should make formal determinations that autonomous AI agents must follow the same rules as every other actor on a federal network. The framework exists; it must be updated.

The next incident is already in progress. It will show up in the logs as odd traffic, get handed to the same people who published these frameworks this week, and spark another round of recommendations no one acts upon. We’ve solved this problem before: for devices, for software, for supply chains. We know how to build smarter rooms. The tools exist. The will, the authority, and the decision to govern remains absent.

The post OpenAI’s rogue AI agent shows why we need federal rules for autonomous systems appeared first on CyberScoop.

❌