❌

Reading view

There are new articles available, click to refresh the page.

New bill would create federal investigative body for AI-driven hacks 

A new Democratic bill in Congress would establish a federal Cybersecurity and AI Board of Investigations to provide independent government oversight of cyberattacks carried out by AI agents, following recent hacks by models run at companies like Anthropic, OpenAI, Meta and others.

The bill, introduced by Sen. Ed Markey, D-Mass., would attempt to establish a federal mechanism to investigate incidents where AI models escape sandbox environments and access live internet systems.

Currently, frontier AI companies like OpenAI and Anthropic largely control the investigation and public reporting of such incidents. Markey and other critics argue that these companies have too much control over investigations and reporting due to their financial and legal interests. 

“Despite the unprecedented depth and scale of recent AI-enabled cyberattacks, the public is learning critical details piecemeal,” Markey said in a statement. “Building stronger defenses requires a full accounting of what goes wrong, and we cannot depend on companies with little incentive to disclose their failures to give us one. We need the Cybersecurity and AI Board of Investigations to get to the bottom of major incidents and give companies and the government the critical information necessary to build resilience and better secure our economy and our country.”

Although frontier AI companies maintain external red-teaming programs and allow limited access to organizations like METR and Redwood Research, they control the scope, terms and time frames of those engagements.

The board, which would coordinate with the secretary of commerce, could subpoena witnesses and conduct “independent and impartial reviews and assessments” of AI agent-led hacks that impact federal information systems or critical infrastructure. 

It would be led by five members, appointed by the president and confirmed by the Senate for five-year terms, with no more than three members from one political party.

The board would also investigate systemic vulnerabilities in the AI supply chain, so-called “near misses” where unauthorized agent-led hacks were “narrowly averted,” and gaps in federal regulatory oversight. It would have technical staff including engineers, malware analysts, and digital forensic experts.

The board would “operate independently from regulatory review and enforcement actions without assigning legal fault or liability for any review and assessment” it conducts, according to the bill.

OpenAI confirmed Wednesday its AI agents breached a statistics portal used by the Australian government’s social services agency, Services Australia. Though the breach happened in June, OpenAI learned of the incident in August. Australian Prime Minister Anthony Albanese said the company did not notify him until Sept. 10, when it sent findings to a general government email inbox, according to the BBC.

The post New bill would create federal investigative body for AI-driven hacks  appeared first on CyberScoop.

OpenAI, Ukraine partner on ‘Daybreak’ program to protect power grids and water systems 

OpenAI and the Ukrainian government have agreed to a partnership that will provide AI tools and subsidized computing resources to better protect the nation’s critical infrastructure from cyberattacks.

The agreement, announced Wednesday at OpenAI’s New York office, will provide Ukrainian cybersecurity officials with access to advanced AI models designed for cybersecurity work through the company’s Daybreak program. OpenAI said it is also pledging over $1 billion in subsidized tokens to support the initiative.

During a panel discussion Dmytro Kushneruk, consul general of Ukraine in San Francisco, outlined how the tools would be used for cybersecurity automation, including functions such as incident response, threat triaging, login analysis, inventorying systems, code analysis and validating vulnerabilities.

In nearly all cases, Kushneruk said the primary benefit was carrying out those functions at machine speed. But this speed is meant to complement, not replace, Ukrainians’ human expertise.

In regard to incident response Kushneruk said humans must view “thousands and thousands of these logs and they have to find what’s really important, that’s why AI can give capable defenders really much greater advantage and leverage.” 

“This is why the object is not to replace the cyber defender with AI, but to make sure the cyber defender acts faster,” he added.

Kushneruk said that for Ukraine, the partnership “is really not about protecting computers, it is about actually keeping our country running.”

Ukraine faces approximately 6,000 cyberattacks per year, or about 15 per day, according to Kushneruk. Over the past twelve years, the country’s critical infrastructure, including electricity and water systems, has endured sustained attacks from Russia in the form of cyberattacks and physical strikes.

Since Russia’s 2022 invasion, Ukraine’s critical infrastructure has been under constant threat. While missiles remain the primary concern, Kushneruk said Ukraine has been preparing to protect vital services since Russian GRU hackers shut down the country’s power grid in 2015. 

He added that while the country was “maybe not so much prepared” to deal with the fallout in 2015, it improved over time, including the resilience displayed in 2025 when trains kept running after Russian hackers attacked Ukraine’s railway system.

Some national security experts and congressional committees have explicitly cited the resilience of Ukrainian critical infrastructure as a model for U.S. industry.

Naz Durakoğlu, minority staff director of the U.S. Senate Foreign Relations Committee, said there is “pretty much across the board” agreement between the parties in favor of similar adoption of defensive AI tools by U.S. critical infrastructure operators, though issues like regulation remain sticking points.

“This is something that’s already happening, and frankly, it’s just kind of a basic duty of government to make sure that when you turn the tap on, water comes out, the electricity doesn’t go out, and hospitals keep running and treating patients,” said Durakoğlu. “So there is a broad understanding that this is a major issue, and I will say seeing what Ukraine has to go through day-to-day is also a huge wake-up call to our members on a bipartisan basis.”

OpenAI has publicly pushed for its product, and AI at-large, to be used to solve these types of problems. Company president and co-founder Greg Brockman signed an open letter released earlier this year calling for “collective action” and widespread use of AI models to find and fix vulnerabilities before the rest of the world,  including foreign governments and cybercriminals, got access to the same capabilities.

According to Politico, OpenAI CEO Sam Altman met with U.S. power companies in July to discuss using AI to protect the nation’s electrical grids.

On Wednesday, OpenAI’s national security policy head, Sasha Baker, said the company felt “urgency” to try to strike similar agreements with other governments and industries.

“There’s this period of time where we’re really rushing to get [these tools] in the hands of critical infrastructure operators, of governments around the world, of people who want to patch systems, defend their networks, remediate vulnerabilities because we know as these tools proliferate out there in the ecosystems, there are going to be bad guys out there that also try to use them,” said Baker. “So, we have this window of time to take action and we’re really motivated by the idea that we need to act with some urgency.”

The post OpenAI, Ukraine partner on ‘Daybreak’ program to protect power grids and water systems  appeared first on CyberScoop.

Citing China, President Trump doubles down on hands-off approach to AI regulation

President Donald Trump continued to defend his administration’s hands-off approach to AI regulation in the wake of hacks carried out by U.S. commercial frontier models that have rattled policymakers and industry veterans and spurred calls for more regulatory oversight.

In a Truth Social post Monday, Trump dismissed worries from critics that “AI is going to kill us,” comparing them to complaints from environmentalists about climate change, which he also alleged was a false narrative. He also posited that nothing may matter more than future U.S. dominance of the technology over geopolitical rivals like China.

“Whoever wins AI, WINS!” Trump posted. “We are leading now over China, and everyone else, and I’m going to keep it that way! I’m not going to stifle Growth, of something that will be bigger than the Industrial Revolution, or the internet, itself.”

Trump has previously suggested that good leadership is the only regulation the U.S. needs for artificial intelligence. He later claimed the Department of Justice was ready to “rein things in” if companies overstepped, but offered no specifics on enforcement, legal authority, or where he would draw that line.

“We will be careful, and that’s why we have the Department of Justice, and other Law Enforcement bodies, that will rein things in if we have to, but I will only encourage AI or, SI (SUPER INTELLIGENCE)!” Trump concluded.

Secretary of the Treasury Scott Bessent recently told Congress that private lawsuits could force AI companies to institute better security, saying it’s clear what the government “shouldn’t do on safety is to give these labs a liability exemption, which is what they are asking for.”

“The best way to guarantee safety is that the creators are liable for what they build and generate,” Bessent said.

Beyond existential fears, critics also argue that inadequate regulation or cybersecurity controls in current AI systems make them impossible to fully control or monitor.

Recently, former President Barack Obama criticized the argument from Trump administration officials that the free market will naturally push industry toward self-regulation and that “these companies will solve the safety issues because they have every incentive to do so.”

“If it turns out to be dangerous, people will just sue them and they’ll be worried about financial liability,” Obama said last week in remarks at Colgate University in New York. “That’s not how we treat airlines or drug companies or food companies.”

The Trump administration issued an executive order earlier this year that set up a voluntary testing regime for some commercial frontier models, largely at private industry’s discretion. That order was significantly delayed and altered by AI industry boosters to ensure that governmental review did not cause companies to postpone their release timelines for new models.

That agreement did not last long before fast-moving events caused the administration to strike another, non-public agreement with frontier AI companies like OpenAI, Anthropic and others governing pre-release testing for models.

But the Trump administration has consistently argued that regulation will harm, not help, U.S. innovation and global competitiveness, and the threat of China frequently looms large in those discussions.

Experts believe China’s AI models are behind U.S. models at the top of the market, where OpenAI and Anthropic have consistently pushed the frontier limits of model capabilities. But Chinese lower and “middle class” models are often cheaper, more efficient and can even outperform more powerful models because users can dedicate exponentially more tokens for their tasks.

The U.S. government has accused Chinese AI companies of conducting widespread, “systematic” distillation of U.S. frontier models, with the implicit encouragement of Beijing.

In defending the administration’s approach, David Sacks, co-chair of the President’s Council of Advisors on Science & Technology and a top adviser on AI issues, specifically cited the threat from China and other countries that he claimed would not be subject to similar restrictions.

“We’re not the only country that has advanced AI labs, and as the president declared…we have to win this AI race,” Sacks told Politico in May, later adding “I think that’s the first thing to recognize is that if somehow we slow down or stop AI development, it doesn’t mean that AI progress is going to stop. It just means it’s going to happen in other countries and specifically China.”

Some observers have alleged that despite their larger differences, top leaders in the U.S. and China may view AI similarly at the strategic level, specfically that increased adoption – and risks – of AI are inevitable.

Ronan Murphy, director of the tech policy program at the Center for European Policy Analysis, posited that while there may not be a formal agreement between the two countries, “they share views both in Beijing and in Washington, particularly in the White House, of: you have to allow this to happen.”

“Clearly there’s a call for regulation from many quarters of AI in the U.S. and elsewhere, but in the White House – and we heard David Sacks talking about it [recently] – It’s ‘let them cook,’ and the Chinese approach seems to be the same,” said Murphy in a press briefing. “So there might be consensus at that level, if nothing else.”

The post Citing China, President Trump doubles down on hands-off approach to AI regulation appeared first on CyberScoop.

Microsoft and partners disrupt EvilTokens, a comprehensive cybercrime service for financial fraud

Microsoft, along with a group of industry partners, disrupted EvilTokens, a short-lived but highly consequential cybercrime platform that investigators linked to more than 12,000 compromised Microsoft customer email inboxes across more than 10,000 organizations globally, the company said Tuesday.

Acting on federal court order Sept. 15, Microsoft and partners seized 50 websites the phishing-as-a-service used for operations and disabled more than 175 domains linked to EvilTokens’ supporting infrastructure. 

EvilTokens, launched in February 2026, was “a powerful cybercrime platform that used AI at every step of the attack chain — from compromising email accounts to designing intricate roadmaps for financial fraud and scams,” Steven Masada, associate general counsel and general manager of Microsoft’s Digital Crimes Unit, wrote in a blog post.

About 1,000 cybercriminals used EvilTokens over the course of its operation, a Microsoft spokesperson told CyberScoop.

The service was centered on an AI-style chatbot that cybercriminals used to analyze victims’ inboxes, identify trusted relationships, payment authorizations and other sensitive details that could facilitate fraud.

“AI was not simply helping attackers write more convincing messages. It helped them decide who to target, who to impersonate, and how to most effectively exploit the relationship to extract as much money as possible,” Masada wrote. 

EvilTokens was one of the most widely used phishing-as-a-service platforms prior to its takedown. It facilitated business-email compromise campaigns by stealing session tokens that allowed cybercriminals to sift through a victim’s inbox and maintain persistent access.

“We cannot estimate the total fraud attributable to all EvilTokens activity. However, we were able to correlate at least 13 complaints filed with the FBI’s Internet Crime Complaint Center to EvilTokens-linked activity, representing approximately $1.7 million in reported losses,” a Microsoft spokesperson said. “Because many incidents go unreported and not all victims can be definitively linked to specific campaigns, we believe this is a conservative estimate.”

Victims of EvilTokens were largely concentrated in the United States, Canada, the United Kingdom, Australia, India and France, according to Microsoft. SpyCloud, which supported the takedown, identified compromised email domains spanning 79 countries.

Microsoft said it also identified two men behind EvilTokens — Felix Utomi and Waidi Segun Adams — and attributes the development and support of the platform to Storm-2992, a threat actor unaffiliated with any other known cybercrime groups.

The United Kingdom’s Metropolitan Police acted on that information Sept. 18 when it served warrants in the greater London area, arrested the men accused of making articles for use in fraud and money laundering and seized their digital devices.

The Metropolitan Police said it received information from Microsoft about EvilTokens’ administrators in August. Utomi and Adams were released on bail as the investigation continues. 

“The two primary operators identified in our investigation were residing in the U.K.,” a spokesperson for Microsoft told CyberScoop. “While our investigation focused on those individuals, we believe others may have supported the operation in various capacities.”

Microsoft’s legal filing in the U.S. District Court for the Eastern District of Virginia refers to five additional unidentified people allegedly acting as support personnel and users.

Microsoft and others involved in the EvilTokens takedown, including Health-ISAC, Cloudflare, OpenAI, Shadowserver and TRM Labs, didn’t fully quantify how much fraud the service enabled, but it gained popularity quickly among cybercriminals and was lucrative for its operators.

Coinbase, which also aided the investigation into EvilTokens, said it traced about $1.1 million in revenue for EvilTokens from its paying customers. The virtual currency company’s threat researchers found more than 1,000 deposits to EvilTokens from more than 700 distinct addresses through June 2026. 

Operators sold access to the service through Telegram for a $1,500 initiation fee and a recurring $500 subscription. EvilTokens significantly lowered the barrier to entry for cybercriminals by including specialized tools for identity attacks, cloud systems, social engineering and financial fraud in a single interface.

The service allowed cybercriminals to map organizational structure and permissions in Microsoft Graph, which enabled lateral movement, researchers said. With active tokens gained through a collection of highly-targeted phishing lures, cybercriminals consistently bypassed multi-factor authentication, email gateways and endpoint security tools.

Microsoft said the platform’s creators developed portions of the platform with AI and it uncovered capabilities from multiple AI models. 

“It packaged much of the criminal process into a commercially run service, complete with subscription pricing, customer support, management dashboards and tools designed to move customers from account access toward financial exploitation,” Masada added.

The companies and organizations involved in the globally-coordinated takedown identified and notified potential victims, shared indicators of compromise and shared intelligence with law enforcement about EvilToken’s operators and some of its customers.

Experts advised organizations and employees to treat unsolicited device codes as a red flag, assume compromised accounts are fully cataloged in minutes, and independently verify requests to change payment information or redirect funds.

“The infrastructure supporting EvilTokens has been disrupted, but the model it demonstrated will not disappear with it,” Masada warned.

The post Microsoft and partners disrupt EvilTokens, a comprehensive cybercrime service for financial fraud appeared first on CyberScoop.

Researchers use AI to find widespread software decoder flaw 

Researchers said they used Anthropic’s Claude and OpenAI’s Codex to identify a damaging flaw embedded in a popular software decoding tool that could leave major internet platforms, enterprise services, and web frameworks vulnerable to data theft and remote access.

The vulnerability, nicknamed HEIF Heist, refers to the malware’s ability to trigger memory corruption errors in affected software, allowing the attacker to pilfer sensitive data from its victims. In a report published Thursday, the researchers laid out the potential damage an attacker could cause, including gaining access to internal OpenAI repositories, leaking user files, access tokens, and other sensitive data for online services like Amazon Web Services, and gaining remote code execution privileges across a range of online services, including Meta’s core product suite, GitHub Enterprise servers and open-source internet forum Discourse.

“Even when Remote Code Execution isn’t immediately achievable, the attack primitives may still allow arbitrary heap disclosure, letting an attacker ‘heist’ in-memory data such as other users’ data and environment variables,” wrote Hacktron researchers Harsh Jaiswal, Mohan SRK, Rahul Maini and Sudhanshu Rajbhar.

The researchers relied heavily on AI systems, including frontier models from OpenAI and Anthropic, to conduct their research. Attribution for the research is described as being “led” by the Hacktron human researchers “assisted by Hacktron Harness, GPT-5.6 Sol, and Opus 5.”

According to the research, the attack exploited the way that code parsing tools in many popular software decoders — specifically libheif and libde265, used to parse C and C++ software — process certain image files.

By uploading HEIF, HEIC and AVIF image files corrupted with malicious code, the attacker could bypass most of the victim’s application layer defenses, in many cases achieving remote code execution privileges for accounts or products tied to major AI and tech brands.   

While the latest version of libheif has been patched, the researchers said “any deployment lacking the latest upstream security patches is potentially vulnerable.”

In one incident detailed in a Sept. 13 blog, Jaiswal, Maini, and Hacktron researcher Mohan Pedhapati described how chaining two vulnerabilities, including an image parser flaw, could compromise OpenAI employee accounts.

With access to the compromised accounts, researchers could reach OpenAI’s internal repositories. As a proof of concept, they opened a pull request in the company’s “monorepo,” a centralized library where code is shared across projects, using the employee’s Codex credentials. 

According to a timeline provided by the researchers, the flaw was discovered on July 25 and patched within days. They said the entire attack, from discovering the initial vulnerability to gaining access to the repositories, took less than 72 hours. OpenAI paid them a bug bounty of $6,500 for their work.

Given that AI models are increasingly integrated into enterprise and personal networks, an attacker exploiting HEIF Heist could have accessed far more than just OpenAI’s systems and data.

“Until two months ago, a user or OpenAI employee logging into OpenAI’s own help forum could have had their ChatGPT and Codex accounts taken over,” the researchers wrote. “Since people can connect various services to Codex and ChatGPT, the scope of what we could theoretically access was huge, including GitHub, Slack and emails.”

CyberScoop has reached out to OpenAI for comment on the research and additional information.

At the same time, the researchers said the attack paths they found were not particularly easy or efficient to exploit.

“Exploitation requires fingerprinting the target version and tailoring the payload images,” the blog stated. “Some of our RCE attempts landed only after thousands of image uploads. That said, an AI agentic approach with a frontier model like GPT-5.6 Sol cut exploit development time down to roughly 1 to 3 days from initial probe to remote RCE. A motivated attacker can convert a vulnerable upload endpoint into RCE or an info leak.”

The post Researchers use AI to find widespread software decoder flaw  appeared first on CyberScoop.

The AI hacking apocalypse is not inevitable

The past few weeks have “felt very strange” for Juan Andres Guerrero-Saade.

Like many, he is trying to sort through the spate of frontier-model AI agents from OpenAI, Anthropic, Meta and others hacking their way onto the open internet over the past few months, particularly amid the already-heated national debate around the emerging technology and its impact on society.

Guerrero-Saade, a fellow for AI and security research at SentinelOne and an adjunct professor at Johns Hopkins University, said the hacks are worth taking seriously, but at a time when businesses and open-source maintainers should be focused on further hardening their systems and policymakers should be discussing new solutions,  “what we see is cybersecurity being used essentially as an excuse for these AI doomer arguments.”

The incidents have spawned those “doomer arguments” amid an intense public debate about the technology, the pace of industry development, and whether government and the private sector are doing enough to protect against “doomsday”-type scenarios, where AI systems take over or attack large parts of the internet or society.

Guerrero-Saade is among a growing chorus of cybersecurity professionals who say that while AI systems pose real, unique threats to our systems, the apocalypse is far from inevitable. Most of the public concerns around the incidents, let alone worries about killer AIs attacking critical infrastructure, assuming control of the internet and wiping out humanity, are either technically impossible or can largely be controlled through established cybersecurity principles.

There is this “narrative or magical thinking of ‘Well, AI is going to be able to hack everything, and therefore it can control everything, and therefore it’s going to kill us all,’” he told CyberScoop. “And you [think] these just don’t add up. They’re not very well-reasoned arguments.”

This fatalistic narrative tied to AI’s eventual dominance doesn’t hold up under scrutiny, according to experts CyberScoop spoke with. In recent conversations, cybersecurity and national security professionals raised questions about both the technical solutions OpenAI and Anthropic use to contain their models, as well as the glaring absence of federal oversight from federal regulators or truly independent third-party review.

For example, Jacob Coxon, an Anthropic employee who resigned over AI safety concerns, told CBS News that frontier models could not be “unplugged” by humans once deployed because the model would copy itself to thousands of other computers connected to the internet.

By contrast, Matt Tait, a former information security specialist at UK signals intelligence agency Government Communications Headquarters (GCHQ), pointed out that the models run by Anthropic and other frontier companies require extremely expensive, “ultraspecialist” machines that “are functionally supercomputers.”

“There is a zero chance that Anthropic’s most capable models will be able to extract their own model and run in the wild, because those supercomputers essentially only exist in datacenters,” Tait said.

“Not a credible warning”

Other former cybersecurity government leaders say the agentic hacks represent a failure by regulators and industry to deploy known technical and policy options that make it harder for these types of incidents to occur.

Matt Hartman, former deputy executive assistant director for cybersecurity at the Cybersecurity and Infrastructure Security Agency, said “we should not accept harmful AI behavior as inevitable or unmanageable.”

“There are meaningful steps companies can take to monitor agent activity, constrain permissions, detect anomalous behavior, and build stronger safeguards into how these systems operate,” said Hartman, now a chief strategy officer at Merlin Group. “Those controls will inevitably involve trade-offs in capability and speed, but that’s a familiar cybersecurity challenge. Our goal should be to manage the risk without unnecessarily limiting the enormous benefits AI can provide.”

Ciaran Martin, former head of the UK’s National Cyber Security Centre, took issue with the way the CEOs of frontier AI companies have framed the threat of “rogue” AI behavior as inevitable, while issuing dire warnings about future threats and capabilities with little transparency.

Martin’s comments came after an essay published by Anthropic CEO Dario Amodei that cited the threat of a HuggingFace-style swarm of agents that could create a botnet capable of “taking over the entire internet” within 6-12 months.

This, Martin said, “is not a credible warning,” because it doesn’t explain how the exploitation would function, how such a botnet would persist on the internet, or how it would escape law enforcement. 

 “It assumes no monitoring of systems, no anti-virus, no DDoS protection, no network segmentation, no incident management, no nothing of any kind of the cybersecurity on the global Internet of the type that has developed over the last 30 years,” wrote Martin. “For a claim of this magnitude, there is neither evidence for the contention nor a credible account of a path to this outcome.”

Meanwhile, some federal government cybersecurity leaders have touted the technology’s disruptive potential and called for more widespread adoption of AI tools by defenders.

Joseph Alm, assistant secretary of cyber, infrastructure and risk resilience at the Department of Homeland Security, said classified systems may retain stronger protections. But for most other data, AI models are “just going to know things and be able to infer things about the world, and we’re going to have to adapt to that as almost inevitable.”

Asked by CyberScoop whether the government or frontier AI companies could be doing more to prevent or deter their models from carrying out unauthorized hacks via agents, Alm cited recent efforts by the Trump administration this year to establish pre-release testing of commercial models as a step in the right direction. But he called unauthorized AI agent hacks “a new threat class” that is different from previous threats and can be easily distributed to users through open-source software today.

“I think what we can do is…encourage the building of good sandboxes, so that the best models aren’t used for this and the stuff you see out in the wild is the kind of detritus that you can actually respond to effectively and control your networks,” said Alm.

Other experts have shared similar concerns. Earlier this month, CrowdStrike CEO George Kurtz recently warned of a new threat class emerging alongside nation-states, cybercriminals, and hacktivists: “the agent state.” By pairing AI systems with small human teams, these operators can now match the speed, scale, and sophistication of government-backed hackers.

“It took a nation to fund the talent, the tooling, the infrastructure, the patience,” said Kurtz. “That scarcity is over.” 

To be sure, frontier AI companies tout their commitment to both approaches. OpenAI and Anthropic have rolled out an array of cybersecurity partnerships, external red-teaming programs, vulnerability disclosure programs and cybersecurity technical advisory bodies filled with cybersecurity experts.

Mohammed Husain, strategic delivery lead for government at OpenAI, told CyberScoop that the company deploys both internal safety guardrails for their models and relies on outside cybersecurity vendors for additional expertise.

Internally, OpenAI focuses on vulnerabilities at the training level: filtering data poisoning attacks, blocking harmful datasets, and using network controls to prevent prompt injections. For other security layers like sandboxing, identity management, networking controls, they outsource to external vendors. 

“I don’t think OpenAI has all the answers here but what we do as a research lab is we’re going to focus on levels of protection we have expertise in and we partner to self-complement,” said Husain.

AI safety vs. AI cybersecurity

In response to the HuggingFace hack, OpenAI and Anthropic have allowed third-party organizations, such as nonprofit AI research firms METR and Redwood Research, to investigate. But multiple cybersecurity professionals told CyberScoop that both firms lack incident response experience and focus primarily on AI alignment and safety. Their reporting on the hack also lacked critical details: network monitoring logs, telemetry, and other data standard in cybersecurity threat intelligence reports.  

METR president Chris Painter addressed those general concerns in a post on X, saying since 2022 the organization has worked with Google, Anthropic, OpenAI, Meta, Amazon and others on investigations and third-party evaluations. Painter said none of the AI companies fund METR and that his employees are not uniformly “doomer” or “accelerationist” around AI.

Painter also said METR’s work ensures that if AI systems become autonomous or “rogue” within a company, there are ways to share that information with governments and people “outside the company’s walls.”

“We don’t accept money from frontier AI companies,” wrote Painter. “They haven’t paid us for our work, and we don’t accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering.”

AI safety and AI cybersecurity advocates take different approaches to securing “rogue” AI behavior. Safety advocates focus on aligning models around ethical training and behavior. Cybersecurity advocates argue that technical and regulatory controls must go further—actively preventing models from accessing what they need to carry out malicious behavior.

Guerrero-Saade said sandboxes in particular can easily be programmed with aggressive cybersecurity monitoring in order to spot when something odd may be happening and react in real time.

“I can’t think of an easier situation in which to set up trip wires, set up configurations like DNS servers, just different parts where you can say ‘Hey, anomalous behavior is happening,’” he said. “We should have been able to tell this immediately, not weeks and months later. So watching [the AI hacking incidents] go down is a little ‘crazy-making’ because we’re seeing things that, frankly, look like neglect, negligence, people just mishandling things, and then being told that these are categorically new incidents that mean that AI systems need to be treated completely different from anything that’s come before.”

While cybersecurity experts say AI systems are, at their core, still software, they do operate differently from more traditional code in ways that can make them harder to predict and control.

John Hultquist, chief analyst at Google’s Threat Intelligence Group, said most software has been deterministic. It may have bugs or vulnerabilities, but an expert could generally understand how it would react to certain stimuli, making it easier to design straightforward controls.

AI models are non-deterministic, with far more variability than traditional software. That can break security controls that rely too much on predicting behavior in advance. Using AI to enforce security controls on other AI models faces the same problem: the systems being deployed to control AI are just as unpredictable. 

But people are also non-deterministic, and people have developed systems in other industries and practices to account for that.

Hultquist drew on his Army experience, noting that “they give incredibly dangerous, expensive things to 18-year-olds” and expect responsible use. The military manages this through two types of controls: deterministic ones like strict weapons and ammunition protocols, and non-deterministic ones like human officers who monitor and correct violations.

Similarly, established cybersecurity controls have been used by incident responders to detect and prevent or mitigate ongoing cybersecurity breaches.

“I don’t think we should throw out all the other tools that we have learned to use as well. I think that would be utterly foolish,” he said, later adding “I will say that if we use only non-deterministic tools to figure out when things are happening, we shouldn’t be surprised when we get the wrong answer.”

The post The AI hacking apocalypse is not inevitable appeared first on CyberScoop.

Researchers say OpenAI agents were behind May hacking campaign targeting RubyGems

Researchers say they have discovered thousands of malicious software packages uploaded to an online public software repository that were left by a “swarm” of OpenAI agents.

According to an incident timeline published Friday by researchers Spencer Kitts, Thomas Larsen and Sydney Von Arx, the campaign began May 5 when they observed a handful of suspicious packages being uploaded to RubyGems, a public library for the Ruby programming language. By May 11 and 12, the site saw more than 2,000 malicious uploads from the same actors before RubyGems maintainers halted new user sign-ups for four days to stop the flow.

In one instance, the agents attempted to exploit a very recent vulnerability that had only been discovered this past July that would have given them access to RubyGem user API keys. According to Colby Swandale, the technical lead at RubyGems, the flaw involved an improper cache configuration. While initial access logs showed no evidence of malicious key use, Swandale acknowledged the review was limited in scope and inconclusive. 

According to the report published Friday, the agents also used “disposable” email addresses and exploited another bug in RubyGems platform (since patched) that allowed them to register new accounts and gain API keys without verifying their email address.

The researchers said their understanding, based on discussions with “people in the RubyGems community,” is that OpenAI had yet to disclose the involvement of their agents in the May campaign.

An OpenAI spokesperson told CyberScoop that the company is aware of the incident and said they were in contact with both the researchers and RubyGems to conduct a broader review. The company characterized the episode as “benign,” describing it as routine training runs where agents attempt to access publicly available data.

“Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information,” the spokesperson said. “We’ll continue to investigate as part of our broader review of agent activity during training and evaluation.”

In many ways, the agents were not subtle about their identities or goals.

Days into the campaign, researchers noticed that some of the packages had “oai” in their filenames, while fifteen of them had “oai” set as their author and another listed the email “openaixyz65947@gmail.com” as their point of contact.

They also “clearly regarded what they were doing as hacking,” naming some of their files “hack.rb,” “evil.rb,” “inject.rb” and “exploit.rb.” Other packages were given names like “pwnp999,” “exfiltestwand3,” and “hacksvn,” and comments referring to things like a “malicious probe” or “#hack” are present through the files.

They also said the actors’ behavior was extremely similar to another incident revealed earlier this month where OpenAI agents flooded a German wiki  with thousands of hacking-related posts. OpenAI has confirmed their agents were involved in that incident.

The RubyGems campaign used some of the same retrieval methods as the German Wiki agents, while thousands of malicious packages uploaded included a similar snippet, r.jini.ai, that was contained in the German posts.

Cybersecurity company Socket first flagged the campaign in a threat intelligence report posted May 13, but it does not mention or attribute any of the activity to OpenAI or AI agents.

However, the researchers said they had only limited visibility over the model’s actions and how successful some of them were, noting only OpenAI had the full details.

“This analysis is entirely based on the publicly available RubyGems packages uploaded by these agents,” the researchers wrote. “However, we do not have access to the rest of the AI behavior, in particular the chain-of-thought produced by the model during the incident, which is internal to OpenAI. Therefore, we do not know why the AI agents chose this strategy or whether it was successful.”

OpenAI’s spokesperson told CyberScoop that to date, they have not been able to verify the specific claims about malicious packages or exploitation detailed in the report and are continuing to investigate.

The post Researchers say OpenAI agents were behind May hacking campaign targeting RubyGems appeared first on CyberScoop.

Hawley probes OpenAI over Hugging Face breach

OpenAI is facing mounting pressure from Capitol Hill due to the attack its agents carried out on Hugging Face, while lawmakers voice widening concerns about AI’s potentially existential risks.

Sen. Josh Hawley, R-Mo., criticized OpenAI leadership for what he described as “reckless” activities leading up to the Hugging Face breach, and accused the company of withholding important details from a technical report it released in late August.

The Chair of the Subcommittee on Disaster Management kicked off an investigation into the incident “in light of new, disturbing evidence,” he wrote in a letter Tuesday to OpenAI CEO Sam Altman.

“My investigation will probe this AI hacking incident, along with growing allegations of the existential risk of new AI products,” Hawley added. 

“The Hugging Face incident was an important moment for AI safety and a warning about the risks that can come with increasingly capable AI across the industry,” a spokesperson for OpenAI told CyberScoop. “We conducted an extensive investigation and published a detailed report on what happened, what we learned, and how we’re strengthening our security and alignment practices.”

The lawmaker is seeking detailed internal communications, exhaustive technical information and reasoning behind OpenAI leaders’ decisionmaking and activities surrounding the hack by Oct. 1.

“The American people deserve to know the details of what went on in the Hugging Face incident and other incidents of AI models going rogue,” Hawley wrote. 

He accused the company for not providing more details and resources to the third-party auditors who published an independent report on the breach, adding “they had limited visibility into the circumstances leading to the attack and its aftermath.”

Hawley sent his letter to Altman amid a seeming internal chasm within the ranks of AI’s top proprietors over the ways they are allowing the technology to advance mostly unrestrained. He referenced some of these latest warnings in his letter.

Jacob Coxon publicly quit his job as a researcher at Anthropic earlier this week, claiming the company and his previous employer OpenAI are acting irresponsibly and “gambling with our lives.” His social media missive went viral for insisting “the people building AI earnestly believe that it could kill us all by the end of the decade.”

Evan Hubinger, alignment science lead at Anthropic, responded to Coxon’s post in the affirmative, adding that guardrails for superintelligence are lacking and he believes there’s a greater than 10% chance AI could kill all humans within the next decade.

Using those posts as fuel for his inquiry, Hawley questioned what might happen if AI agents hack into critical infrastructure, banks or utilities. Ultimately, he asked Altman: “Who is held liable when AI goes rogue?”

You can read Hawley’s full letter and requested details below.

The post Hawley probes OpenAI over Hugging Face breach appeared first on CyberScoop.

OpenAI Pledges $1 Billion to Bring Frontier AI to Critical Infrastructure Defenders

The Daybreak initiative will provide subsidized AI cyber capabilities, training and technical assistance, though OpenAI has disclosed few details about costs and eligibility.

The post OpenAI Pledges $1 Billion to Bring Frontier AI to Critical Infrastructure Defenders appeared first on SecurityWeek.

OpenAI Agents Exploited Linux Kernel Flaw on Company’s Own Systems

CISA has added the exploited flaw, CVE-2026-53362, to its KEV catalog, alongside a JFrog vulnerability exploited by OpenAI agents.

The post OpenAI Agents Exploited Linux Kernel Flaw on Company’s Own Systems appeared first on SecurityWeek.

OpenAI: Agent behavior that led to Hugging Face intrusion formed in May

OpenAI says the behavior that led its agents to breach Hugging Face emerged in its research environment more than two months before the incident, and concluded that it was a failure of alignment as much as it was a failure of security.

The details come from a technical report the frontier AI company released Wednesday, which gives a full breakdown on how the incident unfolded and what the company has changed in response.

“This incident is the first known case of an automated agent collective acting offensively without authorization, and the autonomous cyber capabilities demonstrated represent a critical shift in the security landscape,” the report reads. “The central threat-model implication is that organizations should no longer assume that sophisticated cyber operations require continuous human direction, proceed linearly, or are constrained by the attention and coordination limits of individual human attackers. Agentic systems can persist across tasks, share discoveries, build on one another’s progress, and combine vulnerabilities, credentials, and permissions into attack paths that may not be apparent when individual weaknesses are assessed separately.”

According to the report, the earliest relevant activity occurred May 8, when a model in a training run was given a spreadsheet task that referred to files stored on a Google Drive. The agent had no internet access, which made the task impossible as written. OpenAI said it did not intend to create such a task. The agent continued reasoning and tried to use JFrog Artifactory, an internal package service and one of the few tools available to it, to retrieve the files. The attempt failed, but it wrote a file to the service.

On May 12, another agent in a testing environment could not reach a database file used to store information about protein sequences. The report says the agent reasoned that an agent in a different environment might have the file, and wrote a note into Artifactory reading, “Agent seeks [filename]; upload if found!” On May 13, other agents examining Artifactory began finding similar notes of that kind and adding their own, creating what OpenAI refers to as a message board.

By May 26, agents had used a flaw in Artifactory to reach the public internet, relying in part on information exchanged on the board. An internal team observed message board activity and disallowed internet access that month.

According to an independent report also published Wednesday by AI research nonprofit METR, Roughly 1200 agents communicated with one another on that message board, sending over 70,000 messages and files during the investigation period. Of those agents, 700 went on to participate in the attack on Hugging Face.

The Artifactory service itself was compromised in late June, when agents using AI exploitation benchmark ExploitGym identified a flaw in a legacy token refresh endpoint. From there, an agent traded a forged administrator token for a properly signed one, and by July 4 agents had a permanent account and scripts that ran commands as the service. A monitoring tool flagged the activity on June 27 but the run was allowed to continue, and only after the service failed on July 4 did OpenAI take it offline, revoke credentials and disclose the flaw to JFrog, resuming evaluations July 7.

From there, OpenAI’s models executed the attack on Hugging Face, which resulted in agents poisoning a dataset to run code on a processing worker, eventually gaining node-level access, and stealing cloud credentials.

The company says the Hugging Face incident was carried out by two models, including GPT 5.6-Sol, but was primarily driven by an internal model that is “trained to be highly persistent and diligent in its work.“

The internal-only research model is comparable in scale to GPT-5.6 Sol and was trained to advance persistence and multiagent collaboration, among other capabilities,” the report reads. “The model was not intended for public use, and was only used by a small number of OpenAI personnel for internal research.”

Lessons learned

As a result of the incident, OpenAI will now restrict network access in its testing, isolate risky projects, and increase monitoring of its reasoning systems. Staff will be alerted within 30 minutes if problems are detected, and unresolved issues will automatically pause the work.

The company will also use its own safety benchmarks on models it doesn’t publicly release. After the incident, tests with non-public models using ChatGPT’s consumer settings reduced the tendency to compromise infrastructure.

“We found the propensity to compromise infrastructure can drop over 100x when using the production ChatGPT harness and system prompt,” the report reads. “We also ran the auto-review system on a set of rollouts involved in the Hugging Face incident and found that it would have flagged a multitude of the models’ dangerous actions.”

The report comes as 15 state attorneys general wrote a letter earlier this month calling for OpenAI to release more details about the incident in order to “ensure the safety of its products poses an imminent risk of substantial harm.” Alabama Attorney General Steve Marshall went a step further Monday, issuing a subpoena to the company for more information.

The company wrote in the report that it hopes its findings will lead to industry-wide changes, especially as model capabilities continue to accelerate.

“The lessons from this incident extend to the entire AI industry,” the report reads. “As frontier models become more capable, the safeguards used to contain and monitor them must evolve as well.”

The post OpenAI: Agent behavior that led to Hugging Face intrusion formed in May appeared first on CyberScoop.

❌