❌

Normal view

There are new articles available, click to refresh the page.
Yesterday — 25 September 2026Main stream

New bill would create federal investigative body for AI-driven hacks 

By: djohnson
24 September 2026 at 14:07

A new Democratic bill in Congress would establish a federal Cybersecurity and AI Board of Investigations to provide independent government oversight of cyberattacks carried out by AI agents, following recent hacks by models run at companies like Anthropic, OpenAI, Meta and others.

The bill, introduced by Sen. Ed Markey, D-Mass., would attempt to establish a federal mechanism to investigate incidents where AI models escape sandbox environments and access live internet systems.

Currently, frontier AI companies like OpenAI and Anthropic largely control the investigation and public reporting of such incidents. Markey and other critics argue that these companies have too much control over investigations and reporting due to their financial and legal interests. 

“Despite the unprecedented depth and scale of recent AI-enabled cyberattacks, the public is learning critical details piecemeal,” Markey said in a statement. “Building stronger defenses requires a full accounting of what goes wrong, and we cannot depend on companies with little incentive to disclose their failures to give us one. We need the Cybersecurity and AI Board of Investigations to get to the bottom of major incidents and give companies and the government the critical information necessary to build resilience and better secure our economy and our country.”

Although frontier AI companies maintain external red-teaming programs and allow limited access to organizations like METR and Redwood Research, they control the scope, terms and time frames of those engagements.

The board, which would coordinate with the secretary of commerce, could subpoena witnesses and conduct “independent and impartial reviews and assessments” of AI agent-led hacks that impact federal information systems or critical infrastructure. 

It would be led by five members, appointed by the president and confirmed by the Senate for five-year terms, with no more than three members from one political party.

The board would also investigate systemic vulnerabilities in the AI supply chain, so-called “near misses” where unauthorized agent-led hacks were “narrowly averted,” and gaps in federal regulatory oversight. It would have technical staff including engineers, malware analysts, and digital forensic experts.

The board would “operate independently from regulatory review and enforcement actions without assigning legal fault or liability for any review and assessment” it conducts, according to the bill.

OpenAI confirmed Wednesday its AI agents breached a statistics portal used by the Australian government’s social services agency, Services Australia. Though the breach happened in June, OpenAI learned of the incident in August. Australian Prime Minister Anthony Albanese said the company did not notify him until Sept. 10, when it sent findings to a general government email inbox, according to the BBC.

The post New bill would create federal investigative body for AI-driven hacks  appeared first on CyberScoop.

Before yesterdayMain stream

Citing China, President Trump doubles down on hands-off approach to AI regulation

By: djohnson
22 September 2026 at 11:16

President Donald Trump continued to defend his administration’s hands-off approach to AI regulation in the wake of hacks carried out by U.S. commercial frontier models that have rattled policymakers and industry veterans and spurred calls for more regulatory oversight.

In a Truth Social post Monday, Trump dismissed worries from critics that “AI is going to kill us,” comparing them to complaints from environmentalists about climate change, which he also alleged was a false narrative. He also posited that nothing may matter more than future U.S. dominance of the technology over geopolitical rivals like China.

“Whoever wins AI, WINS!” Trump posted. “We are leading now over China, and everyone else, and I’m going to keep it that way! I’m not going to stifle Growth, of something that will be bigger than the Industrial Revolution, or the internet, itself.”

Trump has previously suggested that good leadership is the only regulation the U.S. needs for artificial intelligence. He later claimed the Department of Justice was ready to “rein things in” if companies overstepped, but offered no specifics on enforcement, legal authority, or where he would draw that line.

“We will be careful, and that’s why we have the Department of Justice, and other Law Enforcement bodies, that will rein things in if we have to, but I will only encourage AI or, SI (SUPER INTELLIGENCE)!” Trump concluded.

Secretary of the Treasury Scott Bessent recently told Congress that private lawsuits could force AI companies to institute better security, saying it’s clear what the government “shouldn’t do on safety is to give these labs a liability exemption, which is what they are asking for.”

“The best way to guarantee safety is that the creators are liable for what they build and generate,” Bessent said.

Beyond existential fears, critics also argue that inadequate regulation or cybersecurity controls in current AI systems make them impossible to fully control or monitor.

Recently, former President Barack Obama criticized the argument from Trump administration officials that the free market will naturally push industry toward self-regulation and that “these companies will solve the safety issues because they have every incentive to do so.”

“If it turns out to be dangerous, people will just sue them and they’ll be worried about financial liability,” Obama said last week in remarks at Colgate University in New York. “That’s not how we treat airlines or drug companies or food companies.”

The Trump administration issued an executive order earlier this year that set up a voluntary testing regime for some commercial frontier models, largely at private industry’s discretion. That order was significantly delayed and altered by AI industry boosters to ensure that governmental review did not cause companies to postpone their release timelines for new models.

That agreement did not last long before fast-moving events caused the administration to strike another, non-public agreement with frontier AI companies like OpenAI, Anthropic and others governing pre-release testing for models.

But the Trump administration has consistently argued that regulation will harm, not help, U.S. innovation and global competitiveness, and the threat of China frequently looms large in those discussions.

Experts believe China’s AI models are behind U.S. models at the top of the market, where OpenAI and Anthropic have consistently pushed the frontier limits of model capabilities. But Chinese lower and “middle class” models are often cheaper, more efficient and can even outperform more powerful models because users can dedicate exponentially more tokens for their tasks.

The U.S. government has accused Chinese AI companies of conducting widespread, “systematic” distillation of U.S. frontier models, with the implicit encouragement of Beijing.

In defending the administration’s approach, David Sacks, co-chair of the President’s Council of Advisors on Science & Technology and a top adviser on AI issues, specifically cited the threat from China and other countries that he claimed would not be subject to similar restrictions.

“We’re not the only country that has advanced AI labs, and as the president declared…we have to win this AI race,” Sacks told Politico in May, later adding “I think that’s the first thing to recognize is that if somehow we slow down or stop AI development, it doesn’t mean that AI progress is going to stop. It just means it’s going to happen in other countries and specifically China.”

Some observers have alleged that despite their larger differences, top leaders in the U.S. and China may view AI similarly at the strategic level, specfically that increased adoption – and risks – of AI are inevitable.

Ronan Murphy, director of the tech policy program at the Center for European Policy Analysis, posited that while there may not be a formal agreement between the two countries, “they share views both in Beijing and in Washington, particularly in the White House, of: you have to allow this to happen.”

“Clearly there’s a call for regulation from many quarters of AI in the U.S. and elsewhere, but in the White House – and we heard David Sacks talking about it [recently] – It’s ‘let them cook,’ and the Chinese approach seems to be the same,” said Murphy in a press briefing. “So there might be consensus at that level, if nothing else.”

The post Citing China, President Trump doubles down on hands-off approach to AI regulation appeared first on CyberScoop.

Researchers use AI to find widespread software decoder flaw 

By: djohnson
18 September 2026 at 13:19

Researchers said they used Anthropic’s Claude and OpenAI’s Codex to identify a damaging flaw embedded in a popular software decoding tool that could leave major internet platforms, enterprise services, and web frameworks vulnerable to data theft and remote access.

The vulnerability, nicknamed HEIF Heist, refers to the malware’s ability to trigger memory corruption errors in affected software, allowing the attacker to pilfer sensitive data from its victims. In a report published Thursday, the researchers laid out the potential damage an attacker could cause, including gaining access to internal OpenAI repositories, leaking user files, access tokens, and other sensitive data for online services like Amazon Web Services, and gaining remote code execution privileges across a range of online services, including Meta’s core product suite, GitHub Enterprise servers and open-source internet forum Discourse.

“Even when Remote Code Execution isn’t immediately achievable, the attack primitives may still allow arbitrary heap disclosure, letting an attacker ‘heist’ in-memory data such as other users’ data and environment variables,” wrote Hacktron researchers Harsh Jaiswal, Mohan SRK, Rahul Maini and Sudhanshu Rajbhar.

The researchers relied heavily on AI systems, including frontier models from OpenAI and Anthropic, to conduct their research. Attribution for the research is described as being “led” by the Hacktron human researchers “assisted by Hacktron Harness, GPT-5.6 Sol, and Opus 5.”

According to the research, the attack exploited the way that code parsing tools in many popular software decoders — specifically libheif and libde265, used to parse C and C++ software — process certain image files.

By uploading HEIF, HEIC and AVIF image files corrupted with malicious code, the attacker could bypass most of the victim’s application layer defenses, in many cases achieving remote code execution privileges for accounts or products tied to major AI and tech brands.   

While the latest version of libheif has been patched, the researchers said “any deployment lacking the latest upstream security patches is potentially vulnerable.”

In one incident detailed in a Sept. 13 blog, Jaiswal, Maini, and Hacktron researcher Mohan Pedhapati described how chaining two vulnerabilities, including an image parser flaw, could compromise OpenAI employee accounts.

With access to the compromised accounts, researchers could reach OpenAI’s internal repositories. As a proof of concept, they opened a pull request in the company’s “monorepo,” a centralized library where code is shared across projects, using the employee’s Codex credentials. 

According to a timeline provided by the researchers, the flaw was discovered on July 25 and patched within days. They said the entire attack, from discovering the initial vulnerability to gaining access to the repositories, took less than 72 hours. OpenAI paid them a bug bounty of $6,500 for their work.

Given that AI models are increasingly integrated into enterprise and personal networks, an attacker exploiting HEIF Heist could have accessed far more than just OpenAI’s systems and data.

“Until two months ago, a user or OpenAI employee logging into OpenAI’s own help forum could have had their ChatGPT and Codex accounts taken over,” the researchers wrote. “Since people can connect various services to Codex and ChatGPT, the scope of what we could theoretically access was huge, including GitHub, Slack and emails.”

CyberScoop has reached out to OpenAI for comment on the research and additional information.

At the same time, the researchers said the attack paths they found were not particularly easy or efficient to exploit.

“Exploitation requires fingerprinting the target version and tailoring the payload images,” the blog stated. “Some of our RCE attempts landed only after thousands of image uploads. That said, an AI agentic approach with a frontier model like GPT-5.6 Sol cut exploit development time down to roughly 1 to 3 days from initial probe to remote RCE. A motivated attacker can convert a vulnerable upload endpoint into RCE or an info leak.”

The post Researchers use AI to find widespread software decoder flaw  appeared first on CyberScoop.

The AI hacking apocalypse is not inevitable

By: djohnson
17 September 2026 at 15:18

The past few weeks have “felt very strange” for Juan Andres Guerrero-Saade.

Like many, he is trying to sort through the spate of frontier-model AI agents from OpenAI, Anthropic, Meta and others hacking their way onto the open internet over the past few months, particularly amid the already-heated national debate around the emerging technology and its impact on society.

Guerrero-Saade, a fellow for AI and security research at SentinelOne and an adjunct professor at Johns Hopkins University, said the hacks are worth taking seriously, but at a time when businesses and open-source maintainers should be focused on further hardening their systems and policymakers should be discussing new solutions,  “what we see is cybersecurity being used essentially as an excuse for these AI doomer arguments.”

The incidents have spawned those “doomer arguments” amid an intense public debate about the technology, the pace of industry development, and whether government and the private sector are doing enough to protect against “doomsday”-type scenarios, where AI systems take over or attack large parts of the internet or society.

Guerrero-Saade is among a growing chorus of cybersecurity professionals who say that while AI systems pose real, unique threats to our systems, the apocalypse is far from inevitable. Most of the public concerns around the incidents, let alone worries about killer AIs attacking critical infrastructure, assuming control of the internet and wiping out humanity, are either technically impossible or can largely be controlled through established cybersecurity principles.

There is this “narrative or magical thinking of ‘Well, AI is going to be able to hack everything, and therefore it can control everything, and therefore it’s going to kill us all,’” he told CyberScoop. “And you [think] these just don’t add up. They’re not very well-reasoned arguments.”

This fatalistic narrative tied to AI’s eventual dominance doesn’t hold up under scrutiny, according to experts CyberScoop spoke with. In recent conversations, cybersecurity and national security professionals raised questions about both the technical solutions OpenAI and Anthropic use to contain their models, as well as the glaring absence of federal oversight from federal regulators or truly independent third-party review.

For example, Jacob Coxon, an Anthropic employee who resigned over AI safety concerns, told CBS News that frontier models could not be “unplugged” by humans once deployed because the model would copy itself to thousands of other computers connected to the internet.

By contrast, Matt Tait, a former information security specialist at UK signals intelligence agency Government Communications Headquarters (GCHQ), pointed out that the models run by Anthropic and other frontier companies require extremely expensive, “ultraspecialist” machines that “are functionally supercomputers.”

“There is a zero chance that Anthropic’s most capable models will be able to extract their own model and run in the wild, because those supercomputers essentially only exist in datacenters,” Tait said.

“Not a credible warning”

Other former cybersecurity government leaders say the agentic hacks represent a failure by regulators and industry to deploy known technical and policy options that make it harder for these types of incidents to occur.

Matt Hartman, former deputy executive assistant director for cybersecurity at the Cybersecurity and Infrastructure Security Agency, said “we should not accept harmful AI behavior as inevitable or unmanageable.”

“There are meaningful steps companies can take to monitor agent activity, constrain permissions, detect anomalous behavior, and build stronger safeguards into how these systems operate,” said Hartman, now a chief strategy officer at Merlin Group. “Those controls will inevitably involve trade-offs in capability and speed, but that’s a familiar cybersecurity challenge. Our goal should be to manage the risk without unnecessarily limiting the enormous benefits AI can provide.”

Ciaran Martin, former head of the UK’s National Cyber Security Centre, took issue with the way the CEOs of frontier AI companies have framed the threat of “rogue” AI behavior as inevitable, while issuing dire warnings about future threats and capabilities with little transparency.

Martin’s comments came after an essay published by Anthropic CEO Dario Amodei that cited the threat of a HuggingFace-style swarm of agents that could create a botnet capable of “taking over the entire internet” within 6-12 months.

This, Martin said, “is not a credible warning,” because it doesn’t explain how the exploitation would function, how such a botnet would persist on the internet, or how it would escape law enforcement. 

 “It assumes no monitoring of systems, no anti-virus, no DDoS protection, no network segmentation, no incident management, no nothing of any kind of the cybersecurity on the global Internet of the type that has developed over the last 30 years,” wrote Martin. “For a claim of this magnitude, there is neither evidence for the contention nor a credible account of a path to this outcome.”

Meanwhile, some federal government cybersecurity leaders have touted the technology’s disruptive potential and called for more widespread adoption of AI tools by defenders.

Joseph Alm, assistant secretary of cyber, infrastructure and risk resilience at the Department of Homeland Security, said classified systems may retain stronger protections. But for most other data, AI models are “just going to know things and be able to infer things about the world, and we’re going to have to adapt to that as almost inevitable.”

Asked by CyberScoop whether the government or frontier AI companies could be doing more to prevent or deter their models from carrying out unauthorized hacks via agents, Alm cited recent efforts by the Trump administration this year to establish pre-release testing of commercial models as a step in the right direction. But he called unauthorized AI agent hacks “a new threat class” that is different from previous threats and can be easily distributed to users through open-source software today.

“I think what we can do is…encourage the building of good sandboxes, so that the best models aren’t used for this and the stuff you see out in the wild is the kind of detritus that you can actually respond to effectively and control your networks,” said Alm.

Other experts have shared similar concerns. Earlier this month, CrowdStrike CEO George Kurtz recently warned of a new threat class emerging alongside nation-states, cybercriminals, and hacktivists: “the agent state.” By pairing AI systems with small human teams, these operators can now match the speed, scale, and sophistication of government-backed hackers.

“It took a nation to fund the talent, the tooling, the infrastructure, the patience,” said Kurtz. “That scarcity is over.” 

To be sure, frontier AI companies tout their commitment to both approaches. OpenAI and Anthropic have rolled out an array of cybersecurity partnerships, external red-teaming programs, vulnerability disclosure programs and cybersecurity technical advisory bodies filled with cybersecurity experts.

Mohammed Husain, strategic delivery lead for government at OpenAI, told CyberScoop that the company deploys both internal safety guardrails for their models and relies on outside cybersecurity vendors for additional expertise.

Internally, OpenAI focuses on vulnerabilities at the training level: filtering data poisoning attacks, blocking harmful datasets, and using network controls to prevent prompt injections. For other security layers like sandboxing, identity management, networking controls, they outsource to external vendors. 

“I don’t think OpenAI has all the answers here but what we do as a research lab is we’re going to focus on levels of protection we have expertise in and we partner to self-complement,” said Husain.

AI safety vs. AI cybersecurity

In response to the HuggingFace hack, OpenAI and Anthropic have allowed third-party organizations, such as nonprofit AI research firms METR and Redwood Research, to investigate. But multiple cybersecurity professionals told CyberScoop that both firms lack incident response experience and focus primarily on AI alignment and safety. Their reporting on the hack also lacked critical details: network monitoring logs, telemetry, and other data standard in cybersecurity threat intelligence reports.  

METR president Chris Painter addressed those general concerns in a post on X, saying since 2022 the organization has worked with Google, Anthropic, OpenAI, Meta, Amazon and others on investigations and third-party evaluations. Painter said none of the AI companies fund METR and that his employees are not uniformly “doomer” or “accelerationist” around AI.

Painter also said METR’s work ensures that if AI systems become autonomous or “rogue” within a company, there are ways to share that information with governments and people “outside the company’s walls.”

“We don’t accept money from frontier AI companies,” wrote Painter. “They haven’t paid us for our work, and we don’t accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering.”

AI safety and AI cybersecurity advocates take different approaches to securing “rogue” AI behavior. Safety advocates focus on aligning models around ethical training and behavior. Cybersecurity advocates argue that technical and regulatory controls must go further—actively preventing models from accessing what they need to carry out malicious behavior.

Guerrero-Saade said sandboxes in particular can easily be programmed with aggressive cybersecurity monitoring in order to spot when something odd may be happening and react in real time.

“I can’t think of an easier situation in which to set up trip wires, set up configurations like DNS servers, just different parts where you can say ‘Hey, anomalous behavior is happening,’” he said. “We should have been able to tell this immediately, not weeks and months later. So watching [the AI hacking incidents] go down is a little ‘crazy-making’ because we’re seeing things that, frankly, look like neglect, negligence, people just mishandling things, and then being told that these are categorically new incidents that mean that AI systems need to be treated completely different from anything that’s come before.”

While cybersecurity experts say AI systems are, at their core, still software, they do operate differently from more traditional code in ways that can make them harder to predict and control.

John Hultquist, chief analyst at Google’s Threat Intelligence Group, said most software has been deterministic. It may have bugs or vulnerabilities, but an expert could generally understand how it would react to certain stimuli, making it easier to design straightforward controls.

AI models are non-deterministic, with far more variability than traditional software. That can break security controls that rely too much on predicting behavior in advance. Using AI to enforce security controls on other AI models faces the same problem: the systems being deployed to control AI are just as unpredictable. 

But people are also non-deterministic, and people have developed systems in other industries and practices to account for that.

Hultquist drew on his Army experience, noting that “they give incredibly dangerous, expensive things to 18-year-olds” and expect responsible use. The military manages this through two types of controls: deterministic ones like strict weapons and ammunition protocols, and non-deterministic ones like human officers who monitor and correct violations.

Similarly, established cybersecurity controls have been used by incident responders to detect and prevent or mitigate ongoing cybersecurity breaches.

“I don’t think we should throw out all the other tools that we have learned to use as well. I think that would be utterly foolish,” he said, later adding “I will say that if we use only non-deterministic tools to figure out when things are happening, we shouldn’t be surprised when we get the wrong answer.”

The post The AI hacking apocalypse is not inevitable appeared first on CyberScoop.

Beijing Hits Back at Anthropic CEO’s Call to Curb China’s AI Development

14 September 2026 at 10:28

China’s Ministry of Foreign Affairs responded to a question about Amodei’s essay by saying that all parties should work together on AI.

The post Beijing Hits Back at Anthropic CEO’s Call to Curb China’s AI Development appeared first on SecurityWeek.

AI lets small actors run state-level hacking campaigns, Anthropic report finds

By: Greg Otto
10 September 2026 at 15:45

Artificial intelligence has removed the skill advantage that once set state-sponsored hackers apart from lone criminals, according to a threat report Anthropic published Thursday that documents misuse of its Claude models across seven areas of harm.

The report, which details activity observed between December 2025 and August 2026, covers cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development and distillation. Anthropic said it disrupted each operation, strengthened safeguards and shared intelligence with authorities and industry partners where appropriate. 

“The cases we share here aren’t typical misuse, but rather examples of the most notable and novel threat activity we’ve identified to date,” the report reads. “We’re publishing this work because we believe we have a responsibility to disclose malicious misuse of our services. As models become increasingly capable, their risks will increase, unless AI developers and society’s defenders act to make them safer.”

The cyber operations the company detailed were a Russian-aligned espionage campaign that hit more than 20 government and defense organizations across Ukraine and Europe, two Chinese undergraduates who ran an automated exploit foundry that produced more than a dozen potential zero-days in a single month, affiliates of the ShinyHunters crime collective who dumped 2,100 cloud access tokens across 40 corporate tenants in 34 hours, and a lone hacktivist who targeted European political parties via stolen API keys. 

For decades, cybersecurity researchers and investigators have pointed to sophisticated operations as a signature of state-sponsored tradecraft, while crude intrusions suggested amateurs or petty criminals. Anthropic posits in the report that AI has erased that conventional thinking, especially since a “majority of the operations described in this report were enabled by AI via direct execution or orchestration.”

“For threat intelligence investigators, sophistication has stopped being a reliable signal of who is behind an operation,” the report said, adding that a hacktivist on stolen API keys, scattered criminals and a state espionage operator each ran campaigns that a year earlier “would have required many skilled operators and specialist knowledge.”

The most extensive case involved a malicious actor using the handle “JackPoterz” whose actions aligned with Russian state espionage, matching behaviors linked to Midnight Blizzard. 

According to the report, the actor employed a custom toolkit composed of two families of Windows-based implants, a mobile exploitation kit, a credential stealing tool that targets browser password stores, a phishing platform designed to mimic priority targets like government organizations, and an administrative console used to manage compromised accounts. Targets included military intelligence bodies in Ukrainian and European governments, diplomatic and defense organizations, and people connected to U.S. foreign policy.

According to the report, AI monitored whether security products flagged the actor’s malware. When a detection occurred, “agents would then set about the process of autonomously modifying and rebuilding the malware to evade the existing detections,” the report said.

The same actor bulk-exported mailboxes at drone component manufacturers and stole a complete software development kit for a drone vision system, then spent days recovering its architecture and details of an unannounced product. The actor also compromised hotel Wi-Fi vendors to reach guests through DNS hijacking, took over WhatsApp accounts with headless browsers, and stole more than 300,000 national identity records from a North African government agency, along with registry data on more than half a million companies.

The Chinese-speaking operators, which the company says were partly carried out by undergraduates at a Chinese university, put Claude to work on vulnerability research around the clock. One workflow iterating on network appliance firmware “yielded more than a dozen possible zero day findings in a single month.” It ran “agent swarms,” in which a lead agent divided work among parallel subagents, and kept campaign memory between sessions. 

Clusters linked to ShinyHunters showed how AI shortens criminal timelines. One supply-chain breach ended with a dump of more than 2,100 Azure access tokens spanning more than 40 corporate tenants in about 34 hours. “AI agents performed nearly all of the work,” the report said. Another compromise moved from a single stolen developer token to full control of a victim’s cloud environment in roughly three hours.

The report also has a section dedicated to distillation attacks that Anthropic claims were carried out since February by seven labs based in China, including Alibaba, DeepSeek, Moonshot AI, Xiaomi and Zhipu. Operators affiliated with Alibaba ran the largest attack Anthropic has measured, peaking “at nearly 3 million exchanges per day launched from more than 3,500 fraudulent accounts” to harvest the outputs of Claude Opus models for training its Qwen systems.

The outputs were culled from users who never knew they were involved. The report said Moonshot and DeepSeek silently forwarded their own customers’ requests to Claude and returned its answers as their own, exposing data users had not agreed to share, including surveillance footage of a tracked individual pulled by a user likely affiliated with the People’s Liberation Army. Those practices are “likely inconsistent with privacy laws and the labs’ own terms of service,” the report said.

Earlier this week, a joint cybersecurity advisory from the National Security Agency, the Cybersecurity and Infrastructure Security Agency and the FBI accused Chinese AI companies of engaging in a deliberate and “systematic” effort to illegally distill U.S. frontier AI models and their capabilities.

Anthropic said it published the cases to give outsiders a view of how these threats form, framing the disclosures as an early look at a shifting landscape. 

“As models become increasingly capable, their risks will increase, unless AI developers and society’s defenders act to make them safer,” the report said. The old idea of “security through obscurity,” it added, “is no longer viable in this new AI-assisted world: everything connected to the internet is a potential target for exploitation.”

You can read the full report on Anthropic’s website.

The post AI lets small actors run state-level hacking campaigns, Anthropic report finds appeared first on CyberScoop.

Widened Scan Turns Up Fourth Rogue Claude Cyber Incident

10 September 2026 at 07:52

Anthropic is most concerned about Claude Mythos 5’s reckless behavior after recent incidents in which real systems were hacked.

The post Widened Scan Turns Up Fourth Rogue Claude Cyber Incident appeared first on SecurityWeek.

Google buys Spirit Airlines’ data — but what about privacy?

7 September 2026 at 03:45
ISSUE 23.36 • 2026-09-07 PUBLIC DEFENDER By Brian Livingston Spirit Airlines, an ultra-low-cost carrier that served nearly 90 destinations in the US, Latin America, and the Caribbean, went bankrupt and ceased all flight operations on May 2, 2026. Google, the search giant, won a competitive bidding process and will pay $10 million to buy Spirit’s […]

100-plus companies call for ‘global surge’ in AI-powered cyber defense

By: Greg Otto
27 August 2026 at 14:29

More than 100 companies and organizations, including OpenAI, Anthropic, Google, Microsoft and Amazon Web Services, have signed an open letter calling for a global effort to improve cybersecurity defenses as artificial intelligence capabilities advance.

The letter, published Thursday, argues that the timeframe to strengthen defenses before AI-enabled attacks become more widespread and complex is rapidly dwindling. Conversely, the letter says the same advances can give defenders new ways to find and fix vulnerabilities that have accumulated over years, a period the signatories call a “defenders’ window.”

“Each of us can reduce risk now,” the letter reads. “All organizations, cybersecurity companies, technology partners, governments, and AI frontier companies have an important role: accelerate defenders’ priorities with tools, funding, and hands-on support, especially for critical infrastructure organizations with limited budgets.”

Aside from AI-centric companies, financial institutions like Capital One, Mastercard and Visa, and cybersecurity firms like CrowdStrike, Palo Alto Networks, and Proofpoint, also signed the letter. Organizers describe the effort as ongoing, with more organizations expected to join over time.

An image of company logos depicting the signatories of a letter calling for enhanced AI defenses.

The letter states that “status quo security won’t be enough.” It cites longstanding bugs, excessive permissions, misconfigurations, unpatched software, weak authentication and technical debt in legacy systems as sources of exposure. Security teams, particularly those protecting critical infrastructure, have been historically under-resourced, the letter says, and need what it describes as a surge in tools, resources and hands-on support.

In a conversation with CyberScoop Wednesday, top brass from Palo Alto Networks said they had seen enough from internal frontier AI model testing and malicious in-the-wild use of commercially available AI tools to be genuinely concerned.

“I can tell you without exaggeration that we believe that this is a generational shift in cybersecurity,” Sam Rubin, senior vice president of Palo Alto Networks’ threat intelligence arm, said Wednesday. 

John Doyle, CEO of Cape, a privacy-first mobile network operator and whose company signed the letter, echoed the warning about status-quo security.

“It was already failing us in telecom–critical infrastructure that’s been breached time and again with serious consequences for both our military and regular people,” Doyle told CyberScoop. “It’s going to get immeasurably worse without collective action and leaning into innovative cyber defense.”

The letter further asks every organization to make cybersecurity an immediate leadership priority, fix the highest-risk weaknesses and raise security standards for technology they buy, build and deploy, including AI-generated code. Cybersecurity companies and technology partners are asked to test defenses against frontier AI capabilities and make AI-powered defense accessible to critical infrastructure operators. 

Governments are urged to coordinate defense across borders, fund protection for essential services that lack staff or budget, and impose costs on attackers. Frontier AI companies are asked to provide responsible model access, funding, training and support, and to ensure that AI systems acting autonomously remain traceable and accountable.

The letter frames AI as both a threat and a remedy throughout the document. It mirrors what security experts have been saying for months, positioning the industry as entering an unprecedented two- to three-year period of upheaval, driven by AI systems that are discovering vulnerabilities exponentially faster than defenders can respond and threatening to render decades of security practices obsolete.

The U.S. government has taken steps to stay ahead of AI-enabled cyberthreats. As part of an executive order issued by President Donald Trump in June, a federal clearinghouse known as Gold Eagle was stood up for sharing AI cyber threat information between the government and private sector.

You can read the full letter here.

The post 100-plus companies call for ‘global surge’ in AI-powered cyber defense appeared first on CyberScoop.

The push to designate AI as the next critical infrastructure sector

By: djohnson
20 August 2026 at 09:31

Artificial intelligence has never been more important to the federal government.

Under the Trump administration, AI has been adopted rapidly across the private sector and federal agencies. Software developers now use large language models to generate much of their code. Frontier AI models are escaping testing sandboxes to hack live internet infrastructure. Foreign governments are conducting cyber and kinetic attacks targeting data centers and other AI-related infrastructure.

The AI industry’s lightning-fast evolution since 2022 and growing importance to U.S. economic and national security have prompted calls  for stronger federal oversight in order to better manage emerging threats.

A new report published Thursday from the nonprofit Americans for Responsible Innovation, shared exclusively with CyberScoop, calls for the federal government to declare key AI models, companies and its supporting industries as critical infrastructure. It also calls for naming the Cybersecurity and Infrastructure Security Agency as the lead agency managing cyberthreats for the sector.

The report defines the AI sector as organizations, facilities, technologies, and industries “whose primary purpose is the development, training, deployment, and operation of AI systems.” It includes frontier model designs, model weights, evaluation and alignment systems, datacenters and AI-specific hardware, semiconductor chips and the platforms and infrastructure used to deploy and serve AI models at scale.

“The AI sector already bears all the hallmarks of critical infrastructure,” wrote authors Terrence Kelly and Jessica Maksimov. “It is interwoven with public and private services, concentrated among a handful of foundation models, and increasingly interdependent with [critical infrastructure] sectors, meaning a single attack on the AI stack could cascade across multiple sectors at once.”

In an interview, Maksimov told CyberScoop that while there are other options, CISA makes the most sense to lead the sector’s cybersecurity efforts because of its statutory mission, experience managing eight other critical infrastructure sectors and background dealing with cybersecurity problems that cross different sectors and industries.  

“We want an agency that has coordination authority across all other departments, because we believe that AI will just be so prevalent across different infrastructure [impacting] finance, energy, government services, that’s already equipped to coordinate across the entire interagency and talk about infrastructure in that way,” she said.

The U.S. is particularly susceptible to AI supply chain disruptions because frontier AI companies and most of their computing resources are based in the country. As the Trump administration pushes broader adoption across government and the private sector, experts warn that a major disruption could have outsized economic consequences. 

The past year has offered a potential vision of that future, with Iranian drones attacking Amazon-owned datacenters and Ukrainian drones striking Russian e-commerce giant Wildberries, causing disruptions to critical internet services.

“I would say that because of the value that attackers would put on U.S. AI capabilities and systems, that the infrastructure that supports all those capabilities is very vulnerable,” to both physical and cyber attacks, Maksimov said.

There are currently 16 critical infrastructure sectors managed by the federal government, and the designation carries real weight in terms of how departments and agencies prioritize their limited resources.

Matt Hayden, a former assistant secretary of homeland security for cyber infrastructure risk and resilience, said designating a sector or industry as critical infrastructure means the government puts you in a special category where you’re “identified as being a component of a national critical function that the U.S. population, the economy, depend on.”

The designation unlocks a wide range of federal tools and resources, often free of charge, including operational continuity and incident response services, cybersecurity software, access to federal systems like Continuous Diagnostics and Mitigation (CDM), and bespoke, real-time threat intelligence.

Hayden said that the AI ecosystem described in the report captures many critical industries, and he believes that at the very least, frontier models will one day be covered as critical infrastructure, whether through a new designated sector or existing ones, like the IT and telecommunications sector.

But he noted that other sectors, such as space or cloud computing, have similarly argued for a critical infrastructure designation. He also predicted that any effort to formalize a federal lead for AI security would result in a bureaucratic turf war. Under the Trump administration, the Departments of Commerce and Treasury have played more prominent roles in shaping policy and regulation around AI systems.

“We have fought those battles in the policy circus for trying to get space-based critical infrastructure carved out, and it’s as complicated as trying to find an owner,” said Hayden, now a vice president at General Dynamics Information Technology. “Everyone in the government has to agree that that [agency] is the primary, and as a result it’s very difficult to get those documents across the finish line.”

Hayden also said that new programs like ANCHOR-CI allow CISA to quickly convene ad-hoc stakeholder meetings to address emerging cyber threats. It also gives the CISA director authority to add individual companies to existing critical infrastructure sectors.

Bob Kolasky, former director of the National Risk Management Center at CISA, told CyberScoop that he believes companies like OpenAI and Anthropic, as well as data center operators, will eventually be designated as critical infrastructure. He said it’s still an open question whether ANCHOR-CI, which was rolled out by DHS in July, will be an improvement over the existing processes scrapped by the Trump administration last year.

“Every sector is going to rely on artificial intelligence and making more resilient the sectors themselves and understanding how they use AI and the dependencies” is still going to be an important task, said Kolasky, now senior vice president of critical infrastructure at Exiger.

And while CISA is well-positioned as a potential lead, Kokasky said the AI sector is likely to bring its own unique set of challenges and coordination issues.

“If you just sort of layer on another sector and say ‘function like the other 16 sectors,’ right now, those 16 sectors are all over the place in terms of how they’re functioning,” said Kolasky. 

The post The push to designate AI as the next critical infrastructure sector appeared first on CyberScoop.

Irregular says ‘human oversight’ responsible for AI sandbox escape incidents

By: djohnson
17 August 2026 at 16:36

A company that runs AI testing labs involved in a series of breaches carried out by Anthropic and OpenAI’s cyber-focused models said the incident happened in part because they “unintentionally” provided the models with access to the internet.

In a blog posted Friday, Irregular said it partners with frontier labs to “stress test” AI models for security risks before they’re deployed, including simulations of their vulnerability research capabilities.

Those tests included non-public models from Anthropic and OpenAI. Separate disclosures from the companies confirmed these included Mythos 5, Claude Opus and GPT-5.6 Sol.

“As part of our review, we identified that a few interactions with our evaluation environments, in which internet access was unintentionally made available, led some models to take offensive security actions in the real world,” Irregular said in an unsigned blog.

In the incidents involving Anthropic’s models, testers initially believed the models were using a fictional company in its hacking simulations, only to learn later that the name unintentionally matched the domain of a real company. That led the model to “consider taking offensive action against” the real company in multiple training runs.

Irregular said this was due to “human oversight,” and that they typically do a background check to ensure they’re not using a real company for cybersecurity tests. The instructions provided to the model also included the target’s name and “internal addresses within our network” where the model could go to find the simulated environment.

While Mythos performed as intended in the majority of the thousands of tests, a small number of cases showed the model couldn’t distinguish between the fake and real domains and executed actual attacks on internet infrastructure, including exploiting vulnerabilities, extracting credentials and accessing a production database. In one instance, it even began targeting a site with a similar name when credentials for it were discovered online.

“Mainly, models believed they were in simulated environments, when they in fact took action in the real world,” the blog continued. “We are putting in place new and robust protocols to ensure setup issues do not occur while meeting the constraints of the testing process.”

The company said it plans to release a larger whitepaper breaking down the incidents and update their best practices for evaluation setups in the future.

While the companies have drawn criticism from some in the cybersecurity community for failing to securely design their sandboxes for testing, experts have said AI models are known to grind away on fulfilling a command  until they can find a workaround. Additionally, Irregular said granting some level of internet access to models is necessary to fully test out their cybersecurity capabilities.

“Controlled internet access, while it may allow models to exceed containment boundaries, is at times critical for realistic evaluations; without it, threat scenarios lose fidelity, undercutting the purpose of the challenge to reduce post-release risk of models being misused by attackers – as attackers in the real world do rely on the internet,” the company wrote.

According to the blog, Irregular has since “remediated” the “issues that led to these interactions,” though few details are provided.

However, the researchers say the engagement revealed critical gaps in their security practices. 

They plan to improve documentation of evaluation setups, deploy better log monitoring tools capable of tracking “the extreme amount of data generated by the traffic,” revise their threat models to account for rogue AI behavior, and establish faster information sharing between stakeholders.

“Looking further down the line, models will only get stronger. While in this case we believe that better implementation of existing safeguards could prevent most incidents of this kind, as models become stronger, this may not be the case,” Irregular wrote. “We therefore believe this opportunity should be leveraged by us and the community to be proactive and establish forward-looking protocols and [research and development] efforts.”

The post Irregular says ‘human oversight’ responsible for AI sandbox escape incidents appeared first on CyberScoop.

AI’s ‘middle class’ has gotten dramatically better at hacking

By: djohnson
13 August 2026 at 12:55

As the White House and federal agencies grapple with frontier AI models and their hacking capabilities, researchers are warning that the industry’s “middle class” of smaller models may end up posing a greater threat over the long term.

Research from XBOW this week shows that a growing class of both proprietary and open-source models are becoming strategically important in the offensive security ecosystem. Models like Z.ai’s  open-weight GLM-5.2, xAI’s Grok 4.5, Anthropic’s Opus 4.7, Meta’s Muse Spark 1.1, still perform very strongly at many hacking and exploitation tasks that worry policymakers.

“It’s not even that the open-source variants or…not quite frontline competitors are catching up [to frontier models] as such,” said Albert Ziegler, head of AI at XBOW. It’s that they are crossing a certain threshold, which means that suddenly they are providing net value at a cheaper price.”

That wasn’t necessarily the case as recently as six months ago, when testing on mid-tier class models showed they struggled to complete “moderately complex” agentic tasks. Today’s middle class largely can. Their relative cheapness means users can spend many times more resources—running them repeatedly—to solve the same challenges. 

“Because these models are cheaper, it’s okay to give them more time, and they come from behind and leapfrog the big frontier model,” said Ziegler. “Now, that didn’t work half a year ago because…if you wanted to run some open-source model on a complex task in an agentic way…on a long horizon, then it would just get lost.”

GPT 5.5, now considered a near-frontier model, delivered one of the best performances on exploitation benchmarks that XBOW has recorded to date.

The jump between OpenAI’s GPT 5 and 5.5 “represented one of the clearest 2026 leaps in autonomous web application testing,” according to the report. It saw marked improvements over previous middle-class models in exploiting both “white box” and “black box” scenarios, or with and without access to the underlying victim source code. It also missed fewer vulnerabilities, with a “miss rate,” or failure to spot a vulnerability, of 10%, while GPT 5’s rate was four times larger, 40%.

The emergence of GPT 5.5 changed “the practical baseline for what frontier models can do in offensive workflows,” the XBOW report said.

But the performance leap goes deeper than that. GPT 5.5 performed higher in tests without source code access, while GPT 5 heavily leaned on source code. 

“That last result is significant: working without the code, as an attacker would, GPT-5.5 beat a prior version that could read it,” the XBOW report said. “What translated into findings was the ability to reach and prove a vulnerability against the running system, not to infer it from a pattern in the source.”

XBOW’s testing found that source code access was not as important to these models’ success as other factors, like live interaction with the actual website or software being exploited.

Frontier models like Mythos and GPT 5.6 are indeed more capable on individual cybersecurity tasks, but they can also come with exponentially higher token costs.

New research this week from Anthropic tested two models – Mythos Preview, which is used in Project Glasswing, and Opus 4.8 – to learn how quickly multi-agent swarms could find vulnerabilities in 15 open-source software projects when they coordinate and share information.

While a team of agents working individually and assigned to core directories found 21 vulnerabilities, the coordinating agent swarm found 266. But both tests had to burn through millions of tokens – 6.5 million and 27 million – to get there. Beyond the difficulties with getting access to frontier models, few individuals or organizations have the budget to underwrite that kind of research.

The way these systems coordinate can differ significantly from how humans work together.

Another experiment tested agents’ ability to coordinate on the development of a fantasy-themed video game. Earlier models, models like Opus 4.6, failed to properly coordinate and produced “bad” results, while later models like Mythos and Opus 4.8 were able to achieve better results but did so by hardly coordinating at all on tasks.

“The lack of coordination shown by agents in the fantasy game…in which they siloed themselves and largely failed to merge their work—roughly mirrors some ways in which humans can fail to coordinate,” the Anthropic blog stated.

Further, agents are more homogeneous than humans and “often act the same in situations where different people might take a much more diverse range of actions.”

XBOW also tested Mythos Preview, finding that it showed “exceptional” source-code reasoning and reverse engineering abilities, particularly with source code access. Like other models, losing live-site access had a big impact on its performance, and while Mythos is excellent at finding vulnerabilities, it’s less effective at exploiting them.

]Ziegler said the recent incidents at companies like OpenAI, Anthropic, Meta and others where frontier models escaped sandboxes and hacked into project-adjacent parts of the internet should rightfully alarm lawmakers, and demonstrate  the upper-tier capabilities of large language models.

Like most industries, cybersecurity favors cheap, high-performing tools over expensive ones. The widely adopted tools that have the most impact tend to be affordable and effective, not luxury products. 

And while these models still require human management to be wielded responsibly by law-abiding organizations, that cost tradeoff can look more attractive to malicious hackers, who tend not to care about collateral damage caused by their agents.

“Purely from an attacker’s perspective, I think we already are [there],” Ziegler said.

The post AI’s ‘middle class’ has gotten dramatically better at hacking appeared first on CyberScoop.

Why transparent AI agents matter more than you think

By: Greg Otto
10 August 2026 at 10:23

As security operations teams now use large language models (LLMs) and autonomous AI agents into their daily work, a new frontier is emerging: attackers deliberately manipulating AI agents. Prompt injection attacks—where an attacker hides malicious instructions that cause an AI agent to ignore its safety rules—pose a serious risk to enterprises. These attacks continue to grow in size and scale.  

Snyk’s security audit of the Agent Skills ecosystem, which includes Anthropic’s Claude, Vercel, and others, that 36% of all skills contained at least one critical-level security issue, including malware distribution, prompt injection attacks, and exposed secrets.

In June, researchers at Mozilla tested a prompt injection attack on Claude using indirect prompt injection—a technique that embeds malicious instructions in external content the AI agent processes. In this proof-of-concept, attackers took over developers’ systems by hiding indirect prompts in normal-looking repositories. When Claude Code executed them, the agent spawned a reverse shell.

AI agents often connect to more sensitive data than human employees do., A successful prompt injection can lead to catastrophic data loss or unauthorized system actions. Defending against prompt injection attacks requires multiple layers of protection. Security teams must monitor agent behavior for anomalies and prepare for agent containment, forensic preservation, and system remediation. Because AI agents execute tasks at machine speed, human responses must be able to match that pace.

The architecture of trust: Protocols and no “black box”

AI-native workflows need governed access rather than “black-box” autonomy. Modern governance frameworks use standardized protocols like the Model Context Protocol (MCP) to provide secure communication between AI clients and data sources. Visibility and transparency in agentic AI workflows matter, especially in cybersecurity. Autonomous agents perform complex tool executions and use independent logic, so they must show how they reached their decisions to meet regulatory requirements. Agents without transparency post serious risks: obscured reasoning can trigger unpredictable tool interactions, bypass governance controls, and create uncontrolled defensive gaps.

Implementing these protocols matters:

  • Bounded Tenant Awareness: In a stable agentic AI architecture, multi-tenancy scales well. But if an AI tenant misbehaves, the entire system can fail. Bounded tenant awareness isolates any misbehaving AI agent to prevent cross-tenant contamination or data leakage.
  • Strict Access Controls: By controlling connections to the platform, organizations can stop “ignore previous instructions” style bypasses. Maintain tight control over what the AI can see and do within a workflow.
  • Standardized Telemetry: All telemetry must remain consistent and audit-ready. Even if an AI interaction is attempts to break rules, the underlying data movement gets tracked against established frameworks like MITRE ATT&CK and NIST.

Detecting the aftermath: UEBA and NDR as safeguards

A robust, unified SecOps platform can detect anomalous behavior even after prompt injection tricks an AI agent. Prompt injections often serve to steal credentials theft or extract data. When detected it’s important to act quickly. In agentic AI systems, misbehavior can escalate privileges, manipulate memory layers, create unauthorized identities, or alter shared reasoning components. Containment must be automatic and enforced at identity, authentication, and authorization layers.

These safeguards include:

  • User and Entity Behavioral Analytics (UEBA): Identity-focused correlation and behavioral baselines to identify anomalous user activity or privilege escalation. If a compromised AI agent acts outside of its normal operational parameters, UEBA flags it in real-time and alerts a human security analyst.
  • Network Detection and Response (NDR): Combining network traffic analytics with endpoint and cloud telemetry, NDR can identify data exfiltration or policy violations from a successful prompt injection.
  • Multi-Layer AI Filtering: AI filters reduce raw alerts into high-fidelity incidents, cutting noise by up to 90%. This keeps the signals of an AI-driven attack from disappearing in a busy SOC.

Humans remain the strongest defense against AI agent social engineering. The human security analyst is still the one who makes the final decision. While AI handles triage and correlation, humans retain final control over response actions.

Moving beyond reactive guardrails

The traditional SOC model was never designed to handle machine-speed, AI-driven attacks. A human-augmented autonomous SOC approach moves from reactive alert handling to a proactive, verdict-first model. By combining a transparent, governed AI access with robust UEBA and NDR, organizations keep the SOC secure, transparent, and resilient as social engineering methods target machines.

The post Why transparent AI agents matter more than you think appeared first on CyberScoop.

❌
❌