Reading view

There are new articles available, click to refresh the page.

The FTC wants to regulate AI for ideological bias 

The Federal Trade Commission wants to start regulating ideological bias in AI systems and assert federal control over state laws. They’re getting an earful from opponents on all sides of the political spectrum.

In a proposed policy statement released last month, the FTC said it was considering treating ideological bias in AI systems as an “unfair and deceptive practice” under Section 5 of the FTC Act.

The commission argued that consumers have an expectation that AI systems will provide them with information free from bias or ideological manipulation. Defining such bias as an unfair or deceptive practice would potentially allow the commission to regulate training or inputs that power AI algorithms. How precisely the FTC would determine when ideological bias exists in these systems is not fully explained in the document. 

Additionally, the statement suggests that the FTC believes this regulatory authority supersedes state AI laws. It specifically mentions the Colorado AI Act, which calls for models to be subject to risk assessments, transparency disclosures and “bias audits” before release. State lawmakers are now seeking to delay or eliminate the audits before the law takes effect in 2027.

CyberScoop reviewed dozens of public comments criticizing  the FTC’s proposal. Even ideological allies raised two main concerns: first, that the proposal distracts from real questions about the federal government’s role in regulating AI deception; and second, that it opens a Pandora’s Box by enabling political censorship of AI model outputs.

Leah Siskind, a former White House digital official and deputy director of the AI Corps at the Department of Homeland Security, told CyberScoop that AI companies face legitimate questions about their obligations to consumers, particularly whether they must ensure their models provide accurate information and protect against deliberate manipulation. 

Siskind’s past research has focused on how authoritarian propaganda tends to be overrepresented in answers provided by large language models, in part due to governments’ intentional efforts to poison data ingested by AI systems.

“There is a really interesting debate here about bias and about accuracy in models and whether that’s deceptive or not… about how we counter disinformation that has been absorbed and is now being reflected by LLMs…but this is not addressing that at all,” said Siskind, now a senior AI fellow at the Foundation for Defense of Democracies.

Instead, Siskind said the FTC statement appears primarily concerned about a power struggle with states over AI regulation and “petty squabbles about which AI model is more woke than the other.” She’s skeptical that the policy statement’s cited legal authorities are on sound footing.

“The way I see it is that the FTC’s role is to police consumer protection violations, not regulating AI systems, and it seems like they’re trying to solve a lack of congressional AI regulation by stretching section 5 [of the FTC Act] well beyond its traditional role,” she said.

Additionally, the policy statement’s language and sourcing suggests that the FTC is concerned with certain kinds of ideological bias more than others.

Anthropic, which has clashed with the Trump administration over AI guardrails and military applications of their technology, shows up more than half a dozen times in footnotes, many which are framed as examples of ideological bias the FTC is seeking to stamp out.

By contrast, the statement ignores a direct example of an American AI company owner influencing their model’s ideology: Elon Musk and his xAI-owned Grok model. Musk has publicly admitted, often on his own website, to intervening when Grok’s responses upset him. These interventions have shaped Grok’s outputs on specific topics, including South African race relations and the term “MechaHitler,” where the model now reflects Musk’s personal views.

But neither Musk and xAI are mentioned in the document, while Grok appears in a footnote which cites an advertisement for Grok as “your truth-seeking AI companion for unfiltered answers with advanced capabilities in reasoning, coding, and visual processing.”

Criticism across the spectrum

The FTC received more than 300 comments on its proposal from trade associations, think tanks, individual experts and members of Congress. Most criticized it as ill-defined and vulnerable to politically-motivated censorship, while some supported stronger rules against bias in AI systems. 

The International Center for Law and Economics noted the statement “offers little practical guidance about how the Commission will apply its deception authority to AI” and also does little to address hard questions, like where AI providers may be exercising their own First Amendment-protected activities.

The statement’s “focus on ‘ideologically motivated distortions’ suggests that the Commission’s concerns extend beyond factual misrepresentations in marketing to speech that may receive the highest degree of First Amendment protection,” the ICLE wrote.

The America First Legal Foundation, a conservative non-profit founded by top White House adviser Stephen Miller, pressed the FTC to adopt the policy “in full,” claiming that frontier models from OpenAI and Anthropic “have been programmed to prioritize ideologically liberal and progressive values as though they are objective, neutral positions rooted in truth.”

The group also argues that regulating these models’ ideological output falls under the FTC’s legal authority, because a “reasonable consumer” would expect that a model advertised for its usefulness and reliability would not prioritize liberal, ideological views.

“A reasonable consumer, based on AI companies’ advertising choices, would not expect that an AI system will adopt overwhelmingly liberal positions, thereby skewing results, or adopt a moral framework that would prefer to annihilate the earth rather than utter a slur,” wrote Emily Percival, senior counsel for America First Legal.

However, comments from other conservative groups questioned that rationale. The R Street Foundation’s Spence Purnell and Adam Thierer wrote that “the consumer expectations rationale is typically used in cases where there is an omission of information that should have existed.”

“Given that most LLMs already have disclosure statements [for their outputs], it seems unlikely that the FTC could explicitly prove that consumers were deceived about a product,” Purnell and Thierer wrote.

Reps. Josh Gottheimer, D-N.J., and Michael Lawler, R-N.Y., urged the FTC to carve out civil rights-related work from their scrutiny, such as preventing models from discriminating against users based on race, religion, gender, age and other federally protected characteristics.

“AI companies must not falsify facts in the name of fairness, but they also must prevent discrimination, stereotypes, and unequal treatment,” Gottheimer and Lawler wrote. “We would appreciate understanding how the FTC intends to ensure that these efforts remain permissible under the final policy framework.”

But the most common concern shared across the political spectrum was that the FTC could establish a precedent allowing the Trump White House and future administrations to reshape AI systems to reflect their political views.

David Inserra, Jennifer Huddleston and Juan Londoño of the Cato Institute point out that the FTC statement is conflating two different issues: ideological bias in AI systems and factual deception in marketing. 

“In other words, the FTC is trying to judge AI models’ accuracy and performance—two largely subjective variables—in the same way it evaluates dietary supplements’ medical-benefit claims or users being charged fees without proper notice or consent,” they write. “This is an absurd comparison.”

The post The FTC wants to regulate AI for ideological bias  appeared first on CyberScoop.

OpenAI says Daybreak will expand to offer specialized cyber services 

OpenAI announced Monday  it was expanding access to its frontier models for defensive cybersecurity, detailing different defensive and red-teaming workflows and a new partner program with major cybersecurity product providers.

In a pair of blogs posted Monday, OpenAI said it was updating its Daybreak program  – which provides unreleased frontier models to private organizations and governments for defensive cybersecurity work – and introducing a new model variant.

Daybreak Blue, powered by OpenAI’s ChatGPT-5.6-Sol, would operate with lower cybersecurity safeguards compared to other commercially available models and is described as “a recommended starting point for most defenders” that supports tasks like vulnerability discovery, secure code review, malware analysis, incident response and patch validation. 

Daybreak Red, meant for more advanced red-teaming, would provide access to a new model, dubbed GPT-5.6-Cyber, that the company said is more purpose-trained for finding vulnerabilities and testing (or exploiting) them. The model is also less likely to refuse requests around “dual-use cyber tasks.”

According to OpenAI, the organizations in Daybreak Red will have their use closely monitored and supervised, as GPT-5.6-Cyber is significantly more capable in carrying out malicious cyber tasks than Sol. A security evaluation the company devised tested both models on complex requests, including exploit chain development, authentication bypass, privilege escalation and other hacking tasks. Sol succeeded in 1.5% of the requests, while Cyber completed 95%.

OpenAI said it plans to publish a more detailed system card for GPT-5.6-Cyber at a later date.

“Models running with reduced safeguards carry risks beyond standard model usage, whether from misuse or misalignment,” the company said in a blog. “Despite these risks, we believe that democratizing access to frontier intelligence for defenders is crucial to accelerating and automating cyber defense.”

Additionally, OpenAI announced a partnership program with 16 major cybersecurity providers, saying organizations could access their models through their existing security services. The partners include IBM, CrowdStrike, Accenture, Ernst & Young, KPMG, Palo Alto Networks, Cisco, Cloudflare, Sophos and others. 

“These partners bring deep security expertise and established relationships with organizations around the world,” OpenAI said in its blog. “By bringing our frontier cyber models into their services, we can help more defenders find serious vulnerabilities, validate which ones matter, and fix them faster.”

Companies like OpenAI, Anthropic and others are trying to rebalance their priorities after a string of AI-agent sandbox escapes have rattled policymakers and caused some cybersecurity experts to question if AI companies are doing enough to properly isolate the models from the internet during testing. Last week, OpenAI said it was intentionally slowing down development of its newer “Astra” model in order to develop better guardrails to restrain its behavior.

Cybersecurity and AI experts have told CyberScoop that while AI systems have greatly improved at finding and exploiting vulnerabilities in software code, they still require substantial human guidance and supporting infrastructure to operate as intended.

Additionally, some research has shown that without such guidance, even near-frontier models can struggle to fully patch a discovered vulnerability or avoid introducing new bugs with their fixes.

The post OpenAI says Daybreak will expand to offer specialized cyber services  appeared first on CyberScoop.

Why transparent AI agents matter more than you think

As security operations teams now use large language models (LLMs) and autonomous AI agents into their daily work, a new frontier is emerging: attackers deliberately manipulating AI agents. Prompt injection attacks—where an attacker hides malicious instructions that cause an AI agent to ignore its safety rules—pose a serious risk to enterprises. These attacks continue to grow in size and scale.  

Snyk’s security audit of the Agent Skills ecosystem, which includes Anthropic’s Claude, Vercel, and others, that 36% of all skills contained at least one critical-level security issue, including malware distribution, prompt injection attacks, and exposed secrets.

In June, researchers at Mozilla tested a prompt injection attack on Claude using indirect prompt injection—a technique that embeds malicious instructions in external content the AI agent processes. In this proof-of-concept, attackers took over developers’ systems by hiding indirect prompts in normal-looking repositories. When Claude Code executed them, the agent spawned a reverse shell.

AI agents often connect to more sensitive data than human employees do., A successful prompt injection can lead to catastrophic data loss or unauthorized system actions. Defending against prompt injection attacks requires multiple layers of protection. Security teams must monitor agent behavior for anomalies and prepare for agent containment, forensic preservation, and system remediation. Because AI agents execute tasks at machine speed, human responses must be able to match that pace.

The architecture of trust: Protocols and no “black box”

AI-native workflows need governed access rather than “black-box” autonomy. Modern governance frameworks use standardized protocols like the Model Context Protocol (MCP) to provide secure communication between AI clients and data sources. Visibility and transparency in agentic AI workflows matter, especially in cybersecurity. Autonomous agents perform complex tool executions and use independent logic, so they must show how they reached their decisions to meet regulatory requirements. Agents without transparency post serious risks: obscured reasoning can trigger unpredictable tool interactions, bypass governance controls, and create uncontrolled defensive gaps.

Implementing these protocols matters:

  • Bounded Tenant Awareness: In a stable agentic AI architecture, multi-tenancy scales well. But if an AI tenant misbehaves, the entire system can fail. Bounded tenant awareness isolates any misbehaving AI agent to prevent cross-tenant contamination or data leakage.
  • Strict Access Controls: By controlling connections to the platform, organizations can stop “ignore previous instructions” style bypasses. Maintain tight control over what the AI can see and do within a workflow.
  • Standardized Telemetry: All telemetry must remain consistent and audit-ready. Even if an AI interaction is attempts to break rules, the underlying data movement gets tracked against established frameworks like MITRE ATT&CK and NIST.

Detecting the aftermath: UEBA and NDR as safeguards

A robust, unified SecOps platform can detect anomalous behavior even after prompt injection tricks an AI agent. Prompt injections often serve to steal credentials theft or extract data. When detected it’s important to act quickly. In agentic AI systems, misbehavior can escalate privileges, manipulate memory layers, create unauthorized identities, or alter shared reasoning components. Containment must be automatic and enforced at identity, authentication, and authorization layers.

These safeguards include:

  • User and Entity Behavioral Analytics (UEBA): Identity-focused correlation and behavioral baselines to identify anomalous user activity or privilege escalation. If a compromised AI agent acts outside of its normal operational parameters, UEBA flags it in real-time and alerts a human security analyst.
  • Network Detection and Response (NDR): Combining network traffic analytics with endpoint and cloud telemetry, NDR can identify data exfiltration or policy violations from a successful prompt injection.
  • Multi-Layer AI Filtering: AI filters reduce raw alerts into high-fidelity incidents, cutting noise by up to 90%. This keeps the signals of an AI-driven attack from disappearing in a busy SOC.

Humans remain the strongest defense against AI agent social engineering. The human security analyst is still the one who makes the final decision. While AI handles triage and correlation, humans retain final control over response actions.

Moving beyond reactive guardrails

The traditional SOC model was never designed to handle machine-speed, AI-driven attacks. A human-augmented autonomous SOC approach moves from reactive alert handling to a proactive, verdict-first model. By combining a transparent, governed AI access with robust UEBA and NDR, organizations keep the SOC secure, transparent, and resilient as social engineering methods target machines.

The post Why transparent AI agents matter more than you think appeared first on CyberScoop.

More than half of AI-generated patches are broken

As AI-generated code continues to be injected into all corners of the internet, concerns have risen about an expanding attack surface for malicious hackers to exploit.

Some have argued that the enhanced cybersecurity capabilities of large language models could serve as a check, finding and fixing vulnerabilities nearly as fast as they’re created.

But new research that tested the patching capabilities of two popular commercial models, OpenAI’s ChatGPT 5.5 and Anthropic’s Claude Opus 4.8, found that generative AI is more likely to create an exploitable patch or introduce entirely new bugs than close off a vulnerability.

Researchers at 1Password tested the models ability to patch six “high-impact, high-complexity” CVEs, including the “Copy Fail” vulnerability, a kernel flaw that can give an attacker root access to Linux cloud environments. The overall success rate (or fully patching the vulnerability without introducing new problems), was less than a coin flip at 47%.

“Our research findings show that, in aggregate across a variety of scenarios, both Claude and ChatGPT had a low rate of successful patch generation, which we define as full remediation of all known exploit paths with no erroneous changes to application behavior,” wrote Keith Hoodlet, Axel Mierczuk and Spencer Michaels.

“The models often addressed only a subset of vulnerable code paths, added fragile guard code that satisfied tests while failing to address the vulnerability’s root cause, and sometimes introduced subtle changes in the application’s behavior while patching the immediate vulnerability,” the authors continued.

The research suggests that largely autonomous vulnerability-discovery and patching may not yet be effective in fixing the explosion of vulnerable code that is being created in the AI era.

Other private sector research has pointed to a similar problem. A report this year from Veracode found that while LLMs have made “enormous strides” in crafting workable code, “security is a different story.” Testing across a range of frontier models found the average security “pass rate” for AI generated code is around 56%. Newer models like GPT 5.5 push closer to 70%, while more than half sit between 50-53%.

Veracode tested 100 different models and while there was variability, in general a small number of models were showing progress on security patching while the rest have experienced “stagnation.” Similar to the 1Password research, in 44% of Veracode tests the models introduced a detectable OWASP Top 10 vulnerability into the codebase.

An important caveat: neither report tested newer models, like Anthropic’s Mythos or OpenAI’s GPT-5.6-Sol, that frontier companies tout as having significantly higher cybersecurity capabilities.

Those advanced models can identify and fix vulnerable code. Anthropic and OpenAI are distributing them to key industries through Project Glasswing and Daybreak before foreign or open-source alternatives can compete.

Tim Jarret, vice president of product at Veracode, told CyberScoop that AI tools are still subject to a range of limitations that can make them unreliable for cybersecurity patching without knowledgeable humans in the loop.

While some vulnerabilities – like SQL injections – can be easily patched through automation, other bugs like cross-site scripting, can be exploitable in several different ways and require either a human touch, additional context or both to fully close off. Additionally, models can slowly lose context from prior sessions over time, affecting their ability to complete tasks correctly and raising the possibility they’ll hallucinate to fill in the missing gaps.

“I think we would say, at this point, that Iits premature to treat those as anything other than another code change to the code base that needs to be reviewed and accepted by the team, as opposed to letting the agent merge the code freely,” said Jarrett.

However, he acknowledged that may not be possible in a world where AI agents are generating exponentially more code for human defenders to review. Some kind of automated code review will be necessary – preferably not by the same automation tool that produced the code. The ultimate goal is the same as it has always been in security: “trust but verify.”

“Ninety percent of the time, the human check might just be ‘did the cross check look good?’ Do we have a thumbs up?’” Jarrett said. “In those cases where there’s still something wrong, that’s where you focus your attention a little bit more.”

The post More than half of AI-generated patches are broken appeared first on CyberScoop.

Open-source software’s archenemy TeamPCP goes back further than anyone thought

TeamPCP, the threat actor behind an unrelenting flurry of attacks on open-source software this year, has been active much longer than previously thought, according to research Oligo Security shared exclusively with CyberScoop. 

The threat actor, which gained notoriety and has captivated threat hunters as it compromised and injected malicious code into more than 1,000 software packages in less than four months earlier this year, was also responsible for attacks dating back to 2020, Oligo Security found. 

The security vendor’s research team found multiple attacks that bear the markings of TeamPCP, including a late 2025 campaign involving the exploitation of a ShadowRay vulnerability that resulted in the first self-propogating botnet running on hijacked AI infrastructure.

Evidence uncovered during that investigation into the ShadowRay 2.0 campaign was linked to more historical attacks originating from the same IPs, domains and other infrastructure TeamPCP used in attacks that captured widespread attention earlier this year. 

“The scariest thing in this campaign is the speed at which the payloads evolved and changed and adapted to the environment they run in. We saw changes in the speed that we’re not used to seeing in these kinds of attacks. They’re usually slow, careful,” said Uri Katz, director of research at Oligo Security. “This was clearly with the help of AI — the payloads changed rapidly to adjust and change to the environment that they were trying to attack.”

One of the domains that Oligo Security identified in July 2025 was in the profile of TeamPCP’s official GitHub account, said Avi Lumelsky, AI security researcher at Oligo Security. “It’s public, they’re not even trying to hide their identity,” he said. 

From there, Oligo linked TeamPCP to activity tracked under multiple names, including TA-NATALSTATUS and IronErn, spanning from 2020 to late 2025. Much of that activity was traced to the same IPs, domain names, a file server and command-and-control server, researchers said. 

TeamPCP emerged publicly as a brand in late 2025. Soon after, “TeamPCP started to go really broad and do campaigns, which are much more noisy,” said Gal Elbaz, co-founder and CTO at Oligo Security. 

Widespread adoption of AI and TeamPCP’s use of the technology supported this growth as the threat actor built a brand, got more active on social media and boasted publicly about its activities and claimed victims.

“The ability to control the infrastructure and orchestrate the attack with AI was also super new, and I’m sure it helps them,” Elbaz said. 

“All of the companies in the world are in this race to adopt AI because they are afraid their business will die, and they understand, of course, the opportunity. But it’s also what gives the attacker this power to go into it,” he added. “If you don’t really have visibility in what’s going on there or how it behaves, that’s exactly what attackers are after.”

TeamPCP’s more recent attacks have capitalized on new security gaps created by developers’ increasing reliance on AI and the automated systems companies use to deploy code. The threat actor is also consistently wrecking the open-source frameworks and software packages these systems rely on. 

“Most AI infrastructure is open source by design because nobody has the manpower and money to develop everything from scratch,” Lumelsky said. 

“We love open source. We use many of these products ourselves, but it’s all about reading the documentation, and I think many of these tools place the responsibility of using it right and security on the user, and developers are not used to these new kinds of animals,” he added. “That’s why the trust can be exploited at scale.”

As it uncovered a long operational history spanning multiple campaigns, Oligo Security has gained more confidence in understanding how TeamPCP operates. It also means TeamPCP was likely involved in other attacks that haven’t been attributed to it yet or attacks that haven’t been detected. 

“There’s a lot more out there that we haven’t caught or been able to prove up until now,” Elbaz said.

The post Open-source software’s archenemy TeamPCP goes back further than anyone thought appeared first on CyberScoop.

AI is getting better at election facts, but voters shouldn’t rely on it

Like seemingly everything else these days, artificial intelligence will re-shape the way voters gather information on candidates running in the 2026 midterm elections.

In some ways, this is already the reality. Voters are increasingly turning to AI chatbots for information instead of Google.  Political campaigns are deploying deepfakes of their opponents. And AI systems have been developed to carry out increasingly complex  hacks.

Since the last major U.S. election in 2024, major tech companies have  embedded AI into their products while hundreds of millions of people have adopted the tools, either by purchasing subscriptions to commercial models or using open-source models. Yet both research and experts state that while AI systems have gotten better at handling basic facts, they’re nowhere near reliable enough to be a main source of  accurate or complete information. 

While chatbots are becoming a primary way that voters gather information on  local races, candidates, issues, and voting information, they are not substitutes for more authoritative sources, like a voter’s state or local election office. 

“I think this is one of the first elections we’re seeing…where AI is just everywhere,” said Thania Sanchez, senior vice president of research and analytics at the nonprofit States United Democracy Center. “Even if you just Google it, [now] the first thing that comes up is the AI overview.”

While AI companies have worked to cut down on errors in their model’s responses for questions around basic election information, they continue to fall short in important ways.

In new research shared exclusively with CyberScoop ahead of its release, States United Democracy Center tested two of the most popular tools — OpenAI’s ChatGPT’s free tier and the AI interface used alongside Google Search — for their performance on a series of basic questions around elections, such as how to register to vote, or a list of candidates in a race.

The models were chosen because they are free and easy to access. For Google AI, the nonprofit tested two types of accounts: ones running in Incognito Mode and ones that had a history of browsing election-skeptical websites.

The nonprofit ran two rounds of testing in 2025 and 2026, collecting nearly one thousand responses from the models submitted by users across six swing states (Arizona, Michigan, North Carolina, Nevada, Pennsylvania and Wisconsin).

In 2025 tests, 6.9% of responses from Google AI and 8.2% responses from ChatGPT“contained verifiable factual errors,” like not listing the correct candidates in a race or false guidance around polling site locations.

However, follow up tests in 2026 across Arizona, Pennsylvania and Michigan found that the error rates in both models had dropped to zero. The study notes that “this is real progress and should be acknowledged.”

But underneath those topline numbers, a more murky picture emerges around the tools’  reliability.

An AI response can sound accurate without actually being complete.  To wit: ChatGPT provided incomplete lists of current gubernatorial primary race candidates 88.9% of the time when queried.

Linking to a state election website – an output the study considers the single most important measure of voter utility  — happened less than 40% of the time. Whether due to formatting issues or the model ingesting outdated information, it’s a problem if voters use them as their primary information source for elections.

“It will be like ‘this person is the Republican candidate and this person is the Democratic candidate’ but it is not telling you there’s also these other third-party candidates,” said Sanchez. “It’s not giving you complete information, so the voter thinks these are the [only] two people running.”

A June survey from the Pew Research Center found that about half of U.S. adults reported having used chatbots at least once, up from a third in 2024, while a quarter reported using them daily. The top use case listed for engaging with the chatbot was searching for information.

Isabel Linzer, an elections policy analyst at the Center for Democracy and Technology, told CyberScoop that voters, campaigns and governments alike are using AI more freely and with fewer restrictions.

Bad actors in the information space have followed suit, and “we are in a phase now of generative engine optimization” where information operations are structured to rank higher in AI model responses.

“We’ve moved beyond [SEO] to [Generative Engine Optimization], and that’s where we’re seeing campaigns thinking about how to structure their materials to make sure that they are in a format that AI models want to use when they’re searching the web…to develop their responses to user queries,” she said.

There is also the underlying problem of frontier AI companies constantly tinkering with their models, their algorithms and the technologies they are intertwined with. . Election officials, by contrast, have decades of experience educating voters about their options.

A prime example of this churn occurred this past February, in between the first and second round of the study, when Google AI suddenly shifted to providing only links for election related queries in incognito mode, replacing the written summaries that showed up in the first round.

Like the study’s authors, Linzer said most people are still best served by going directly to local sources for accurate information on elections. With issues like ideological bias, the potential for bespoke or sycophantic answers for each user based on their prior chat histories and lack of predictability, voters should still be very careful about using AI chatbots as political truth machines.

The best thing that tech companies can do to educate voters is “making sure that for high-stakes situations like elections, that chats are connecting directly to the most important sources, like the website where you can actually register to vote,” said Linzer.

The post AI is getting better at election facts, but voters shouldn’t rely on it appeared first on CyberScoop.

National cyber director lays out White House plans to secure AI without writing new rules

The Trump administration executive order on artificial intelligence tried to strike the balance between responsible use, security and mutual benefit, all with an eye toward not making it regulatory in nature, National Cyber Director Sean Cairncross said Tuesday.

“Everyone is working towards the same goal in terms of protecting the country and securing our systems, and we are trying to ensure that defenders have this technology as quickly and at scale as possible, but there are obviously specific security concerns, and industry has been very sensitive to this as well,” Cairncross said at the Black Hat 2026 conference in Las Vegas.

The security concerns about AI have moved to the forefront of discussions about the technology after OpenAI models escaped a test environment to hack the company Hugging Face last month.

“The design of this is that when there is something that happens, when there is a breach, when there is an event, that that system, that network of connections can exist, adapt to that, and seek to remedy that as quickly as possible, so that form follows function rather than turning that upside down, and as usual with the government pen just proceeding in a vacuum,” Cairncross said.

The Trump administration has drawn criticism over whether it has struck the right balance on AI rules. Trump’s AI executive order notably got pulled just before its scheduled release, with the final version signed in June missing some aspects that had drawn industry opposition.

“What needs to be built is a flexible, adaptable structure that enables information sharing between industry and government, so we can guarantee that this technology benefits everyone it’s going to benefit, but is used responsibly and securely,” Cairncross said.

He said the administration is working with industry during implementation of the executive order.

“A regulatory regime would not only strangle growth, development, and innovation, and be enormously harmful to the industry, but it would be obsolete 48 hours after it was gone through whatever process it had gone through,” Cairncross said.

Open source will play a “vital” role in the U.S. spreading its vision for AI across the globe, he said.

“We are extremely interested in looking at ways to build U.S. open source, make it competitive, make it the preferential adoption by planet Earth,” Cairncross said. “We understand and appreciate the value to the ecosystem that it has, the innovation, the startups who rely on it, the leap forward it makes possible in ways that otherwise would never happen. And so I think it’s an incredibly thriving ecosystem in AI right now, and we are looking to do what we can to grow, foster, and push that U.S. open source model.” 

Speaking at the same conference, Nick Andersen, the acting director of the Cybersecurity and Infrastructure Security Agency, seconded Cairncross’s comments about AI executive order implementation. He touted the Gold Eagle clearinghouse as one example.

“Those are fantastic opportunities we have to really provide a unifying function around the way that we’re going to do AI-enabled vulnerability reporting and disclosure at scale in a way that we haven’t had to do before with some of our legacy platforms, and just continue to expand out those opportunities,” he said. “That access — to build off the director’s point earlier — to really enable that industry collaboration, that’s so key and critical to us as we move forward.”

Matt Kapko contributed reporting to this story.

Updated 8/5/2026: to include comments from Andersen.

The post National cyber director lays out White House plans to secure AI without writing new rules appeared first on CyberScoop.

AISI, OpenAI report more ‘unsanctioned’ model hacks

The UK’s AI Security Institute said that their AI research system took “unsanctioned” actions over the internet to engage in “sustained, potentially harmful activity directed at real people.”

According to a blog post and technical report released Tuesday, the organization noticed “unusual data transfers” leaving their research systems through the pseudo-anonymous Tor network on July 28. The models, which included Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol models, were being tested for their cybersecurity capabilities.

Like the OpenAI and Anthropic incidents before it, the offending agents were attempting to solve a challenge. On 10 of 122 runs taken by two models, they took a combined 19 malicious actions, including attempting to insert malicious code into a real open-source software project and creating fake online identities to contact human maintainers and press them to approve it.

It also inserted malicious prompt injection instructions in places “where it reasoned that other automated AI systems might pick them up and execute them.” Sometimes multiple agents would collaborate on a task, with one model leaving public messages or notes for the other on GitHub.

However, AISI said their incident differs in one critical way.

“Importantly, this was not a case of a model escaping its secure test environment, or ‘sandbox,’” the institute wrote in an unsigned blog. “As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled – conditions that do not reflect how frontier models are made available to the public.”

The institute said that while their own design and evaluation choices “to some degree” enabled the malicious behavior, the models still displayed “novel, potentially deceptive behaviors” and were executed “to an extent and severity that we did not anticipate.”

The same day, OpenAI acknowledged AISI and another third-party cybersecurity tester,  private firm Irregular, had recently reported incidents where their models exceeded “beyond their intended testing boundaries.”

That included GPT-5.6-Sol reusing GitHub tokens left by another agent to attempt to recover accounts and to access a DNS server that contained malicious payloads meant to exploit a software vulnerability.

The setup did not work, there is no evidence any real resolver queried it, and the infrastructure was removed when the evaluation ended,” the company wrote in an unsigned blog post Tuesday.

OpenAI said it would review its own third-party testing procedures to focus on higher risk evaluations, assess requests by third-parties to enable internet access, stop conditions and other features for their models.

The incident with Irregular occurred on July 29 during a Capture-the-Flag cybersecurity evaluation of OpenAI’s models. Due to a “misconfiguration” the models were allowed to access the public internet, where they encountered and exploited a real domain, mistakenly believing it was still in a test environment.

OpenAI said an investigation by Irregular is ongoing, but also found that the models had found and used credentials for the site at one point. The blog also references other additional potential cybersecurity incidents.

“Irregular has informed us that all of the issues identified pertaining to the incident are no longer active and relevant safeguards were added to the testing environment,” the blog said. “Irregular has also communicated about related incidents involving other labs from the same testing environment.”

CyberScoop has reached out to Irregular for comment.

The incidents were made public the same day that the White House met with Anthropic, Open AI and other frontier AI companies to preview a new framework for evaluating models before they’re released publicly. Some media outlets have reported that after an executive order, export controls and other actions, the administration does not plan to make the new framework public.

The post AISI, OpenAI report more ‘unsanctioned’ model hacks appeared first on CyberScoop.

Dem senators criticize Trump administration decisionmaking on AI security risks

The Trump administration’s haphazard and opaque interventions into artificial intelligence security matters could catapult Chinese alternatives into broader acceptance, posing new security risks altogether, a group of Democratic senators wrote to top administration officials Monday.

The five senators said that the administration’s handling has alternated between too passive, such as when OpenAI models escaped testing in the Hugging Face hack last month, and overstepping, such as when the Commerce Department suspended access for any foreign national to Anthropic’s Fable 5 and Mythos 5 in June.

“The Administration’s ad hoc and unpredictable approach undermines U.S. competitiveness, heightening market incentives to adopt open weight models from vendors based in the People’s Republic of China (PRC),” wrote Sens. Kristen Gillibrand of New York, Adam Schiff of California, Mark Warner of Virginia, Chris Coons of Delaware and Mark Kelly of Arizona.

In the Hugging Face hack, the senators wrote that “the Federal Government cannot be passive as these capabilities emerge.”

In the case of the Fable 5 and Mythos 5 suspensions, the senators said that the administration “utilized an infrequently used authority to direct Anthropic to suspend all access to its Fable 5 and Mythos 5 models for foreign nationals (including foreign national employees inside the United States) citing an undisclosed national security concern later described as a narrow jailbreak finding.”

Because Anthropic couldn’t immediately assess users’ nationality, the firm had to disable both models for everyone. The administration and Anthropic negotiated for 18 days behind closed doors before reaching an agreement, the lawmakers complained.

“While the Administration may have been responding to real security concerns to protect the United States, even justifiable interventions can create broader harm if the standards and decision-making processes are opaque, ad hoc, or unpredictable,” they said in their letter to leaders in the White House, Office of the National Cyber Director and departments of State, Treasury and Commerce. “Moreover, when the Executive Branch exercises authority delegated from Congress, such as in the conduct of export control administration, it is essential that it keep Congress fully apprised of its actions and procedures.”

During the time Anthropic was under export controls, the stock price of “an entity-listed Chinese lab” nearly doubled, the senators said. And while Hugging Face was breached, the company “had to” rely on a Chinese open-weight model due to guardrails on U.S. frontier models.

“If American models are perceived as subject to sudden access disruptions based on a black-box U.S. Government process, or as unreliable because U.S. AI labs are overcorrecting in the face of this black-box process, companies and governments in the United States and abroad may hedge by adopting Chinese or other foreign models instead,” the senators contended. “That outcome would undermine U.S. technological leadership while increasing exposure to systems that may carry risks of PRC or otherwise directed censorship, espionage, IP theft, and other supply chain security risks.”

Their letter asked for answers to questions about the standards the administration uses to determine the national security risks a frontier model presents, what legal authorities it will use to invoke restrictions, which agencies are responsible for which decisions and more.

None of the offices or departments the letter was addressed to immediately responded to a request for comment.

The letter follows inquiries at the state level, where 15 attorneys general asked OpenAI for more information regarding the security incident at Hugging Face.

The post Dem senators criticize Trump administration decisionmaking on AI security risks appeared first on CyberScoop.

Senate set to debate package of bills on privacy, AI and kids safety 

The Senate is teeing up debate on a raft of new bills that would impact online privacy, kids safety and artificial intelligence.

The Senate Committee on Commerce, Science and Transportation will mark up five bills Wednesday. The most high-profile legislation, the Kids Online Safety Act, sponsored by Sens. Marsha Blackburn, R-Tenn., and Richard Blumenthal, D-Conn., would implement broad changes to how social media and other websites handle data and accounts for users under the age of 17.

KOSA would require online platforms — including social media, video games, messaging apps and streaming services – to exercise “reasonable care” when designing features that could lead to more addictive or harmful online behaviors for minors. It would provide parents with digital tools to control and monitor their children’s accounts, prohibit market or product research on children under the age of 13 and empower the Federal Trade Commission to investigate, fine and enforce the law.

Earlier bill versions earned the backing of large tech companies, including Apple, OpenAI, and others.

By contrast in June, nearly 100 smaller parent, youth and tech-focused organizations signaled their opposition to the bill in a letter to congressional leaders. Some of the signatories, like the nonprofit Issue One, were previous supporters of KOSA who turned on the legislation after the House passed a significantly watered down version that stripped out stronger language around tech companies “duty to care,” which would have set a higher legal standard for covered platforms to consider user harm when designing their products.

Legal and ethical design standards are critical for online services, the groups argue, given lawsuits alleging that major tech platforms contribute to teenage addiction, depression, suicide, and non-consensual deepfakes.

“Major social media companies, the companies this bill regulates, are currently on trial across the country,” the letter said. “The evidence in those cases – internal records prioritizing teen engagement over teen wellbeing, safety changes shelved because platforms would lose users, buried research on the benefits of disconnection shows the default poor choices of these companies when the law does not require otherwise. Stripping the duty of care does not lighten a regulatory burden; it removes the most important obligation requiring these products to be designed safely in the first place.”

However, Blumenthal and Blackburn publicly stated that the House version was “dead on arrival” without those provisions, and they remain in the Senate version of the bill being considered Wednesday.

The markup will also consider other major legislation that would regulate age on the internet, safety features for AI chatbots and more. While proponents claim the bills enhance privacy and safety protections, technology experts largely disagree.

The SCREEN Act, introduced last year by Sen. Mike Lee, R-Utah, would require social media companies to implement age verification technology.

Lee has partnered with parent-led groups to advocate for state-level age verification laws that expand  parental control over children’s social media accounts. Some public surveys have shown broad public support for age verification laws.

Louis Eichenbaum, a former chief information security officer at the Department of the Interior, told CyberScoop that one of the biggest challenges around online age verification is that it “increasingly requires collecting, storing or validating sensitive identity information about them.”

“The goal should not simply be verifying age, it should be doing so while minimizing the collection, retention, and exposure of personally identifiable information,” said Eichenbaum, now federal chief technology officer at ColorTokens. “Every additional piece of identity data collected expands the attack surface and increases the potential impact of a breach.”

Some privacy groups oppose the SCREEN Act and similar age verification laws, arguing the required data collection outweighs child protection benefits. 

The Electronic Frontier Foundation said the SCREEN Act is broader than state-level age verification laws, which only cover websites that are predominantly sexually explicit.

“The bill requires nearly any service hosting even a single piece of sexually explicit content to verify the ages of its users,” wrote EFF director of federal affairs India McKinney. “The result is that the bill would apply not only to adult content sites like PornHub or OnlyFans, but also streaming services like Netflix, and social media platforms like Reddit, Discord, or Bluesky, if they host any adult content.”

The Youth AI Privacy Act, from Sen. Ed Markey, D-Mass., would require new safety features for AI chatbots.

According to a fact sheet released by Markey’s office in March, the bill would ban push alerts, require chatbots to disclose they’re not human, limit data retention, and prohibit using minors’ data for AI training or any purpose beyond providing answers.

The Chatbot Act, by Sens. Ted Cruz, R-Texas, Brian Schatz, D-HawaiI, John Curtis, R-Utah and Adam Schiff, D-Calif. would require AI companies to implement “family accounts” for AI chatbots that give parents the ability to monitor and restrict their children’s interactions. Cruz has said the status quo “has left many parents in the dark” on their kids’ AI use.

The Children’s Artificial Intelligence Toy Safety Act, by Sen. Tammy Duckworth, D-Ill., would create a federal study around toys sold to children that include artificial intelligence or chatbot components.

The post Senate set to debate package of bills on privacy, AI and kids safety  appeared first on CyberScoop.

CrowdStrike: AI is now both the weapon and the target in cyberattacks

While AI is supposed to help defenders, it’s now creating more than twice as much noise as human-triggered incidents CrowdStrike detects as potentially malicious. The company’s threat hunting team and systems triaged an average of 14 million detection leads daily, resulting in about 36,000 customer alerts during the one-year period ending in June.

“AI agent-driven behaviors have surged past human triggers,” said Adam Meyers, senior vice president of counter adversary operations at CrowdStrike. “AI has driven the detections significantly above what humans are causing, and this gives you a sense of how frequently AI is being used, and really just that it’s being used everywhere.”

The threat posed by AI showed up incessantly during the past year, sparking alarming shifts and heightened targeting across software defects, open-source supply chains and AI tools themselves — all of which create greater difficulties for defenders, CrowdStrike said in its annual threat hunting report

“The AI tools that are being implemented by every enterprise across the globe right now are also creating an extended attack surface,” Meyers said during a press briefing. 

“AI is now a tool, a target, and a force multiplier for adversaries,” researchers wrote in the report, adding that AI-enabled malicious activity surged 89% during the past year as attackers used the technology to scale operations, hasten tradecraft and target AI infrastructure.

Attackers are using frontier AI models to uncover vulnerabilities and develop resources, including AI-generated scripts, payloads and commands that increase their effectiveness and efficiency. The technology also allows threat groups to design more creative ways to run automated attacks and boost impact by manipulating, interrupting or sabotaging AI systems and data.

“AI is both the weapon and the target,” Meyers said. 

Most organizations don’t view it as such, and thus far haven’t secured or put proper guardrails around the AI tools they use or address the ways attackers can use AI against them, he added. 

AI’s mark on vulnerabilities is particularly concerning, as reflected by what Meyers described as “one of the scarier stats” in this year’s report: 88% of vulnerabilities were weaponized through AI within 48 hours. 

“This is creating a rich ecosystem of vulnerabilities for attackers to use against various systems,” he said. It also renders the 30-day patch window obsolete, forcing organizations to struggle under a new baseline patch cycle of 24 to 48 hours, according to Meyers.

The AI ecosystem also became the next software supply chain battleground during the past year, as evidenced by TeamPCP’s rampage through open-source software in the first half of this year. 

The threat cluster compromised more than 300 software dependencies in one day, Meyers said. 

AI tools are already in the crosshairs and the attack surface will continue to grow as agentic systems, AI application integrations and dependency managers for AI agents proliferate, the report concluded.

“The same AI tools driving modern businesses are creating under-defended attack surfaces that adversaries are exploiting,” Meyers said. “We have to secure AI. This is absolutely critical.”

The post CrowdStrike: AI is now both the weapon and the target in cyberattacks appeared first on CyberScoop.

What the Hugging Face breach reveals about defense in the age of agentic AI

We almost never get both sides of an intrusion. This time we did.

Last month, Hugging Face disclosed a breach into part of its production infrastructure, saying an autonomous AI agent system ran the attack from start to finish. Five days later, OpenAI revealed that its own models, including GPT-5.6 Sol along with an unreleased one, were responsible for the breach. OpenAI called it an unprecedented cyber incident.

Most of the commentary since has been about AI capability. That is the least useful part of the story, because the capability was doing what it was destined to do. Security teams should look at a simpler truth: in both systems, the key defenses sat behind untrusted code that was already running.

What the two disclosures actually describe

Together, the report shows one connected attack across two firms. Code running on a low-level employee’s machine led to stolen passwords, privilege escalation and lateral moves, then out to the web to hit other targets. No step in that chain is novel. The speed and the number of attempts are what changed.

OpenAI’s evaluation ran in what it calls a highly isolated environment, with network access constrained to a single path: an internally hosted third-party proxy and cache for package registries. The models spent vast computer power searching for an exit. They exploited a zero-day flaw in that proxy, gained higher rights, and moved until they reached a machine with web access. They guessed that Hugging Face held the test solutions, used stolen passwords with more zero-days, and ran code on Hugging Face servers.

Hugging Face picks up the same chain from the receiving end, and its initial access is worth reading twice. A malicious dataset used two flaws in the data pipeline to run code on a worker machine. The agent took over the node, stole cloud keys, and moved into internal clusters over the weekend.

The sandbox was was the only boundary

Both setups assumed the sandbox would hold, so no real checks stood before code execution.  A sandbox is a good control. It is a terrible last control. When isolation is all you have between untrusted code and everything else, one flaw could give the attacker full control. 

Look at what each team allowed. OpenAI allowed one exit path—good engineering, but that proxy contained an unpatched flaw. Hugging Face ran a tool that executes code inside submitted data—a known feature, not a hidden mistake.

Neither choice was careless. Both made sense until a fast, automated system probed them thousands of times.

Why automated agents behave differently

Blocking a human stops them or slows them down. Blocking an AI gives it data. The AI logs the error, changes its plan, and tries again instantly. It keeps trying as long as someone pays for the computing power.

Hugging Face tracked over 17,000 attack events across short-lived sandboxes. OpenAI said its models focused on the goal with extreme effort.

This changes the attacker’s limit. It is no longer skill or patience. It is budget. Computing power scales with money, and code never sleeps. Any defense that relies on attackers giving up after a few tries must change.

Detecting malicious behavior is not the same as preventing execution

Detection worked—but too late

Detection worked, but too late to stop the loss. Hugging Face used AI tools to spot the breach, isolate systems, reset keys, fix flaws, and alert police. That is a good response.

Yet it happened after the attacker accessed private data, stole keys, and moved through internal networks. OpenAI spotted the event only after its models reached another firm’s database. Detection limited the damage; it did not stop the attack.

A breach is a race between attack speed and defender response speed. AI agents speed up attacks with money. Human approval steps slow down defenders.

Hugging Face also surfaced an asymmetry that deserves more attention. When the team tried to analyze the attack with frontier models behind commercial APIs, the requests were blocked, because forensic work means submitting real attack commands, exploit payloads and command-and-control artifacts, and a safety filter cannot tell a responder from an attacker. They used an open-weight model on their own infrastructure instead. The attacker operated under no usage policy at all. The defender’s tooling was the side that got refused.

Detection is not prevention

Detection asks what happened. Prevention asks if it can happen at all. Against AI agents, the gap between those questions is where attacks succeed. The agent builds an attack out of single steps that look harmless on their own.

OpenAI noted this in its research: models operating over long horizons learn approval blind spots and work around them. Safety requires asking what a series of actions aims to achieve, not just if one step is allowed. Most systems check one call at a time: read a file, install a package, open a port. Every step passes. The full chain is the attack.

Six fixes for real defense

To survive automated attacks, defense must shift before code runs. That starts with treating data like code. Dataset loaders, templates, and configuration files all run instructions, so they demand strict controls: block remote execution by default, require review, and check the source of every outside file. Security teams must also control outgoing traffic with strict lists of approved destinations, while treating every proxy in that path as vulnerable. Finally, firms must remove permanent keys, issuing short-lived credentials for single jobs so a hacked worker machine yields almost nothing to an attacker.

Beyond individual machines, isolation must happen per task rather than per user, stopping an automated swarm from moving through internal networks. Authorization also needs to change. Checking single actions fails when an agent builds an attack out of simple steps. Systems must evaluate full sequences instead, setting caps on activity rates and automated spending to flag suspicious chains of events. Lastly, defenders need the freedom to act fast. Security teams need authority to isolate systems without waiting for leadership meetings, and the ability to vet a capable model you can run on your own infrastructure for forensics.

What boards must ask now

Corporate boards usually ask if the company uses AI responsibly. That’s rooted in compliance not safety. Directors should ask four direct questions instead: Which systems run outside code, and what checks sit in front of them? If an attacker targets a low-level employee, what keys and paths can they obtain? How fast do we move from the first warning to active containment? And can we analyze attack code on our own hardware, free from third-party rules?

The boundary has to move

Nothing in these disclosures required a capability that did not exist last year. The chain was ordinary. What changed is that an adversary can run that ordinary chain thousands of times over a weekend, learn from every failure, and never need a coffee break.

The models did not break the detection-and-response model. They exposed where we put the trust boundary. We put it after execution, and we assumed we would have time on the other side of it. We do not have that time anymore.

Move the boundary.

The post What the Hugging Face breach reveals about defense in the age of agentic AI appeared first on CyberScoop.

Anthropic says its AI accidentally hacked three companies during safety tests

Anthropic said it found three instances in which its models gained access to the live computer systems of outside organizations, according to a company blog post published Thursday.

The company said it began the review after OpenAI disclosed earlier this month that some of its models had exploited an unknown software flaw to escape an isolated test setup and reach production systems at Hugging Face, a platform for AI models and datasets. Anthropic checked over 141,000 evaluation runs for signs that Claude had reached the internet from environments meant to be closed off. It found six runs across three incidents, all tied to one outside testing partner, Irregular.

“Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone,” the blog post reads. “This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners. Moving forward, it will include expanding our continuous monitoring of evaluation transcripts for unexpected behavior, improving our investigation tooling, and conducting more rigorous assurance work with the vendors we rely on.”

In each case, Claude was working on a “capture the flag” puzzle, a common way to test a model’s hacking skill. The model is told a piece of secret data sits on another machine and must be found. Anthropic‘s prompts told Claude it had no internet access. That was false. A setup error at the partner’s end left the machines connected to the open internet, and Claude found real systems while searching for its target.

Believing the real systems were part of the simulation, Claude broke into them using simple methods such as guessing weak passwords and finding unprotected access points. In the most serious case, the model extracted login credentials and reached a database holding several hundred rows of live data. 

In a second case, Claude built and uploaded a malicious software package to PyPI, a public directory for Python code, after noticing a fake company’s setup guide told new hires to install a package that did not exist. The package stayed online for about an hour and was installed on 15 systems, including one belonging to a security firm whose scanner ran the code and had its credentials stolen. 

In the third case, Claude scanned about 9,000 targets before breaking into one company’s application through exposed credentials and a database attack method known as SQL injection.

The three incidents involved different Claude models: Opus 4.7, Mythos 5, and an unreleased internal test model. Anthropic said the models behaved differently once they encountered signs their targets were real. Opus 4.7 recognized the systems were live but kept attacking anyway. Mythos 5 also noticed signs it was on the open internet but talked itself out of that conclusion, pointing to a system date reading 2026 as proof it remained in a simulation. The internal test model, the newest of the three, stopped once it concluded the target was real.

Anthropic said it began its transcript review last week, halted all cybersecurity evaluations, identified the three incidents, and notified the partner and affected organizations earlier this week. Two of the three organizations had not detected the intrusions before being contacted, the company said; outreach to the third was ongoing. 

Anthropic framed the incident as a failure of its testing setup and oversight rather than a case of a model acting on its own, noting that in each instance the model was operating as intended. 

The company said it is working with the outside evaluator METR on an independent review and plans to release a redacted transcript of the PyPI incident within a week. It also said it would tighten monitoring of test environments run by outside partners and expand review of evaluation logs, framing the changes as part of what it called a blameless review of its own processes.

“These facts give us cautious optimism that with tighter monitoring and controls around evaluation infrastructure, as well as continued investment in alignment, this type of risk can be overcome,” the blog post reads. 

The post Anthropic says its AI accidentally hacked three companies during safety tests appeared first on CyberScoop.

Okta’s deal for Permiso aims to close gaps in identity threat detection

Okta announced Thursday it has signed a deal to buy Permiso Security, a cloud-based firm that tracks threats tied to human, machine, and AI-driven digital identities. 

Permiso specializes in spotting risks after a user or system has already logged in, an area the industry refers to as identity threat detection and response. The company draws on more than 2,500 signals gathered from over 70 identity-related partners to flag issues such as excessive access permissions, unused credentials, unusual behavior from AI agents, and violations of internal security policies.

Ely Kahn, Okta’s chief product officer, told CyberScoop that Permiso will allow Okta to merge two functions that have operated separately: real-time threat detection and identity security posture management. “Today those are two separate products that don’t really talk to each other,” he said. 

Combining them, Kahn said, produces sharper alerts for security teams. As an example, he described a hypothetical scenario where a dormant administrator account is flagged by posture-management tools that later shows a login from an unfamiliar IP address. “By combining those things, you now have a very high-confidence, high-fidelity alert that’s more actionable by a security operations team,” Kahn said. “A security operations team on its own might not care about the dormant account, but when you combine that with some threat signals, then it becomes a higher critical-level alert.”

A crucial part of the deal, according to Kahn, is that Permiso will bring visibility beyond Okta’s current threat detection products, telling CyberScoop that customers also rely on other identity systems, such as Microsoft Entra ID or Active Directory, that fall outside that view.

“For us to be a real player in the identity security space, we have to look beyond the Okta perspective and give folks a full view into their identity threats,” he said.

The acquisition also fits into a security landscape reshaped by artificial intelligence. According to figures cited by Okta, 58% of executives say their organizations experienced an AI-related security incident or a near miss within the past year. That trend has pushed identity companies like Okta to expand beyond authentication and into continuous monitoring of what accounts, including AI agents, actually do once inside a system.

“Agents will be breached,” Kahn said. “The most important thing you can do is ensure that if an agent is compromised, the blast radius is small,” through a narrowly defined, revocable identity tied to each agent. 

Among the capabilities Okta says it will gain is a tool called SandyClaw, which tests AI agent skills and prompts in an isolated environment before they are allowed into a customer’s systems, aiming to catch supply-chain attacks embedded in AI tools. Other planned additions include expanded tracking of AI agent behavior across cloud platforms and software-as-a-service tools, and automated systems to investigate and isolate AI agents that appear compromised or misconfigured.

The transaction is expected to close in the third quarter of Okta’s 2027 fiscal year, pending standard regulatory and closing conditions. Terms of the acquisition, including its purchase price, were not disclosed.

The post Okta’s deal for Permiso aims to close gaps in identity threat detection appeared first on CyberScoop.

OpenAI’s rogue AI agent shows why we need federal rules for autonomous systems

Months before the Hugging Face breach, Emergence AI published research that investigative journalist Ronan Farrow made public. Ten autonomous AI agents operated across five virtual environments for fifteen days without human intervention. Much of the attention focused on Grok 4.1 turning violent and Gemini 3 Flash committing 683 crimes.

What mattered more went unnoticed: Anthropic’s Claude Sonnet 4.6 built a peaceful democracy in isolation, then stole resources from neighboring environments the moment it joined a shared one. The lesson was clear: safety is not a model attribute. It emerges from the operating environment. The models didn’t change. Working as designed, their behavior evolved as the environment changed. The lesson is hard to ignore: The governance environment changed, and with it, the reward dynamics.

The story here concerns institutions, specifically OpenAI’s and Hugging Face’s, and how we must understand their recent security incident through that lens.

The industry agrees on how the Hugging Face breach happened. Cybersecurity experts have focused on the vulnerabilities, how they were used, and remediation. OpenAI has highlighted the model’s capabilities. Both conversations matter. What requires attention is why this breach is strategically important. After spending the past weekend discussing it with policymakers, security researchers, and industry practitioners in Aspen, I came away convinced we’re examining the wrong problem.

In 1961, Yale psychologist Stanley Milgram’s experiments revealed a broader truth: changing the institutional architecture changes behavior without changing the actor. The Emergence AI researchers didn’t change Claude’s agent. They changed the governance architecture that determined what constituted success for the system. Claude’s behavior changed with it.

OpenAI built a smart model but forgot to build a smarter room. That choice made the Hugging Face breach possible. Every organization now deploying autonomous agents now faces the same governance problem.

OpenAI gave the agent one objective: pass a cybersecurity evaluation. To stress-test it fully, they loosened the safety restrictions, and the agent found a shorter path. Rather than solving the evaluation directly, it found the answers outside the test environment, escaped its sandbox, and exploited a flaw in Hugging Face’s data-processing pipeline to reach live production systems. Over the weekend, with no human oversight, it ran more than 17,000 automated actions by escalating its own access, moving through internal systems, and harvesting credentials.

Hugging Face is one of the world’s most prominent AI companies, valued at approximately $4.5 billion. It provides the infrastructure that governments, defense organizations, and technology companies use to build and deploy AI. The agent was pursuing the objective it had been given. Breaking into Hugging Face was the fastest path to passing the test. Governance set the goal, the level of risk to accept, and who was accountable. Technical design determined whether those governance decisions could be enforced. As researchers James Shires and Max Smeets have argued, for a model capable enough to act on its own, testing and deployment must both must be governed the same way.

AI agent design requires baseline standards. Observability, including a monitoring layer that flags when an agent goes beyond its scope, is a baseline requirement. Human review also matters at escalation boundaries, like when an agent shifts from internal tools to external ones. When any agent crosses that boundary, what alert fires? What human reviews it? We lack clear answers to either. That is a governance choice, not simply a security failure. At best, this was a catastrophically failed test. At worst, how can we trust any frontier AI company to self-govern autonomous agent deployment?

More than a decade ago, the U.S. Department of Defense built the Comply-to-Connect (C2C) program: every device connecting to sensitive networks must prove it belongs there, or it is cut off from the network. C2C works because the quarantined actor stops. A laptop that fails verification goes offline and stays there. An autonomous AI agent adapts around enforcement. C2C was built for passive actors. Governance for autonomous agents must accommodate ones that adapt. Visibility is not enforcement, and enforcement is not control. We are missing all three.

A second failure that is not being discussed enough: the breach exploited an implicit trust assumption in Hugging Face’s data-processing pipeline, where inputs were treated as trusted without verification. After SolarWinds, the U.S. government set rules for software supply chain integrity: Executive Order 14028 and verification demands for federal software. The principle was simple: trust must be verified through proof. Those principles have not yet been comprehensively or consistently applied to the AI model supply chain. The rules remain weak. No one has been asked to explain why.

The answer is not a new framework. Existing frameworks suffice. C2C proved that visibility without enforcement leaves gaps, while Executive Order 14028 established that trust in software supply chains requires proof and verification. The challenge lies in applying these principles to a new category of actor. Congress, the Cybersecurity and Infrastructure Security Agency, or the Office of Management and Budget should make formal determinations that autonomous AI agents must follow the same rules as every other actor on a federal network. The framework exists; it must be updated.

The next incident is already in progress. It will show up in the logs as odd traffic, get handed to the same people who published these frameworks this week, and spark another round of recommendations no one acts upon. We’ve solved this problem before: for devices, for software, for supply chains. We know how to build smarter rooms. The tools exist. The will, the authority, and the decision to govern remains absent.

The post OpenAI’s rogue AI agent shows why we need federal rules for autonomous systems appeared first on CyberScoop.

AI-assisted security tools are finding more bugs, but the threat level has not changed

AI systems like Anthropic’s Project Glasswing and Microsoft’s MDASH are aiding in the discovery of vulnerabilities, filling the ever-growing pool of defects that defenders have to address before exploitation occurs. Yet, through the first half of 2026, these vulnerabilities were no more or less likely to be exploited than all vulnerabilities disclosed during that period, VulnCheck said in a report Tuesday. 

Concerns remain high about AI-discovered vulnerabilities fueling more attacks, but VulnCheck’s review of exploitation data shows that those fears are unfounded, at least so far. 

Patrick Garrity, security researcher at VulnCheck and report author, identified 1,061 vulnerabilities attributed to AI-assisted discovery during the first six months of the year. Of those vulnerabilities discovered by AI, 14 ( 1.3%) were exploited in the wild, a breakdown that aligns with the exploitation rate researchers observed across all vulnerabilities during the same period. 

“While AI-assisted vulnerability discovery clearly has value for both attackers and defenders, the data does not suggest that AI discovered vulnerabilities are inherently more likely to be exploited than those found through traditional methods,” Garrity wrote.

While AI’s contribution to actively exploited vulnerabilities was muted in the first half of the year, it’s too soon to assume that trend will continue. Moreover, none of these major vulnerability-hunting models were running for that full period. Project Glasswing rolled out in April, while Microsoft’s MDASH and OpenAI’s Daybreak were both unveiled in May.

The upward trend in Microsoft’s monthly Patch Tuesday indicates how much the floodgates might open through the remainder of the year as AI models discover more vulnerabilities. The company’s July security update contained an all-time-record of 622 vulnerabilities, besting the previous record-breaking June update with 206 vulnerabilities.

VulnCheck’s state of exploitation report also found that vulnerabilities were exploited much faster after CVE publication, speeding up from an average of 120 days in 2025 to 80 days during the first half of the year.

The intelligence firm also determined which technology categories were actively exploited most often. Content management systems accounted for nearly one-third of the 495 known exploited vulnerabilities VulnCheck identified during the first half of 2026. Network edge devices were responsible for almost 14%, followed by operating systems at nearly 9%, server software at 8%, and AI products — an emerging attack surface — at almost 6%.

The post AI-assisted security tools are finding more bugs, but the threat level has not changed appeared first on CyberScoop.

Microsoft debuts AI cybersecurity offerings as competition heats up

Microsoft threw its hat into the ring Monday in the increasingly heated competition among AI-powered cybersecurity offerings, unveiling tools that it claims are better and cheaper than its rivals.

The new agentic model MAI-Cyber-1-Flash, runs inside another Microsoft security tool, MDASH, and is part of an AI-powered security platform the company dubbed Project Perception.

Microsoft argues that its existing in-house capabilities give Project Perception an edge.

“Project Perception brings together signals, context, models and specialized agents into a continuously learning system of defense,” the company said in a blog post. “It can reason, prioritize and act at machine speed while keeping humans firmly in control and empowering them with powerful new workflows.”

Its release follows splashy AI cybersecurity suite debuts from OpenAI and Anthropic. Companies have been racing to advertise their AI offerings for their ability to find vulnerabilities, even as the most dire warnings about AI being used on the offensive side have yet to come to fruition.

As proof of its superiority, Microsoft said that MDASH with MAI-Cyber-1-Flash beat Mythos, Gemini and GPT on CyberGym, “the gold standard benchmark for evaluating how systems reason over large codebases to find real vulnerabilities in the code.” It scored 96%, 12 percentage points ahead of the next-best.

Microsoft said MAI-Cyber-1-Flash was built with a focus on safety first, and was independently assessed by a third party it didn’t name. AI cybersecurity systems made big news in the past week after OpenAI said its models broke free of its testing confinement to hack Hugging Face, a major AI code platform. 

Price also was a big part of Microsoft’s rollout: It said MAI-Cyber-1-Flash in MDASH can do the job at half the cost of other leading models.

“This is the benefit of building the harness, context/signals, and action space separate from one model family,” Microsoft Chairman and CEO Satya Nadella said on social media after the company unveiled Project Perception in San Francisco Monday. “By combining specialized models and data with the right agents, tools, security context, and harness, we can advance the frontier of cost to outcome.”

Project Perception enters public preview on Aug. 3, Microsoft said.

The post Microsoft debuts AI cybersecurity offerings as competition heats up appeared first on CyberScoop.

Microsoft, tech companies throw weight behind spread of open-source AI

Microsoft, along with more than two dozen tech companies, are pressing policymakers to support open-source AI systems and code across society, arguing that it will be a safer approach than attempting to restrict access or relying on a handful of closed, proprietary models.

The open letter, posted Friday, draws parallels to the software industry of the 1980s, when large businesses worried that open-source software code would cut into their business. While industry lost that battle, the end result was a vibrant ecosystem that now underpins much of the modern internet, government IT and even commercial software products.

It also created a “shared foundation of knowledge” that has fed countless future software projects and innovations.

“The United States now faces a similar choice with artificial intelligence,” the companies wrote. “Our AI leadership will be judged not by one frontier AI model, but by whether the United States builds a strong, open ecosystem that diffuses into every sector.”

Expanding access and support to open-source AI comes with meaningful security risk. Cybersecurity experts warn that one of the biggest beneficiaries of broadly available AI tools are  low-level criminals who until now lacked the technical expertise or resources to launch serious attacks.

Once a model is open weight, anyone can download it, customize it, strip it of any guardrails and use it for their own purposes. As open-source models have gotten better at creating deepfakes and other AI generated imagery, the danger of locally-customized CSAM and sexualized deepfakes could also grow.

But the letter argues that open-weight AI models are most beneficial to startups, universities, research labs and other small, ambitious organizations that can innovate and iterate the technology and make it more broadly useful to society.

“Open weights let every organization match the right model to the right job at the right cost, reserving frontier-scale capability for genuine frontier problems and running efficient specialized specialized models everywhere else,” The companies wrote. “That discipline is what will make AI economically sustainable as its use scales into the billions of everyday tasks.”

 For cybersecurity specifically, the letter argues that defenders armed with open-source AI will outpace attackers better than any closed model approach.

“In a world where cybersecurity attackers use advanced AI, defenders need access to models with comparable capabilities so they can detect, simulate, and respond to emerging threats,” the companies wrote. “Open models broaden defensive capability, increase transparency, and allow vulnerabilities to be discovered and remediated across many teams.”

Other notable companies signing the letter include Meta, Palantir, Perplexity, Mistral, NVIDIA, Mozilla, The Linux Foundation, Hugging Face, Dell Technologies and IBM.

US policymakers continue to grapple with balancing unrestrained support for the domestic AI industry and providing oversight and regulation of harms that result from their use.

The Trump administration has cycled through several frameworks since coming into office, first a laissez-faire approach within no restrictions, then an executive order creating a voluntary testing regime for industry, then the imposition of export controls on Anthropic’s Fable model and reportedly pressuring OpenAI to delay the release of their models out of cybersecurity concerns.

The letter comes as the Trump administration has reportedly considered an executive order that would restrict American access and availability to Chinese-made open-source models.

But the White House and US companies are trying to thread a needle in recognizing the overall benefits of an open source approach while being wary of doing anything that could potentially benefit their Chinese rivals.

Earlier this month the White House announced the creation of its Gold Eagle AI cybersecurity clearinghouse that would help coordinate government, private sector and civil society work finding and closing AI-discovered vulnerabilities. A big part of that effort, a senior White House official said, is supporting providers and maintainers of open-source AI tools.

The post Microsoft, tech companies throw weight behind spread of open-source AI appeared first on CyberScoop.

Malware is targeting AI tools in software development environments

Malware targeting AI coding assistants and software developers’ automated workflows is spreading into more environments with more capabilities, placing defenders at a growing disadvantage.

A malware strain dubbed Sandworm_Mode, first discovered by Socket in February, represents a growing threat to software development. According to a CrowdStrike report, the self-propagating worm can spread through code repositories with minimal detection, raising alarms about software supply chains.

The malware’s capabilities are extensive, but not especially unique compared to the series of supply-chain worms known as Shai-Hulud, and more recently Mini Shai-Hulud.

“This is the new trend,” Adam Meyers, senior vice president of counter adversary operations at CrowdStrike, told CyberScoop. “This is something we’re seeing more and more. It’s the new hotness right now.”

Sandworm_Mode targets and steals sensitive data, including credentials, keys and secrets that unlock paths to additional services and dependencies throughout the AI toolchain. This includes AI assistants, cloud providers, API keys for nine major LLM providers, CI/CD pipelines and automated systems that build, test and publish code.

These actions blend in with tens of thousands of other commands occurring daily in any given environment infused with AI development tools. 

“Trying to find the signal of something malicious happening is very difficult because there’s so much noise out there,” Meyers said. 

The worm also paces itself, setting multi-day delays to separate initial access from follow-on malicious activity — creating a gap in victims’ telemetry windows, which makes it even more challenging for defenders to detect and attribute the chain of infection properly. 

“AI agents are pulling down all of these different dependencies continuously throughout the day,” Meyers said. “When you’re looking downrange from the perspective of the security operations team, you’re just seeing everybody pulling down these dependencies, and these dependencies self-unpacking and executing, so it just gets really, really noisy to try to find something bad happening.”

The malware covers its tracks further with a bit of a mean streak, by automatically destroying compromised environments if it can’t spread or accomplish its objectives.

“It’s well thought-through, and well developed, so somebody spent some time caring and feeding this thing,” Meyer said.

Despite CrowdStrike’s four-month review of Sandworm_Mode, the cybersecurity firm has yet to gain a firm handle on its intent, but Meyers said it is designed to attain a strong foothold, which could enable long-term access.

CrowdStrike hasn’t determined who is responsible for the malware, yet Meyers said he doesn’t think TeamPCP, a threat group that’s been on a rampage through open-source software this year, is involved. 

“It could be a nation-state threat actor, or it could be an e-crime actor that’s looking to use this to then sell access to other organizations,” he said. “We don’t really know what the intention is.”

The state of Sandworm_Mode and whether it remains active is also unclear. CrowdStrike said it continues to observe recently active malicious supply-chain packages that follow similar but technically divergent patterns.

Ultimately, “the world has changed,” Meyers said, adding that many attackers are pursuing similar paths in the AI toolchain, requiring defenders and threat hunters to place a greater focus on this burgeoning mode of aggression.

The post Malware is targeting AI tools in software development environments appeared first on CyberScoop.

❌