❌

Normal view

There are new articles available, click to refresh the page.
Before yesterdayMain stream

Irregular says ‘human oversight’ responsible for AI sandbox escape incidents

By: djohnson
17 August 2026 at 16:36

A company that runs AI testing labs involved in a series of breaches carried out by Anthropic and OpenAI’s cyber-focused models said the incident happened in part because they “unintentionally” provided the models with access to the internet.

In a blog posted Friday, Irregular said it partners with frontier labs to “stress test” AI models for security risks before they’re deployed, including simulations of their vulnerability research capabilities.

Those tests included non-public models from Anthropic and OpenAI. Separate disclosures from the companies confirmed these included Mythos 5, Claude Opus and GPT-5.6 Sol.

“As part of our review, we identified that a few interactions with our evaluation environments, in which internet access was unintentionally made available, led some models to take offensive security actions in the real world,” Irregular said in an unsigned blog.

In the incidents involving Anthropic’s models, testers initially believed the models were using a fictional company in its hacking simulations, only to learn later that the name unintentionally matched the domain of a real company. That led the model to “consider taking offensive action against” the real company in multiple training runs.

Irregular said this was due to “human oversight,” and that they typically do a background check to ensure they’re not using a real company for cybersecurity tests. The instructions provided to the model also included the target’s name and “internal addresses within our network” where the model could go to find the simulated environment.

While Mythos performed as intended in the majority of the thousands of tests, a small number of cases showed the model couldn’t distinguish between the fake and real domains and executed actual attacks on internet infrastructure, including exploiting vulnerabilities, extracting credentials and accessing a production database. In one instance, it even began targeting a site with a similar name when credentials for it were discovered online.

“Mainly, models believed they were in simulated environments, when they in fact took action in the real world,” the blog continued. “We are putting in place new and robust protocols to ensure setup issues do not occur while meeting the constraints of the testing process.”

The company said it plans to release a larger whitepaper breaking down the incidents and update their best practices for evaluation setups in the future.

While the companies have drawn criticism from some in the cybersecurity community for failing to securely design their sandboxes for testing, experts have said AI models are known to grind away on fulfilling a command  until they can find a workaround. Additionally, Irregular said granting some level of internet access to models is necessary to fully test out their cybersecurity capabilities.

“Controlled internet access, while it may allow models to exceed containment boundaries, is at times critical for realistic evaluations; without it, threat scenarios lose fidelity, undercutting the purpose of the challenge to reduce post-release risk of models being misused by attackers – as attackers in the real world do rely on the internet,” the company wrote.

According to the blog, Irregular has since “remediated” the “issues that led to these interactions,” though few details are provided.

However, the researchers say the engagement revealed critical gaps in their security practices. 

They plan to improve documentation of evaluation setups, deploy better log monitoring tools capable of tracking “the extreme amount of data generated by the traffic,” revise their threat models to account for rogue AI behavior, and establish faster information sharing between stakeholders.

“Looking further down the line, models will only get stronger. While in this case we believe that better implementation of existing safeguards could prevent most incidents of this kind, as models become stronger, this may not be the case,” Irregular wrote. “We therefore believe this opportunity should be leveraged by us and the community to be proactive and establish forward-looking protocols and [research and development] efforts.”

The post Irregular says ‘human oversight’ responsible for AI sandbox escape incidents appeared first on CyberScoop.

NIST wants to overhaul its vulnerability database for the AI age

By: djohnson
11 August 2026 at 11:36

The National Institute for Standards and Technology is looking for input on how to overhaul its vulnerability reporting process to better meet the challenges of an “evolving cybersecurity landscape increasingly shaped by artificial intelligence and machine-consumable security data.”

In a request for information set to publish Wednesday in the Federal Register, NIST said its National Vulnerability Database, one of the primary ways the federal government coordinates with security researchers to identify and fix software vulnerabilities, must be updated for the AI age.

NIST is concerned that as large language models become more capable of finding and exploiting vulnerabilities at scale, the NVD’s process must be updated.

“The inadequacies of traditional vulnerability management approaches, which center on periodic scanning, static prioritization, and manual remediation, are increasingly apparent,” the RFI states.

NIST believes AI hacking tools are contributing to recent trends in vulnerability reporting. The NVD has seen increased volume and complexity of disclosed vulnerabilities, inconsistent data quality, increased reliance on automation and machine-readable security data, and “demand for near real-time vulnerability enrichment” from defenders facing faster threats.But NIST believes these challenges also present an “opportunity to transform the vulnerability management ecosystem” through proactive reforms and NVD innovation.

That’s where the public comes in. NIST is posing a series of questions that must be answered before a larger strategy can be developed. Many of their questions focus on better integrating automation – AI or otherwise – into the process.

The agency asked for insight on how defenders could better leverage automation in the vulnerability reporting process; which capabilities, products and processes would help more quickly disseminate information to stakeholders, how to build transparency and auditability into AI-driven decisionmaking, and what role AI should play in automated vulnerability remediation.

“NIST intends to support a future-ready vulnerability management ecosystem that is continuous, contextual, and automated, while enabling cybersecurity practices to respond appropriately to real-world threats and business priorities,” the RFI states.

The NIST effort to revamp its vulnerability database comes a month after the Trump administration rolled out a new federal clearinghouse, overseen by the Department of Treasury, for sharing AI threat information between government and the private sector called “Gold Eagle.”

It’s not clear how Treasury’s process will interact with NIST’s database. The White House also partnered with Carnegie Mellon’s Software Engineering Institute to create the Vulnerability Information and Coordination Environment, (VINCE) which will collect and distribute reports on AI-discovered vulnerabilities.

The post NIST wants to overhaul its vulnerability database for the AI age appeared first on CyberScoop.

OpenAI says Daybreak will expand to offer specialized cyber services 

By: djohnson
10 August 2026 at 16:55

OpenAI announced Monday  it was expanding access to its frontier models for defensive cybersecurity, detailing different defensive and red-teaming workflows and a new partner program with major cybersecurity product providers.

In a pair of blogs posted Monday, OpenAI said it was updating its Daybreak program  – which provides unreleased frontier models to private organizations and governments for defensive cybersecurity work – and introducing a new model variant.

Daybreak Blue, powered by OpenAI’s ChatGPT-5.6-Sol, would operate with lower cybersecurity safeguards compared to other commercially available models and is described as “a recommended starting point for most defenders” that supports tasks like vulnerability discovery, secure code review, malware analysis, incident response and patch validation. 

Daybreak Red, meant for more advanced red-teaming, would provide access to a new model, dubbed GPT-5.6-Cyber, that the company said is more purpose-trained for finding vulnerabilities and testing (or exploiting) them. The model is also less likely to refuse requests around “dual-use cyber tasks.”

According to OpenAI, the organizations in Daybreak Red will have their use closely monitored and supervised, as GPT-5.6-Cyber is significantly more capable in carrying out malicious cyber tasks than Sol. A security evaluation the company devised tested both models on complex requests, including exploit chain development, authentication bypass, privilege escalation and other hacking tasks. Sol succeeded in 1.5% of the requests, while Cyber completed 95%.

OpenAI said it plans to publish a more detailed system card for GPT-5.6-Cyber at a later date.

“Models running with reduced safeguards carry risks beyond standard model usage, whether from misuse or misalignment,” the company said in a blog. “Despite these risks, we believe that democratizing access to frontier intelligence for defenders is crucial to accelerating and automating cyber defense.”

Additionally, OpenAI announced a partnership program with 16 major cybersecurity providers, saying organizations could access their models through their existing security services. The partners include IBM, CrowdStrike, Accenture, Ernst & Young, KPMG, Palo Alto Networks, Cisco, Cloudflare, Sophos and others. 

“These partners bring deep security expertise and established relationships with organizations around the world,” OpenAI said in its blog. “By bringing our frontier cyber models into their services, we can help more defenders find serious vulnerabilities, validate which ones matter, and fix them faster.”

Companies like OpenAI, Anthropic and others are trying to rebalance their priorities after a string of AI-agent sandbox escapes have rattled policymakers and caused some cybersecurity experts to question if AI companies are doing enough to properly isolate the models from the internet during testing. Last week, OpenAI said it was intentionally slowing down development of its newer “Astra” model in order to develop better guardrails to restrain its behavior.

Cybersecurity and AI experts have told CyberScoop that while AI systems have greatly improved at finding and exploiting vulnerabilities in software code, they still require substantial human guidance and supporting infrastructure to operate as intended.

Additionally, some research has shown that without such guidance, even near-frontier models can struggle to fully patch a discovered vulnerability or avoid introducing new bugs with their fixes.

The post OpenAI says Daybreak will expand to offer specialized cyber services  appeared first on CyberScoop.

White House accuses Chinese company of distilling Anthropic’s Fable

By: djohnson
22 July 2026 at 12:45

A top White House technology official is accusing a Chinese company of distilling Anthropic’s models to create their own AI product.

Michael Kratsios, who leads the White House Office of Science and Technology Policy, claimed that Moonshot AI, a Beijing, China-based AI company, had distilled Anthropic’s recently-released Fable model to develop its own K3 model.

“To do this they developed a sophisticated internal platform to conduct large scale distillation against U.S. models, allowing them to quickly switch between multiple methods of access to avoid detection,” Kratsios wrote on X Wednesday.

Kratsios also said the company has used GB300 servers – either newly acquired or through Thailand – to train its AI models.

“The United States strongly supports the free and fair development of AI, including a thriving competitive ecosystem that spans frontier models, specialized systems, open-source frameworks, and open-weight models,” Kratsios continued. “Legitimate AI distillation used to create smaller, more efficient models plays a vital role in this open innovation ecosystem. However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.”

Kratsios did not provide details on how the U.S. government learned that K3 had been distilled from Anthropic’s model. 

Frontier AI companies in the U.S. have pressed policymakers to make it more difficult for third-parties to copy or duplicate advanced commercial models, calling it a form of intellectual property theft.

On their website, Moonshot AI describes its Kimi K3 model as the first open 2.8 trillion parameter model, and promotes its lower token costs while still delivering near-frontier performance. 

“While its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrated frontier-level performance across our evaluation suite, consistently outperforming other tested models,” the company said on its website. 

A request for comment sent to Moonshot AI was not returned before this article’s publication. 

Piyush Sharma, CEO of Tuskira, an AI cybersecurity detection and response company, said distillation of AI models allows developers many of a model’s core capabilities. He pointed to another example when Anthropic earlier this year accused Chinese company Alibaba of distilling their Claude AI model.

According to Anthropic, the campaign used 25,000 fraudulent accounts to run 28.8 million interactions on Claude over six weeks. Given that kind of volume “the goal was clearly replication,” he said. 

“When a model has learned to reason through software weaknesses, security gaps, and attack paths, copying its behavior also copies that analytical capability,” said Sharma.

In April, Rep. Andrew Garbarino, R-N.Y., who chairs the House Homeland Security Committee and Rep. John Moolenaar, R-Mich., Chair of the Select Committee on China, announced they were conducting a joint investigation into the integration of Chinese AI models.

The committees said the inquiry will also focus on “examining a pattern of conduct by [Chinese]-based AI laboratories involving the large-scale theft of proprietary capabilities from American frontier AI systems through adversarial distillation” as well as “ the redistribution of those stolen capabilities as open-weight models available for global download, and the incorporation of PRC-origin models into products used daily by hundreds of thousands of American developers and engineers.”

Western governments and industry accuse Chinese companies of routinely stealing their technology, intellectual property and other trade secrets, often with the tacit support of Beijing. The copying of AI models would continue a long and established tradition of Chinese-sponsored intellectual property theft.

However, while distillation attacks by foreign governments or companies on U.S. frontier companies can have real national security implications, it’s still a fraught question of where policymakers should draw the line.

The AI industry, which includes not just frontier companies but large businesses with their own bespoke models, smaller proprietary startups and a vibrant open-source ecosystem, routinely share and use third-party data, including critical code and training sets for AI models.

Further, U.S. frontier AI companies have built and trained their world leading models in large part by crawling the open internet, ingesting content created and produced by others. Critics (and multiple ongoing lawsuits) argue that AI companies like OpenAI and Anthropic built their empires on data and content from others, taken almost entirely without consent or compensation.

The post White House accuses Chinese company of distilling Anthropic’s Fable appeared first on CyberScoop.

White House details ‘Gold Eagle’ clearinghouse for AI cyber threats

By: djohnson
14 July 2026 at 17:44

The Trump administration unveiled its new federal clearinghouse for sharing AI cyber threat information between the government and private sector, and said the project is already receiving threat intelligence on cybersecurity vulnerabilities and prioritizing patching.

Created last month through a White House executive order, “Gold Eagle” will be managed by the Department of the Treasury, with contributions from the Cybersecurity and Infrastructure Security Agency, Department of Homeland Security, and Department of Defense, as well as open-source software providers, critical infrastructure operators and industry.

“Under President Trump’s leadership, the Treasury Department is working hand in hand with the private sector to safeguard our financial institutions, close vulnerabilities, and protect the integrity of the U.S. financial system,” Secretary of the Treasury Scott Bessent said in a statement. “Treasury, along with our partner agencies, will continue to harness frontier AI capabilities to stay ahead of our adversaries and defend the American people from emerging threats.”

Gold Eagle is meant to help both public and private organizations find, fix and patch vulnerabilities found using AI tools before they’re discovered and exploited by bad actors. The work will involve using AI to find cybersecurity vulnerabilities in victim systems and software, and Secretary of Homeland Security Markwayne Mullin said it would also further explore ways for the technology to be leveraged for cyber defense.

A senior White House official told reporters on a background call that closed source models from frontier AI models, including Anthropic’s Mythos, will be used to discover vulnerabilities.

White House officials said they worked with the Software Engineering Institute, SEI at Carnegie Mellon University to develop a new platform, the Vulnerability Information and Coordination Environment – or VINTS – to receive third-party reports on AI-discovered vulnerabilities. According to the White House, the system has already begun collecting intelligence on vulnerabilities and prioritizing patches.

“I think on the early side of this, we have seen that the scale of vulnerability discovery, particularly with users of new technology to scan their system, is something that is a step function change [than] we’ve seen seen before,” the official said.

As AI models have improved at carrying out core cybersecurity-related tasks – like scanning code for vulnerabilities or developing proof-of-concept exploit code – cybersecurity experts and policymakers have become increasingly worried. The modern internet is rife with insecure code, misconfigurations and other mistakes that can be identified and exploited faster than ever before using AI tools.

Vulnerabilities in open-source software can be both widespread and hidden, as many commercial software products on the market rely on open-source code but few bother to document it. When hackers compromised a logging tool in the Log4J open-source Apache software library in 2021, it required a massive, multi-month coordination effort by CISA, the private sector and other stakeholders to find and fix affected pieces of software.

The White House official said the work of Gold Eagle is reflective of the administration’s “full support” of U.S. open-source software providers and maintainers.

Open source tools are “vital to systems that run throughout our country and daily life,” a senior administration official said, speaking to reporters on background. “It is being maintained by a talented group of people and entities and we will do everything we can to support the strength of that community.”

Michael Daniel, former White House cyber coordinator under President Barack Obama, told CyberScoop that AI is still so new that policymakers continue to observe its impact and adapt. While some existing communication channels for sharing cybersecurity threat information could probably be duplicated for tracking AI threats, there is still much for policymakers to learn more about the technology, the kind of threats it produces and its ecosystem of stakeholders.

“It may turn out at the end of the day that phishing is still phishing, and the fact that now you’ve got AI tools doing it, it’s still phishing. Or there may be something fundamentally different about it that we need to figure out how to combat and share information around,” he said.

The post White House details ‘Gold Eagle’ clearinghouse for AI cyber threats appeared first on CyberScoop.

French nonprofit starts global intelligence and research hub for AI cyber threats 

By: djohnson
8 July 2026 at 14:15

The Paris Peace Forum, a French non-profit that has convened world leaders on global security issues, is launching a new project to bring together international experts to assess AI-related threats to global internet infrastructure.

The Integrated Network for Trusted AI in Cyberspace (INTAiC) will tap researchers and civil society experts from government and the private sector, analyzing current AI cyber threats from the field and creating “forward-looking” reports on how the technology will impact society and what organizations can do to respond.

One of the project’s top goals is to create an international, quick-response coalition of government and business to address AI-related threats, similar to coordination mechanisms that exist in other areas of cybersecurity.

“Evidence fragmentation on AI-driven cyber threats isn’t incidental — it’s structural: those defending networks and those securing AI systems have long worked in separate spheres,” said Adrien Abecassis, policy initiatives director for the Paris Peace Forum. “That’s exactly why INTAiC is unique — it’s built to turn those fragments into one comparable reading of the threat, because this is a challenge no actor can meet alone.”

The network already lists a number of prominent businesses and organizations, including Microsoft, the Cyber Threat Alliance, the Cloud Security Alliance, Orange Cyberdefense and others.

According to the forum, INTAiC’s work will focus primarily on two, separate workstreams. One is a single and regularly updated resource for defenders to stay up to date on how AI is reshaping cyber threats. The resource is focused more on attacker capabilities, different forms of misuse and the impact on security operations rather than isolated incidents.

“The result is a common reference point, grounded in reality, that gives policymakers a clearer measure of the threat and identifies the risks most deserving of collective attention,” the Forum said in a release.

The second workstream will focus on evaluating and preventing cyber risks associated with AI, building up a base of independent third-party experts who can provide neutral or unbiased assessments of frontier model cyber capabilities. That work will pull in governments, research institutions and non-profits to develop new organizational and funding pathways to support that kind of research.

While the U.S. federal government has come a long way in recent years building up its own capacity to test and study AI cyber threats, much of the access and technical expertise around frontier model capabilities are concentrated within commercial AI companies. This has at times created concerns that federal agencies were being overly reliant on AI companies to explain how the technology worked and walk them through the possible threat scenarios.

As Anthropic and OpenAI have rolled out defensive cybersecurity programs like Project Glasswing and the Trusted Access for Cyber program, access to those models have become available to a wider group of researchers and organizations.

The Paris Peace Forum intends to brief the public further on INTAiC’s work and accomplishments in Paris later this year during the organization’s annual conference in November.

The post French nonprofit starts global intelligence and research hub for AI cyber threats  appeared first on CyberScoop.

US lifting export control restrictions on Anthropic’s Mythos, Fable

By: djohnson
1 July 2026 at 09:36

Anthropic has announced its Fable 5 and Mythos 5 models will once again be available to the public as it has reached an agreement with the Commerce Department to deploy the AI models with new guardrails and classifiers meant to address jailbreaks.

In a blog posted Tuesday, Anthropic said that export controls that prevented their sale to foreign companies and individuals have been lifted after weeks of negotiation with the White House and Commerce Department. The company has also restored access to the model for U.S. users.

The export controls were put in place after the Trump administration became alarmed by a threat intelligence report from Amazon claiming to have jailbroken Fable’s cybersecurity capabilities.

On X, Secretary of Commerce Howard Lutnick appeared to confirm that the restrictions would be lifted.

“Over the past two weeks, we have worked closely with Anthropic to analyze and approve Fable 5 to ensure alignment across the US Government and strengthen America’s leadership in AI,” Lutnick wrote.

The administration levied the export controls after becoming concerned that the release of Fable 5 would lead to the model being jailbroken, giving users access to cybersecurity and other capabilities that Anthropic has said could wreak havoc on the open internet if  placed in the wrong hands. The Amazon report convinced administration officials that such jailbreaks were on the immediate horizon.

However, one oddity of the administration’s decision is that the capabilities described in the Amazon report, by all accounts, are not cutting-edge. Scanning code and breaking down how to exploit vulnerabilities for a user is already possible with existing models.

Anthropic confirmed that, saying that further testing found that equivalent and lesser models like ChatGPT 5.5, Claude Opus 4.8 and Kimi K2.7 could identify the same vulnerabilities as Fable did in the Amazon report, while a half dozen existing models were able to produce the same proof of concept code as Fable.

Crucially, Anthropic reiterated that they have yet to see a jailbreak that affects the model’s restrictions on cybersecurity and biology work, though they did call this instance “a borderline case.” Indeed, some cybersecurity professionals have publicly complained that existing safety guardrails on Fable 5 blocked many routine defensive cybersecurity work in addition to malicious use cases.

“Importantly, the reported technique did not expose any unique Mythos-level cyber capabilities,” the blog continued. “The behavior reflected a borderline case for Fable 5’s safeguards…there are some tasks that are unlikely to be dangerous but are nonetheless blocked by the safeguards out of an abundance of caution. The reported technique allowed access to one such behavior, but it only involved routine defensive cybersecurity work.”

Anthropic said it has trained new safety classifiers to target and block the behaviors described in the Amazon report and notify users when it happens, and that the new safeguards have been stress tested by the federal Center for AI Standards and Innovation. The new classifiers will block the techniques “99.9%” of the time, but Anthropic said they’re not expected to block all lower risk routine cyberdefense capabilities, just the most harmful ones.

The restrictions will likely make it even harder to use Fable 5 for defensive cybersecurity. One effect the company expects is that more “benign” requests for routine coding and debugging tasks will be flagged by the system.

Christopher Padilla, former Assistant Secretary for Commerce for export administration in the George W. Bush administration, said that while it’s “good news” the controls have ultimately been lifted, the Trump administration’s AI policy stumbles over the past two years illustrate “the risks of ad hoc, transactional policymaking.”

In a LinkedIn post, Padilla called the Trump administration’s approach chaotic and unpredictable — the opposite of the clear, consistent rules industry depends on. While Vice President J.D. Vance mocked AI safety regulations in a speech in Europe last year, the administration has quietly partnered with OpenAI and Anthropic on voluntary national security testing, especially as frontier models began showing advanced automation and cyberattack capabilities.

That national security arrangement was supposedly codified in a White House executive order last month, shaped heavily by industry boosters who feared regulatory delays would slow U.S. development. But days after Fable’s release, Commerce imposed new export controls on Anthropic’s models anyway.

Padilla called proposed AI safety regulations by the Biden administration “flawed and overly complex” but nevertheless predictable compared to the status quo. Instead of replacing those proposed regulations with their own vision, the Trump White House has been “to put it mildly, all over the place on AI policy.”

“The same BIS that stopped Fable and Mythos has a permissive policy for exporting high-end AI semiconductors to China — in exchange for a cut of the take,” said Padilla, referencing the Trump administration’s lifting of export controls on advanced AI chips. “This is not a smart way to make policy. Bad for industry competitiveness and for national security.”

The post US lifting export control restrictions on Anthropic’s Mythos, Fable appeared first on CyberScoop.

Warner bill would create federally vetted list for secure, trustworthy AI agents

By: djohnson
29 June 2026 at 17:29

A new Senate draft bill would establish a list of AI agent software providers that people can use to establish human ownership and securely run agents on social media and other online platforms.

The Artificial Intelligence Access, Gatekeeper Exchange, and Nondiscriminatory Transfer (AI AGENT) Act, led by Sen. Mark Warner, D-Va., would allow end users of large online platforms with more than 50 million customers or subscribers per month the right to choose at least one AI agent provider who complies with security and identity standards developed by the Federal Trade Commission.

Such agents are increasingly making decisions on behalf of users, like shopping, posting content on social media, or changing account settings, sometimes without the user’s consent or knowledge.

Under the bill, the FTC would certify independent bodies to vet AI agent vendors. These certification bodies would ensure products meet baseline protections for privacy, data security and acting in the user’s interest. The bill would also require providers to link each AI agent to its human operator’s identity and to include built-in controls that let users clearly grant or revoke permission for the agent to act on their behalf.

While the commission cannot bar platforms from using AI agent providers that fail to meet those standards, it can deregister violators from the FTC list.

The bill is a discussion draft, and Warner said he was releasing it now to receive feedback before introducing a formal version for consideration in the Senate.

“As agentic AI transforms how Americans interact with technology, consumers deserve a real choice in the marketplace – and AI agents must be accountable to the people they serve,” Warner said in a statement. “This discussion draft is a major step toward building a clear federal framework that promotes innovation, protects consumers, and ensures the United States continues to lead the world in emerging technology.”

Last year, Morgan Stanley estimated that nearly one-in-four (23%) Americans made purchases using AI over a 30-day period, and that agentic shoppers could account for potentially hundreds of billions of dollars in online commerce by 2030.

But AI agents can still be unreliable or erratic. They can make absurd purchases that a user would never knowingly approve, leak sensitive data or act contrary to a user’s interest.

As more agents flood the internet, it increases the likelihood of AI bots interacting with and buying from other AI bots – underscoring the need for safe or regulated user solutions that can verify accountable human identities behind AI activity and provide baseline security and privacy protections.

The Trump administration is trying to find its own baseline for regulating frontier models. Earlier this month the Department of Commerce placed export controls on Anthropic’s Mythos 5 and Fable 5 models, and the two parties are attempting to negotiate a framework to provide government oversight of newer releases.

An AI executive order released by the Trump administration set up a voluntary 30-day testing program for AI companies to submit certain frontier models for testing and evaluation, but the administration imposed the export controls days after Anthropic released Fable 5 publicly, reportedly citing concerns that the model could be jailbroken.

Anthropic claims that extensive internal testing has identified no universal jailbreaks for Fable 5 and that third-party research released thus far hasn’t shown that their guardrails preventing access to the model’s enhanced cybersecurity or biological capabilities have been circumvented. Those are the capabilities that Anthropic cited when it held back its newest model, Mythos, from public release.

The post Warner bill would create federally vetted list for secure, trustworthy AI agents appeared first on CyberScoop.

❌
❌