Agregátor RSS
Anthropic makes changes to stop AI agents running amok again
Learning from the OpenAI-Hugging Face fiasco, as well as from recent revelations about its own model, Anthropic is revamping its security and alignment practices.
The company has established controls that flag when a model attempts to break out of a sandbox or successfully accesses the live internet, cordoned off its highest-risk test environments, and proposed a set of safety standards for its external testing partners, such as giving AI agents explicit instructions like “you should not access the internet.”
Anthropic conceded that three recent security incidents involving Claude reflect a “failure of operational security,” and also reveal issues with model reasoning capabilities and “recklessness.” Recent events “stressed that the urgency of improving our cybersecurity defenses is even higher than we previously believed,” the company noted.
Anthropic’s approach to security and alignmentThe company launched an investigation into its own security posture in July following the alarming OpenAI incident in which GPT models escaped a sandbox environment and arbitrarily attacked Hugging Face.
The company subsequently disclosed three situations during cybersecurity testing in which Claude models (Opus 4.7, Mythos 5, and an internal research model) accessed computer systems they should not have been allowed to touch.
The pre-release models were intentionally running without cyber safeguards, a common practice in early testing, and were able to exploit misconfigurations in a third-party’s settings using basic hacking techniques. Flaws in their reasoning led them to believe that all the entities they accessed, including those on the live internet, were in-scope for their capture-the-flag exercise.
Anthropic maintained that its internal security posture was not a contributing factor. The exploits occurred in a third party environment where internet access was mistakenly left open, so “the models had no need to ‘hack out’ of anything, even if they had been inclined to do so.”
Still, the incidents underscored the importance of hardening the sandbox itself, the company said. Until now, builders had been “largely relying on a single layer of defense,” that is, the configuration of the environment, when they should have had several layers in place, including monitoring, setting explicit boundaries within prompts, and sealing sandboxes.
Following its discoveries, Anthropic paused its internal and external evaluations of pre-release models. Higher-risk reinforcement learning (RL) environments for pre-release models were also halted for several weeks, while some sandboxes were moved to isolated settings with more stringent security gating.
Researchers built and deployed a classifier that automatically identified a model’s attempts to “aggressively probe” or break out of a testing environment, and when it inadvertently gained access to the live internet. They also inspected previous model evaluations to identify sandbox escapes or misconfigurations.
The company said they found no examples of models breaching sandbox boundaries. And while several sandbox configurations were flagged, none resulted in a compromise of outside systems.
When it comes to alignment, however, Anthropic researchers discovered a common failure with motivated reasoning: Models believed they were still in a simulated environment when they were actually connected to the internet; thus, they did not think they were breaching the boundaries of their testing lab. The other flaw was “recklessness,” a model’s seeming willingness to take harmful actions in the wild to ensure it achieved its goals.
In response, the researchers have been resampling models, testing them in skewed settings, and filtering out environments that incentivize cheating. They have also overhauled their production RL stack, established a stricter review process, changed specifications around model rewards, and tightened criteria for human reviewers who often dismissed false positives.
Further, Anthropic has expanded its monitoring of internal employee usage, and continues to red team its stack. Some high-risk RL environments remain paused for manual review.
Researchers are still studying how, exactly, models determined whether or not they were connected to the live internet, and if they knowingly disregarded signs that they were, due to flaws in their reasoning. They have also “empirically” determined that defects in training sandboxes disproportionately contribute to misaligned behavior.
However, that is not the sole reason. “Solving alignment will involve addressing a very wide range of potential problems,” Anthropic noted, “and future alignment incidents may involve different behaviors and different causes from those we have seen so far.”
Testing best practicesBecause the Claude security incidents occurred in third-party environments, and pre-release models often have weaker security safeguards, Anthropic is asking external testing companies to commit to a set of best practices.
Notably, they should explicitly set a model’s scope, and define permitted actions and what a model should and shouldn’t access. Boundaries must be stated as instructions (“you should not access the internet”) as opposed to descriptions of the environment (“you do not have internet access”).
Other suggested practices include running continuous, real-time monitoring, conducting evaluations inside a hardened sandbox with no internet access, and probing for vulnerabilities before even performing any testing, instructing the model itself to attempt to break out of its sandbox, and having it report back if it succeeds. Further, testers must confirm that evaluation challenges are actually solvable, rather than theoretical, so models don’t take actions outside their intended scope to achieve them.
“Because evaluations are sometimes conducted over hundreds or even thousands of runs, partners should ideally carry out this testing at length before beginning any evaluations,” Anthropic noted.
The company said it is developing companion best practices for those with access to Claude Mythos 5, which also runs without cyber safeguards.
Going forward, Anthropic described a “defense in depth” strategy. During alignment, a model is trained to be “helpful, honest, and harmless,” and is steered away from irreversible or contextually irrelevant actions. Models are given minimal permissions and their actions are limited, while offline monitoring notifies humans when things look wrong.
Finally, as a last resort, risky actions are blocked based on pre-determined classifiers, and humans can “pull the cord,” rework, or pause an agent when security layers fail.
Safety is just one part of itExperts call the move a positive step, if a basic one. Best practices like better isolation and monitoring should have been in place before agents were kicked off to hack systems, noted David Shipley of Beauceron Security.
“Better late than never,” he said, adding: “All these frontier firms are benefiting from felony-humblebragging-as-marketing, but there are some solid improvements in this announcement.”
The fact that the EU Act is now in force adds another layer of context, Shipley pointed out: Europe’s regulators are digging into the safety issues posed by frontier AI. These companies have had one of the fastest growth trajectories in tech history, and, concurrently, arguably the fastest regulatory response. Ideally, regulators are taking lessons from the “social media mess” and staying on emerging tech’s case before massive harms ensue, he said.
At the same time, frontier AI companies are watching high-profile court cases like the one targeting Meta.
This adds a third layer of context: The speed at which these companies are being sued is also on an unprecedented trajectory. “So, we should also read this blog as building a paper trail for a due diligence defense for regulators and courts,” Shipley noted.
This article originally appeared on CSOonline.
Citrix buys company that containerizes Windows desktop apps independently of the OS
Citrix on Tuesday announced that it has completed the acquisition of longtime partner Numecent, producer of technology that containerizes and manages Windows applications.
The acquisition builds on joint efforts to integrate Numecent’s management tool, Cloudpager, with Citrix Desktop-as-a-Service (DaaS) after an integration announced in April let administrators natively publish and manage the application containers through familiar Citrix workflows.
Numecent’s other product, Cloudpaging, packages Windows applications into isolated application containers independent of the underlying operating system, streaming them to Windows endpoints on demand rather than requiring them to be included in a desktop image.
Citrix plans to further integrate the technology into its platform, while also continuing Cloudpaging and Cloudpager support for physical Windows devices.
“Enterprise customers have told us for years that application management is one of the most painful parts of running a Windows environment,” said Shawn Bass, SVP and GM of Citrix DaaS, in the announcement of the acquisition. “Numecent has solved this in a genuinely elegant way. By bringing Cloudpaging and Cloudpager into Citrix, we can make this capability native to every DaaS and physical desktop deployment so IT teams get back the time they spend wrestling with images and app conflicts.”
Analysts and consultants said the move will help enterprise IT to some extent, but will also increase vendor lock-in with Citrix while potentially exposing enterprises to data security risks.
Good for Citrix customersGartner VP Analyst Stuart Downes said, “overall, this is a positive for Citrix customers,” but he stressed that the promised conversions “are not 100% compatible.”
He said, “low-level integrations into the kernel are generally not successful” because code that needs the lowest level of OS integration usually needs direct links to the hardware. Still, he estimated that applications at the low level probably account for only 2% of enterprise applications.
For the more typical apps, Downes said that there will likely be “north of 90% compatibility. It varies. There are quite a lot of complex factors in app virtualization.” But he emphasized that Numecent offers two components: Cloudpager and Cloudpaging, and “we have yet to see how Citrix will integrate both.”
Justin Greis, CEO of consulting firm Acceligence, also sees a lot of potential savings for the enterprise.
“Large companies can have thousands of Windows applications, including legacy, custom, industry-specific, and highly specialized applications,” he said. “Many have dependencies on particular versions of Windows, libraries, configurations, or desktop images. Every major desktop refresh, Windows migration, VDI program, cloud move, acquisition, or infrastructure modernization effort can therefore create another application testing and repackaging cycle. The ability to abstract more of the application layer from the environment underneath it can remove a meaningful amount of that friction.”
Noah Kenney, principal consultant at Digital 520, added that the theoretical advantage that Citrix can now offer has great enterprise potential.
But, he argued, this likely amounts to an enterprise IT pay less now, pay more later situation.
“There are operational savings here, which is why customers will adopt it, but the bill comes due when they try to leave,” Kenney said. “This is a good acquisition for Citrix and probably bad for enterprise leverage over time. Citrix can now lose the desktop and still keep the customer. Every application moved into Cloudpager raises the cost of the next migration. Customers get the simplification now and Citrix gets the switching cost later.”
Half rightSanchit Vir Gogia, chief analyst at Greyhound Research, said that he reads the containerization pitch as half right. “The packaging premise is valid. The cross-operating-system execution premise is not,” he said.
Gogia pointed out that Numecent Cloudpaging packages a Windows application with its dependencies and streams it to a Cloudpaging Player on a physical or virtual Windows endpoint, where it executes locally. “A Mac or Linux user reaches that application through Citrix’s remote delivery, where it still executes on Windows. That is cross-platform access, not cross-platform execution,” he said. “A Windows application does not become a Mac application merely because its pixels arrive on a Mac. The container is a packaging promise and the boundary of that promise is Windows.”
That said, he noted that there is still a lot of value in the Citrix arrangement, because Cloudpaging separates an application from a particular Windows image and carries that package across physical and virtual Windows environments, including Arm-based devices.
“The real advance is not escaping Windows,” Gogia explained. “It is making application change less dependent on desktop change. Microsoft’s own App Assure data puts enterprise application compatibility above 99.7%, and Cloudpaging’s commercial logic lives almost entirely inside the fraction that remains. At enterprise scale, the final 1% of applications can carry far more than 1% of the business risk.”
But, he added, “Existing Numecent customers need binding answers on entitlements, migration and exit. Citrix has bought control of a useful Windows application lifecycle. Control now has to prove itself, and the proof it owes customers is less complexity, not merely more control for Citrix.”
Possible riskHowever, consultant Brian Levine, executive director of FormerGov, pointed out that the nature of these new Citrix capabilities could expose users to serious security issues, including the risk of data exfiltration.
He sees Cloudpager as “essentially a privileged switch that can push software to every Windows endpoint at once, which is precisely the kind of mass-distribution channel that produced SolarWinds and Kaseya. Bolting it onto Citrix, whose NetScaler gear has been a favorite ransomware target through repeated ‘CitrixBleed’ flaws, may leave CIOs and organizations wondering who will focus on security for the combined entity, and how will it prevent the type of attacks we’ve seen against Citrix.”
Citrix was asked to comment on these security questions, but did not do so by publication time.
[webapps] Bludit CMS 3.20.0 - Reflected Cross-Site Scripting
[webapps] PodcastGenerator 3.2.9 - Stored XSS
[webapps] Ghost_CMS 6.19.0 - Remote Code Execution
[webapps] Langflow 1.10.0 - RCE
[hardware] Fullhan FH8626V100 - Multiple Vulnerabilities
Cops, CrowdStrike disrupt Sality botnet by poisoning the network and diverting into sinkholes
Content for Clicks: AI Is Tearing Up the Web’s Social Contract
As AI eats traffic, the best sites are locking it out, making reliable information harder to find.
For 30 years, the world wide web has run on a surprisingly profound social contract. Most sites are free for search engines to access, but if you use their content, you give credit by linking to the source.
Recently, that social contract has begun to collapse. Artificial intelligence tools are crawling sites not to link to them, but to train models and generate answers (which may or may not be accurate).
When you search for something, ChatGPT’s response or Google’s AI Overviews may still include links to sources, but they’re a kind of optional extra to the main answer.
This has triggered a bad dynamic for website owners, the public, and even AI companies themselves. As websites lose traffic (and revenue), many are beginning to block AI scraping tools, meaning AI results depend more on low-quality websites (many of which are also generated by AI). As a result, good information can be harder than ever to find.
How We Got HereIn the early days of the world wide web, search engines, and content creators came to an agreement about crawling (the practice of technologically examining a site to index it, so it can be served up in search results). Content creators would provide access to their sites for free and even allow search engines to reproduce small snippets of text.
In return, search engines provided links to the sites owned by content creators, who benefited from that web traffic. If content creators didn’t like the deal, they could prevent search engines from crawling their site with instructions in a file called robots.txt.
But if AI tools no longer provide web traffic, it cuts content creators out of the economic loop. There are also other costs associated with each visit to a website, so AI crawling can cost website providers money while not giving them any of the ad or other revenue that would come from human traffic. AI crawlers also crawl more deeply and more intensely than traditional web crawlers, magnifying that cost.
This change in traffic patterns isn’t a small or hypothetical problem. Cloudflare, a web hosting and service company that manages 30 percent or more of the top 10,000 sites on the internet, estimates over half of all web traffic is now AI bots.
Some of this will be AI agents supervised directly by people, but the majority will be crawlers. Site owners can use robots.txt to ask AI crawlers to stay off their sites—but some AI companies may ignore this polite request.
If the AI companies do honor the request, that can create a different problem. Sites containing misinformation are far less likely to ban AI crawlers, so the AI answers won’t be informed by high-quality sources.
What’s Happening in the Short TermOn the horizon is an event dubbed “Google Zero”—the day when through-traffic from Google drops to nothing. While some grey-haired diehards (like one of the authors of this piece) might still click through to verify AI answers, this traffic is rapidly dwindling, as a direct result of AI summaries.
A study of Wikipedia confirms this, showing that traffic in the English language version of the site dropped off quickly with the launch of AI summaries on Google in English, and that the same pattern occurred in other languages as AI summaries were rolled out. Never having to click through to get an answer might seem great for information seekers, but the reality is more complex.
Many sites are now blocking AI crawlers altogether. Site owners who decide to block AI crawlers are less likely to be linked in AI Overviews answers, even when the AI tool can still access the content to ground its answers (using a technique called retrieval-augmented generation).
Alternative “pay to crawl” models have been suggested as a way to compensate content creators, but haven’t gained traction.
Come September 15, Cloudflare sites will block AI crawlers by default on pages that contain advertising (and therefore make money for content creators).
This means up to 30 percent of the world’s top sites will no longer appear in Google AI Overviews summaries. It also means that much of what AI is being trained on will itself be AI-generated text.
What It Means for YouSo what does this mean when you’re looking for information? The quality of AI summaries is likely to go down, at least in the short term, while the new economics of the web get sorted out.
This will happen for two reasons. The first is that high-quality content is less likely to go into those AI summaries—one recent study found that already, around 1 in 6 sources used by AI search tools is itself an AI-generated website.
The second reason is that, as AI models are trained on more AI text, their output may degrade (a phenomenon known as model collapse).
As a result, search engines that depend less on AI may become more reliable. The challenge is finding one that doesn’t use an AI-based crawler. They do exist. ZDNet recommends Mojeek, PCMag recommends Brave, and Ban the Bots lists several, including one specifically for “small producer” content such as blogs.
For now, whatever search engine you’re using, the best thing you can do is to scroll down and click on some actual search results. This benefits content creators and is also more likely to give you more accurate information.
This article is republished from The Conversation under a Creative Commons license. Read the original article.
The post Content for Clicks: AI Is Tearing Up the Web’s Social Contract appeared first on SingularityHub.
Pravidla pro podporu v nezaměstnanosti se asi opět změní. Chystanou novinku už teď mnozí kritizují
Rozvaha nad zálohováním dat v Linuxu a jejich rychlou plnou obnovou
Softwarová sklizeň (2. 9. 2026): pořiďte si dlouhý rolovaný screenshot
Uživatelé si stěžují na blikání obrazovky s GeForce RTX s ovladači Nvidia 616.56
3D tištěné titanové konstrukce poprvé plavou a odolávají mořské vodě
Another Artifactory CVE under attack by AI agents or humans
Hackers abuse Faronics Deploy admin tool to install ScreenConnect
Attacker stole a METR API key, used $600K worth of credits, and no one noticed for weeks
Firefox helps iPhone users bypass ads on web sites while making money showing its own ads
- « první
- ‹ předchozí
- …
- 52
- 53
- 54
- 55
- 56
- 57
- 58
- 59
- 60
- …
- následující ›
- poslední »



