Kategorie
US and South Korea warn of Gunra ransomware targeting govt agencies
Hackers Breach Polish Power Plant Controls via Private Cellular Network and Shut Turbine
BdThemes Supply Chain Attack Poisons JSON to Create Rogue WordPress Admins
Meta’s new local model forces enterprises to recalculate AI costs and ROI
Meta on Monday rolled out a new 30-billion-parameter AI model that is optimized to run on a PC or Mac with a single GPU to offer always-on local agentic workflows rather than relying on the cloud.
The company has dubbed it Muse Glimmer. However, its hardware demands, including a GPU with a minimum of 24GB of VRAM, could make it difficult to justify for deployment at scale.
Although analysts and consultants agree that there is a tremendous enterprise appetite for running models locally, determining whether switching more systems from cloud to local makes fiscal sense is much more complex.
The hardware costs are tricky to calculate even today, with the VRAM needed depending on the particular applications to be run. But the far bigger consideration is that there is no way to determine what RAM costs will look like over the next 12-18 months, and there is an identical lack of visibility into how cloud prices might increase during the same timeframe.
That makes determining the better financial choice impossible.
Agents become capex, not opexNoah Kenney, principal consultant at Digital 520, noted that since RAM costs have increased “exponentially” over the last 12 months, cloud AI providers are also going to have to increase their prices. Beyond that, IT needs to anticipate logistical issues; key questions to ask are, “How quickly can you scale? Can you even get the hardware?”
“Meta just made agents a capital expense instead of an operating one,” Kenney said. “For two years, enterprises have been trained to rent intelligence by the token from someone else’s data center. Muse Glimmer runs the agent on a GPU you own, on the desk, with the meter switched off. That is a direct shot at the business model that cloud AI vendors are built on, and it comes from the one player with no cloud API revenue to protect.”
In its post announcing the new model, Meta pointed out that it has aggressively slimmed it down to try to make it efficient and cost-effective.
“At full precision, a 30-billion parameter model would require over 55 GB of memory — far more than any consumer GPU offers,” Meta said. “We use quantization techniques to compress the model’s weights to approximately 4-bit precision, shrinking the language model to under 20 GB. This leaves enough headroom for the model’s working memory, its KV cache, the perception encoder for image understanding, and the speculative decoding drafter to run simultaneously within a 24 GB or 32 GB envelope. We validated that this compression introduces minimal to no degradation on agentic tasks.”
More options for enterprisesMike Wilkes, enterprise CISO at Aikido Security, said that Meta’s move is significant in that it starts to give enterprises more options.
“The most important thing about Muse Glimmer is not that Meta has produced another capable model, it is that the economics and architecture of AI are beginning to move back toward the edge,” he said. “The financial comparison therefore becomes capital expenditure that can be amortized over several years versus an effectively perpetual per-token or per-request cloud operating expense.”
That means, he said, that an enterprise may rationally pay somewhat more for hardware if doing so gives it predictable AI costs, offline availability, control over model versions, freedom from sudden API pricing or access changes.
Independent cybersecurity and risk advisor Steven Eric Fisher also noted that the specs published by Meta don’t tell the full story.
The problem is that running locally versus in the cloud can generate a lengthy list of related expenses.
“Agentic workloads [in the cloud] can amplify consumption through reasoning, retries, tool calls, context growth, and evaluation, while local deployment [also] introduces hardware, power, lifecycle, support, and utilization costs,” he said, adding that even the RAM requirements need a lot of context.
“Meta’s stated 24GB and 32GB memory targets demonstrate that Glimmer can be loaded and executed on comparatively accessible hardware, but that is not the same as having sufficient capacity for meaningful agentic workloads,” Fisher said, pointing out that once other factors are considered, practical memory requirements can move beyond the 32 GB available on an Nvidia RTX 5090 GPU.
“In enterprise terms, this still places Glimmer primarily in high-end developer, data science, or dedicated AI workstations rather than the standard corporate desktop or laptop,” he said.
Better ROI not guaranteedJustin Greis, CEO of consulting firm Acceligence, agreed.
“I wouldn’t assume that moving inference from the cloud to the endpoint automatically produces a lower total cost of ownership,” he said, noting that variable cloud costs would be traded for the price of deployment, endpoint management, support, security, model updates, and potentially accelerated hardware refresh cycles for the local devices.
“Muse Glimmer is an important milestone because it makes local agentic AI technically viable. But technical viability and enterprise ROI are two different milestones,” Greis said. “I think Meta has crossed the first one. I do not think they have fully crossed the second one yet.”
However, Kenney argued that there are also other elements of the Meta rollout that make meaningful comparisons difficult.
“It is worth remembering that the local model is quantized, compressed to roughly 4-bit precision, while the cloud APIs you are comparing against typically serve full-precision models, so this is not a pure apples-to-apples cost comparison,” he said. “The ROI question is not local versus cloud on price alone. It is also a question of whether a company is willing to use a quantized model for the specific task. A cheaper agent that needs more retries or human correction can erase its savings fast.”
In addition, a local model often neglects to include every service that a cloud provider typically delivers.
“The cloud vendor was quietly handling updates, scaling, reliability, and security patching across your whole footprint. Bring the model in-house and every one of those becomes your problem, multiplied by every machine running it,” Kenney noted. “Most enterprises that consume AI as a service have no muscle for operating a fleet of local models, and that cost rarely gets adequate consideration in ROI conversations. The GPU is cheap. Patching a thousand of them is not.”
Sanchit Vir Gogia, chief analyst at Greyhound Research, also stressed that it can be difficult for IT to comprehensively anticipate the different cost variables.
“Finance leaders are right to feel skeptical. An engine bought for one employee burns capital whether or not it runs. Meta has shown that a thirty-billion-parameter agent can run on a single machine, and that is a real engineering result. What it has not shown is that such an agent works reliably at enterprise scale, or that a fleet of them can be operated safely. Model fit and production fit are different claims. A laptop must still run the employee’s actual job,” Gogia said. “The economics turn on the incremental hardware premium, refresh timing and actual utilization, measured against the price of the remote inference being displaced.”
On the other hand, Arun Chandrasekaran, a distinguished VP analyst with Gartner, said that he found it “very interesting that they have decided to release a smaller model that operates on the edge” and especially liked Meta’s use of the popular Apache license. But he would have preferred that they had released more than just a model.
“Enterprise customers are asking for a car and Meta is delivering an engine,” Chandrasekaran said. “They should have built something more like a platform solution.”
Hackers breached a small Polish energy plant via private APN last year
BdThemes plugins supply-chain hack creates rogue WordPress admins
OpenAI releases ChatGPT 5.6 Cyber, but it's only for approved users
Today’s AI hype-fest is partially IT’s fault
I came across a LinkedIn post the other day that described “hypegineering,” which the poster explained refers to the moment when AI “marketing becomes more innovative than the technology itself.”
That post came from Ralph Aboujaoude Diaz, the global head of GRC for British consumer services company Haleon. Until last year, Diaz worked in operations cybersecurity for consumer goods giant Philip Morris.
In his post, Diaz added that “side effects may include believing mediocre tech is revolutionary, confusing hype with progress, buying solutions to problems you don’t have and defending it like your job depends on it.”
As amusing as that might seem, the problem is frighteningly real. The sad truth is that enterprise IT executives are partly — perhaps mostly — to blame for the current hype around AI.
For many decades, senior IT leaders (pre-dating when it was called MIS) were the technology hype-deflators with razor-shape BS detection skills. That level of skepticism was needed. Tech vendors have always exaggerated and left out critical context when they weren’t outright lying about their products.
Enterprises trusted the IT gate-keepers to ferret out reality from the smoke and mirrors.
But there was a critical difference back then: senior management (especially CEOs, CFOs and board members) didn’t pay much attention to tech. They wanted the benefits, but they left the details to IT management to sort things out. With senior managers, benign neglect can be an incredibly good thing.
As much as we might all think that we want our bosses to really care about our efforts, be careful what you wish for. (For Dilbert fans, think of the pointy-haired boss; a boss who has strong beliefs but no understanding of technology is a nightmare.)
Alas, that is the situation many enterprises find themselves in today. IT management fully understands the level of absurd hype coming from a multitude of AI players. But unlike years and technologies past (RFID? NFC? Biometrics?), deflating hype balloons is easier when it’s comes solely from vendors. When senior managers start spouting this garbage, IT’s hype-deflation ability morphs into contradicting the very people at the top of their own corporate food chain.
That forces IT leaders who want to stay gainfully employed to do a lot of what used to be called lying. “Well, boss, yes these agentic systems have guardrails that will block destruction,” the lie begins, “but we can’t anticipate what any user or attacker might say in a prompt.”
That’s far more corporate acceptable than the actual truth: “Boss, a guardrail that an agent can disregard isn’t a guardrail. It’s more of a mild suggestion.”
Or IT could say, “Well, yes, boss, autonomous agents could exponentially increase efficiency — as long as you’re OK with the risk that they might send our internal data to the competition.”
For the sake of the company, tech leaders and decision-makers need to reassert themselves and perform serious reality checks on all AI efforts. But this is more fraught than merely contradicting a top boss. IT could be seen as embarrassing that top boss by pretty much proving that they were either naive or ignorant enough to be conned by the hype.
Telling them they’re wrong seems nicer than calling them stupid. And yet, someone in IT has to find a politically palatable way to do both.
New StormEncryptor ransomware used by former Medusa affiliate
Shipping 10–50× More Code? Watch This Webinar on Securing AI-Speed Development
China-Linked Hackers Deploy New StormEncryptor Ransomware, Likely via N-central Flaw
Apple’s real memory problem isn’t cost, it’s supply
Apple continues work to get White House go-ahead to source memory from Chinese supplier CXMT, which the US government has placed restrictions on.
For Apple, the issue comes down to simple math. With the cost of making iPhones up 38% because of eye-watering memory price increases — up almost 7-fold since the beginning of 2025 — it makes sense to make pragmatic choices when it comes to sourcing supply, particularly when other computer companies (including HP and Acer) already obtain memory from CXMT.
US resistance to the plan is that because of the way Apple packages memory on chip, the use of that memory might partly contravene the technology transfer restrictions the US has in place. Note: use of off-the-shelf components is fine.
Apple has been lobbying to use memory from the supplier only on devices made and sold in China and would likely argue this is reasonable given that both HP and Acer already use CXMT memory in products sold outside the US.
Why it’s really about supplyThe iPhone maker’s dilemma isn’t just about memory price, it’s also about ensuring it has enough component supply to satisfy demand for its products. This is plausibly a bigger challenge for the company, which is already warning of constrained supply because of lack of available memory. Samsung, Micron, and SK Hynix have allegedly already sold through all their DRAM production for 2027, which only sharpens Apple’s case to secure alternate sources.
With that in mind, it matters that China consumes around 20% of all Apple hardware sold globally. All the company needs is the go-ahead to use RAM from CXMT in those Chinese-made, China-sold devices.
That alone would effectively grow its usable memory supply by the same amount, because the non-restricted memory it currently has to use in Chinese-market products could be freed up for devices sold elsewhere, with CXMT covering the China-sold units instead.
With demand for its products accelerating —even amid a broad industry downturn — Apple really wants to make sure it has enough of the component to meet demand. Those powerful, on-premises 1.5TB RAM-equipped M7 Ultra Mac Studio AI clusters won’t exist unless it’s possible to find memory to put inside them.
The timeline challengeThe proposed arrangement with CXMT might not be a quick fix. Recent reporting indicated the company has already reached its production capacity for 2026, though it is working to double capacity by 2028. Apple has already tested memory chips from the company across products including iPhones and MacBooks.
While Apple this year has applied steep price increases across all its products (except for iPhones), it is now expected to increase the cost of the current iPhone 17 models perhaps as soon as this week.
Market reports already warn that Apple will raise prices on its soon-to-debut iPhone 18 Pro and iPhone Ultra devices next month, and we expect availability to be constrained, once again as a result of RAM shortages. Even if Apple gets the go-ahead to work with CXMT to close the gap, the positive impact of that arrangement is unlikely to kick in before next year.
Enjoy the painElsewhere, big memory manufacturers, including SK Hynix, have announced plans to invest in new DRAM manufacturing capacity. But this will not be operational until 2029, at the earliest.
This suggests constrained product availability and higher prices for the coming months — and the only way Apple, or anyone else, will be able to limit the impact of those price increases across the entire electronics industry will be if they can obtain additional stocks of memory from vendors outside the big three. If they cannot, you can kiss the age of abundance goodbye.
You can follow me on social media! Join me on BlueSky, LinkedIn, Mastodon and subscribe to The Core.
⚡ Weekly Recap: AI Goes Rogue, Metabase 0-Day, MCP Supply-Chain Attacks, and Router Backdoors
CISA: SonicWall SMA1000 flaws now exploited by ransomware gangs
When Credentials Are No Longer Enough: Device Trust in the AI Era
CopyKat Finds 122 Linux Kernel Objects That Can Amplify Memory Corruption
Kimsuky Builds Offline AI Stack to Boost Phishing and Automate Malware Development
Member of The Com sent to prison for blackmail, sextortion
New Passkey Attacks Can Recover Synced Private Keys or Bypass Phishing-Resistant MFA
Linux Vulnerability Prioritization with Threat Intelligence Insight
- 1
- 2
- 3
- 4
- 5
- 6
- 7
- 8
- 9
- …
- následující ›
- poslední »



