A CIO I have known for years called me the day VCF 9.1 dropped. His team had been holding off on a private cloud refresh for almost a year, waiting to see what Broadcom did with the AI story. He read the launch announcement, read the partner brief, then called me with one question.
Is this the release I have been waiting for?
The honest answer is yes, with a caveat. VCF 9.1, released on May 5, is the first release where the private cloud story for production AI lines up with the economics. The numbers are real. The architecture finally fits the workload. And the conversation with your CFO just got a lot shorter.
Here is what actually changed, and what it means for the decision you are making this year.
The Numbers Behind the 40 Percent
Broadcom is leading with two figures. Up to 40 percent lower server costs for AI workloads. Up to 39 percent lower storage TCO. Those are the slide-ready numbers. The interesting part is what is underneath them.
Server cost reduction comes from two architectural changes that landed in 9.1. The first is NVMe memory tiering, which treats high-speed NVMe storage as a memory extension rather than as paging to disk. For inference workloads with predictable working sets, this lets you serve the same model with less DRAM per worker. DRAM is the most expensive line item in any AI host build, and right now it is also the most volatile. Memory prices have spiked through the first half of 2026 as AI demand pulls every available module out of the supply chain. The ability to serve the same workload with less DRAM is not a marketing claim. It is a procurement strategy.
The storage TCO number comes from vSAN compression and global deduplication, which finally ship as a first-class capability across the VCF stack. The deduplication works across the entire vSAN cluster rather than within individual disk groups, which is a meaningful difference when you are storing model weights, KV caches, and training artifacts that have significant redundancy across hosts.
When you put the two together, the per-workload cost of running inference on VCF 9.1 starts to compete with hyperscaler pricing on the workloads enterprises actually want to bring home. The math is not abstract. It is showing up in the conversations we are having with customers right now.
What VCF 9.1 Means for Multi-Tenant Private Cloud
The quieter story in the 9.1 release is the multi-tenant design. Broadcom published a design blueprint for self-service multi-tenant consumption on VCF, and the implications are bigger than they look at first read.
For most of the last decade, multi-tenancy on VMware infrastructure meant heavy customization, custom portals, and a fair amount of duct tape. The 9.1 design treats tenancy as a native concept, with isolated network and storage resources per tenant and a self-service consumption model that does not require a service provider to build the entire experience from scratch.
Why does this matter for an enterprise customer who has no intention of running a service provider business? Because most large enterprises now operate as their own internal service providers. The data science team consumes infrastructure. The application development team consumes infrastructure. The security team consumes infrastructure. Each of those is functionally a tenant, and each of them now has different requirements, different access controls, and different cost centers.
VCF 9.1 gives an internal IT team the tools to operate that way without standing up a parallel platform. For organizations that have been trying to deliver a real private cloud experience to internal teams, this release is the first time the platform supports that pattern without significant engineering investment.
For partners running managed virtual data center services, the new design changes the operating model in the same way. We can deliver tenant-aware infrastructure at lower per-tenant cost. That gets passed back to the customer.
Why the Partnership Story Tells You More Than the Feature List
The feature list in any major release is interesting, but it is not always the best signal of where the platform is going. The partnership announcements that landed alongside VCF 9.1 tell a clearer story.
Three announcements stood out. AMD GPU support, which means open-frameworks AI infrastructure is now first-class on VCF rather than an exception. Arista EVPN interoperability, which signals that the network fabric story is being built with the assumption that customers will not standardize on a single vendor. And the CrowdStrike integration for cyber recovery workflows, which gets at the question that every CIO is asking about ransomware recovery for AI training data.
The pattern across all three is that Broadcom is making VCF the integration point for the enterprise stack rather than the closed system. That is a different posture than what most observers expected eighteen months ago, and it is the right one for the moment.
The security story in 9.1 deserves its own mention. Continuous compliance is now built into the platform rather than bolted on, and integrated cyber recovery means the recovery workflow runs inside the VCF environment rather than through a separate tool. Combined with platform-level security hardening, the answer to the CISO question (what happens when this gets attacked) is now a real answer rather than a slide.
The Decision in Front of IT Leaders Right Now
If you are an IT leader looking at VCF 9.1, three things matter for your decision this year.
The first is whether the workloads you are running can take advantage of the architectural changes. Inference at scale benefits directly. Mixed enterprise environments benefit indirectly. Pure training workloads benefit less. Know which one you are operating before you build the business case.
The second is whether your team has the operational depth to run a modern private cloud at the level VCF 9.1 enables. The platform is more capable. It is also more demanding. The capability gap between what VCF 9.1 can do and what most internal teams have actually done is real, and it is the most common reason a refresh stalls.
The third is timing. Public cloud GPU pricing is not coming down. Memory prices are not coming down. Data sovereignty requirements are not relaxing. The window where the math favors a private cloud refresh for production AI is open right now, and the customers we are talking to are not waiting to see what the next quarter brings.
If you want to walk through what this looks like for your specific environment, the evoila team has been having this conversation every week since the launch. Start with our VMware solutions overview or reach out and we will bring the math.
What is the workload that is going to drive your VCF 9.1 decision this year?