Skip to content
Edouard Topin's Blog
What's new in VCF 9.1 / Series 05/05

What's new in VCF 9.1.1: the Day-2 release

VCF 9.1.1 adds multi-tenant model sharing, an AI operations assistant, VKS observability, native GitOps, EVPN improvements, and a safer upgrade path.

Edouard Topin
12 min read
Editorial diagram of VCF 9.1.1 connecting private AI, operations, Kubernetes, EVPN networking, and GitOps

VCF 9.1 changed the platform’s capabilities. VCF 9.1.1 changes how those capabilities are operated. The version number looks like a maintenance release, but the operational surface is broader: shared AI models, conversational diagnostics, two-second Kubernetes telemetry, native GitOps, expanded identity and certificate workflows, and a more practical path into VCF 9.1.

The useful question is not “what are the release-note bullets?” It is “which changes justify an upgrade, which ones alter the architecture, and which ones are still experiments?” This article answers that question for architects, infrastructure operators, and platform teams.

GA since 3 September 2026AI + VKS operationsTech Preview boundaries

TL;DR

  • Upgrade for operations, not for a new hypervisor headline. The strongest reasons are supported back-in-time upgrade paths, cumulative fixes, a smaller management footprint, VKS observability, and broader fleet management.
  • Do not confuse inclusion with production readiness. Multi-tenant model sharing and the AI Assistant are delivered capabilities; the new VCF Automation GitOps service and vSAN Object Storage are Tech Preview.
  • Monday-morning action: inventory source versions and enabled services, read the 9.1.1 sequencing pre-checks, then validate the upgrade in a representative management domain before exposing any AI assistant or GitOps service to users.

A maintenance release with architectural consequences

VCF 9.1.1 is the first maintenance release for VCF 9.1. It includes cumulative bug fixes and security updates from earlier Express Patches, but Broadcom also uses it to remove adoption friction. That matters because a private-cloud upgrade is rarely blocked by the headline feature; it is blocked by the source version, management footprint, offline content handling, or an unsupported intermediate state.

One important change is support for “back-in-time” upgrade sources. Releases such as VCF 5.2.4, VCF 9.0.2 EP02, and recent vSphere 8.0 Update 3 builds arrived after the original 9.1 baseline and could not follow the expected direct path to 9.1.0. VCF 9.1.1 restores supported paths for those estates. This is less visible than an AI demo, but it may be the feature that actually unlocks a programme.

The VCF Download Tool also becomes more useful in controlled and disconnected environments. A new --latest option filters component binaries to the latest applicable Express Patch, while an artifacts workflow retrieves OCI-based content such as Supervisor updates, VKS versions, Supervisor Services, VCF CLI plugins, VCF services, and Data Services Manager updates. Platform teams can select the Kubernetes releases they approve instead of mirroring everything into an offline depot.

There is one catch: 9.1.1 has explicit sequencing exceptions. Treat the built-in pre-checks as a safety rail, not as a substitute for an upgrade design.

Before updating… Update this first
Other VCF components to 9.1.1 Fleet Lifecycle Management to 9.1.1
Identity Broker or Salt Master/RaaS VCF Management Services to 9.1.1
VCD Migrator VCF Automation to 9.1.1

This is precisely the kind of dependency that disappears in a “patch release” assumption. Freeze the target bill of materials, record the dependency order in the change plan, and keep component-level rollback evidence before starting.

Private AI becomes shared — and operations gains a copilot

The most strategic GA capability is multi-tenant model sharing in VCF Private AI Services. A single Model Runtime can serve models to several tenants or lines of business while their namespaces and private data remain isolated. The economic effect is straightforward: expensive model and GPU capacity no longer has to be duplicated for every team. The governance effect is more important: model lifecycle can remain centralized while consumption stays delegated.

That does not mean “one shared endpoint for everybody.” The architecture still needs explicit tenancy boundaries, quotas, model admission rules, and evidence that prompts, retrieved data, and outputs do not cross namespaces. Shared infrastructure reduces duplication; it increases the importance of isolation tests. The design fits naturally with the control-plane principles described in Private AI on VCF: reference architecture.

VCF Operations also gains an optional AI Assistant for VCF. It is configured locally with a model running through Private AI Services or with a private Google Gemini instance. The assistant can answer natural-language operational questions, correlate health alerts, configuration and logs, help diagnose vSphere and management-service issues, and generate management-pack content for third-party integrations.

The value is not that an LLM “runs the cloud.” The value is compression: an operator can move from several alerts and vMotion errors to a testable diagnostic hypothesis faster. Production adoption still requires a trust model:

  1. define which telemetry and configuration data the model can see;
  2. retain the original evidence beside every generated explanation;
  3. keep remediation behind human approval;
  4. evaluate the assistant on known incidents before using it during an unknown one.

An assistant that produces a plausible answer without traceable evidence can increase MTTR rather than reduce it. Use it as a diagnostic interface, not as an authority.

Kubernetes operations move from polling to evidence

VCF 9.1.1 makes VKS considerably more observable from the infrastructure control plane. VCF Operations can correlate compute, network and storage metrics with Kubernetes events and logs, and the documented real-time path can stream metrics at two-second intervals through OpenTelemetry. This closes a familiar gap: a pod can start, spike, fail, and disappear between traditional five-minute collection windows.

Multi-cluster monitoring uses automated OpenTelemetry collection, while existing Grafana dashboards can be imported alongside native VKS views. That combination lets platform and infrastructure teams inspect the same incident without forcing the application team to abandon its dashboards.

The design tradeoff is data volume. Two-second metrics are not “five-minute metrics, only better”; they multiply ingestion, retention, and cardinality pressure. Start with incident-relevant signals, define retention tiers, and measure the cost before enabling high-frequency collection across every cluster. The goal is to preserve transient evidence, not to collect every label forever.

VKS 3.7 also introduces an Add-on Management Framework. Its practical promise is clearer lifecycle and support ownership for networking, observability, security, and storage add-ons. Add-ons stop being anonymous YAML installed by a project team and become managed platform dependencies with a declared support path.

The Tech Preview is nevertheless architecturally interesting. VCF Automation can install the service, integrate authentication through OIDC, scope operators by region, and attach vSphere Namespaces and VKS clusters as deployment targets. It reduces the glue between infrastructure self-service and application delivery. Evaluate it in a non-production organization with representative identity and multi-region boundaries; do not make it the only production delivery path yet.

Networking, identity, and certificates: the quiet Day-2 wins

EVPN interoperability receives more than a scale bump. Direct VXLAN tunnels between VCF transit gateways and physical leaf switches optimize East-West paths. NAT, NSX Load Balancing, Avi Load Balancing, and DHCP relay become available for EVPN VXLAN distributed connectivity mode. For large multi-tenant fabrics, fewer hairpins mean simpler failure domains and troubleshooting.

VCF 9.1.1 also supports a VLAN-backed distributed transit gateway without a host TEP network. That lowers the entry cost for a VPC design and is compatible with VKS and VCF Automation, but it is not equivalent to a full overlay. Private subnets, SNAT/DNAT/VPN services, and non-IP protocols such as VRRP and multicast are not available in that mode. Use it deliberately for simple connectivity or labs, not as a silent replacement for an overlay design.

Security operations become more dynamic. Active Directory and OpenLDAP group membership can be evaluated at login instead of requiring every user to be pre-provisioned. Password workflows cover more application and management accounts. Certificate lifecycle coverage expands to NSX Edges, vSphere Supervisors, License Servers, cloud proxies, and network collectors, with expiry tracking, renewal-failure alerts, third-party CA integration, and support for non-TLS certificates.

Finally, VMware Salt for VCF Component APIs exposes more than 300 configuration settings using Salt constructs. The opportunity is fleet-wide desired state. The danger is fleet-wide blast radius. Start read-only, version the states, test on a canary domain, and keep drift detection separate from automatic correction until the baseline is trustworthy.

Smaller management footprint, easier labs, fewer excuses

VCF Operations in 9.1.1 introduces a compact form factor advertised with up to 40% lower CPU and memory requirements and a two-node high-availability setup. VCF Management Services are right-sized for Day-0, and a new Small HA option avoids choosing between the smallest footprint and control-plane availability.

Existing 9.1 estates do not shrink automatically after the upgrade. Their current sizing is preserved; adopting the reduced footprint requires the documented post-upgrade right-sizing process after all VCF Management Services components reach 9.1.1. Capacity reclaimed on paper is not capacity reclaimed in the cluster until that step is planned and verified.

The installer also exposes workflows that previously required an API or an override: HTTP and custom URLs for offline depots, single-host deployments, and non-HCL NVMe devices for vSAN ESA. These are meaningful improvements for labs and proofs of concept. They are not a relaxation of production support: production storage devices must still be on the Broadcom Compatibility Guide, and production designs must still meet the documented minimum host configuration.

Upgrade now, evaluate, or wait?

Your situation Recommendation
Running VCF 9.1 with a normal maintenance window Plan 9.1.1 after validating sequencing and cumulative fixes; the operations improvements justify the maintenance cycle.
Blocked on a back-in-time source version Reassess the programme now; 9.1.1 may provide the supported path that 9.1.0 lacked.
Building a new compact or disconnected environment Prefer 9.1.1 for Small HA, offline-depot UI improvements, VCFDT artifacts, and the reduced footprint.
Primarily interested in native GitOps or vSAN Object Storage Evaluate in a lab; both are Tech Preview and should not carry a production dependency.
Planning the AI Assistant Pilot with controlled data access, human-approved actions, and a benchmark of known incidents before broad enablement.

A practical 90-day adoption sequence

The safest way to consume 9.1.1 is to separate the platform upgrade from the activation of its new services. Combining both in one change window makes every incident ambiguous: is the problem caused by the new component version, the new telemetry path, the model integration, or the new operating policy? A phased plan keeps the evidence readable.

Days 0–15 — establish the upgrade envelope. Export the current bill of materials, map every source version, list optional VCF services, and identify offline-depot constraints. Run the official pre-checks, but add your own functional baselines: authentication, certificate expiry, VKS health, vMotion, workload-domain lifecycle, and recovery jobs. These become the before-and-after evidence. Validate the component order and rehearse the rollback decision points rather than merely documenting a global rollback statement.

Days 15–35 — upgrade the platform, change nothing else. Patch Fleet Lifecycle Management first, follow the VCFMS and VCF Automation dependencies, then observe the estate through at least one normal operational cycle. Confirm that existing automations, identity flows, monitoring integrations, and support bundles still behave as expected. If you want the smaller VCFMS footprint, treat right-sizing as a separate change after stability has been demonstrated.

Days 35–60 — activate observability and fleet controls. Onboard a small set of representative VKS clusters to the two-second OpenTelemetry path. Measure ingestion volume, label cardinality, retention, and operator usefulness before expanding. Bring additional certificate and password workflows into VCF Operations in observation mode first. For Salt APIs, compare desired state with reality before enabling correction.

Days 60–90 — expose new consumption services. Pilot multi-tenant model sharing with two teams whose data boundaries are easy to verify. Evaluate the AI Assistant against a catalogue of resolved incidents and require links back to source telemetry. Test the GitOps Tech Preview in a non-production organization with real OIDC groups and namespace boundaries. At day 90, promote only the capabilities for which you have both technical evidence and a named operational owner.

This sequence deliberately optimizes for learning, not feature velocity. It creates a clean answer to the question every change board will ask after an incident: “what changed, and how do we know?”

Pitfalls & things to watch

Four other traps deserve a line in the runbook:

  • Upgrade order: patch Fleet LCM first, then respect the VCFMS and VCF Automation dependencies. A green generic health check does not erase those constraints.
  • AI data boundaries: locally hosted does not automatically mean least privilege. Review prompt, log, configuration, and retrieved-data flows.
  • Telemetry economics: two-second VKS metrics require a cardinality and retention budget.
  • Post-upgrade sizing: the smaller VCFMS footprint is opt-in for an existing 9.1 deployment; validate headroom before and after right-sizing.

Conclusion

VCF 9.1.1 is best understood as the release that makes the VCF 9.1 architecture more operable. It opens blocked upgrade paths, gives platform teams better control of offline artifacts and Kubernetes add-ons, makes transient VKS problems visible, extends fleet identity and certificate workflows, and turns private AI into a shared service rather than a collection of isolated model stacks.

The release is worth planning now, but not consuming blindly. Separate GA from Tech Preview, preserve evidence around the AI Assistant, size the telemetry pipeline, and treat the exceptional component sequence as part of the design. The strongest 9.1.1 deployment is not the one with every toggle enabled; it is the one with a clear production boundary.

Upgrade friction falls

Back-in-time paths, better offline artifact handling, and a smaller management footprint make 9.1 adoption more practical.

Day-2 signal improves

AI-assisted diagnostics, two-second VKS telemetry, and broader credential and certificate workflows reduce operational blind spots.

Boundaries still matter

Tech Preview services, AI trust, telemetry cost, and upgrade ordering belong in the architecture—not in release-day improvisation.

Official and technical sources

Get the next one by email

New articles and series, sent when they are published. No other mail.

One click to unsubscribe, any time.

Back to blog
Share

Related articles

  1. 16 min read

    FinOps cost models: what AWS, Azure and GCP bill — and what VCF calculates

    An EKS cluster-hour, an AKS tier, a GKE Pod request and depreciated VCF hardware are not four values of one variable. What each platform bills, and what VCF calculates instead.

  2. 23 min read

    Network policies and Cilium: building a defensible default-deny

    The NetworkPolicy API ships with Kubernetes; enforcing it is the CNI's job. What Cilium adds, what stays standard, and how to reach default-deny by watching real flows before blocking any.

  3. 18 min read

    Kubernetes RBAC: the foundations, and the pitfalls that survive an audit

    Every one of these pitfalls is published on kubernetes.io. What is missing is the ordering — and the path that leads from a vSphere Namespace straight to cluster-admin.

Follow along

New articles, thoughts, and updates.