Nutanix announced the general availability of Nutanix Enterprise AI (NAI) 2.8, the latest release of its software platform for deploying and governing LLMs and AI agents across hybrid cloud environments.
Among the new features, the release adds:
- MCP gateway within Nutanix’s Agent Gateway
- Consumption controls that track and cap token spending at the agent, user, and team levels
- Inference optimizations to reduce the cost of running private models in production.
Nutanix also announced the general availability of Service Provider Central, a multitenant control plane for service providers, and previewed Nutanix Kubernetes Platform (NKP) 2.19, which is expected in a subsequent release.
Overall, the release addresses a problem that has emerged as enterprises move AI agents from pilot projects into production. Autonomous agents can consume tokens, call external tools, and access enterprise data without the oversight structures that already govern human users and applications.
The announcement also extends Nutanix’s dual-native architecture, the company’s term for running virtual machines, containers, and AI workloads on shared infrastructure with a single governance layer.
Technical Details
NAI 2.8 groups its new capabilities into two areas, agent governance via Agent Gateway and model efficiency via Nutanix Private Inference, along with identity and access management changes that apply across both AI and non-AI workloads on the Nutanix Cloud Platform.
Agent Gateway and MCP Governance
The centerpiece of the release is the general availability of an MCP gateway built into Nutanix’s existing Agent Gateway service, which the company describes as a single front door connecting AI agents to the tools and data they can access.
Those capabilities include:
- Centralized management of both locally deployed and remote MCP servers, replacing a set of point integrations between individual agents and enterprise tools
- Fine-grained permission controls that distinguish read-only access from write access for a specific tool
- Zero-downtime rolling updates for tool deployments
- An accompanying MCP Server for Nutanix Cloud Platform that provides agents with access to infrastructure functions within defined security boundaries
- Activity logs record which data sources and tools each agent accessed; prompt content inspection is not part of the gateway’s default logging
Alongside the MCP gateway, NAI 2.8 adds consumption controls that apply at the individual agent level and aggregate to user and team totals, designed to catch the kind of runaway token spending Cornely described.
Header-based rate limiting (technical preview) allows administrators to set token budgets for individual users without modifying existing identity and user management systems.
Identity and access management gets a Custom Role Builder that lets administrators build roles from a catalog of more than 30 permissions covering user management, licensing, models, endpoints, API keys, and observability. Administrators can also start with a copy of a predefined role and adjust it to fit their organization.
Nutanix Private Inference
On the model-serving side, Nutanix Private Inference adds several capabilities for economically running smaller, self-hosted models:
- Parameter-efficient fine-tuning using Low-Rank Adaptation (LoRA) targets models with fewer than 8 billion parameters to match the accuracy of larger models on domain-specific tasks while keeping compute and memory costs low.
- Tensor parallelism enables multi-GPU serving of a single model, and a separate technical preview extends this capability to pipeline parallelism across multiple GPU nodes for multi-node inference.
- Batch inference and speculative decoding round out the set; Nutanix’s announcement said that speculative decoding, which uses lightweight draft models to predict tokens ahead of the full model, can increase token-generation speed by up to 2.5x.
- KV cache offloading (a separate technical preview) moves key-value cache context from GPU to CPU memory to reduce redundant recomputation and shorten time to first token.
Nutanix also extended support for air-gapped NVIDIA NIM deployments, enabling customers with NVIDIA AI Enterprise licenses to run inference in disconnected, highly regulated environments.
Analysis
NAI 2.8 reinforces Nutanix’s dual-native pitch, the argument that customers can run virtual machines, containers, and now AI agents on a single infrastructure and licensing footprint, rather than setting up a separate AI-specific environment.
For platform and AI engineering teams, the practical value of NAI 2.8 lies in providing visibility into agent behavior that most organizations currently lack:
- Token- and agent-level consumption tracking enables finance and platform teams to attribute AI spend to specific agents, users, or teams
- The MCP gateway centralizes what had been a piecemeal set of point integrations between agents and enterprise tools, reducing the number of places a security team must audit for excessive permissions.
Competitive Landscape
Nutanix competes for enterprise AI infrastructure spending against a robust set of vendors (some of which are also partners) that make comparable dual-stack or private-AI-factory claims, each with varying depth of agent governance.
The table below summarizes how the leading alternatives compare to NAI 2.8.
| Alternative | Model/Approach | Relative to NAI 2.8 |
| VMware Private AI Foundation (Broadcom, VCF 9.1) | Private AI layer built into VMware Cloud Foundation, paired with NVIDIA and partner model catalogs | Comparable virtualization-based private AI deployment and a larger existing installed base, but no generally available MCP-based agent governance layer equivalent to Agent Gateway |
| Dell AI Factory with NVIDIA | Validated hardware and software stack combining Dell infrastructure, NVIDIA AI Enterprise and partner software | Greater hardware breadth and reference-architecture depth, with agent governance handled more by NVIDIA’s own tooling than by a Dell-native layer |
| HPE Private Cloud AI | Turnkey private AI infrastructure combining HPE servers, storage and NVIDIA software under one support model | Similar turnkey positioning for regulated buyers, with less public emphasis to date on token-level consumption controls and custom IAM roles for agents |
| Red Hat OpenShift AI | Kubernetes-native MLOps and model-serving platform integrated with OpenShift | More mature open-source MLOps tooling and broader Kubernetes ecosystem support, without Nutanix’s dual-native VM-and-container governance model or HCI storage integration |
Differentiation is strongest where Nutanix combines its existing dual-native virtualization and Kubernetes foundation with agent-specific governance into a single licensed platform, a level of integration that competitors are still assembling.
It is weakest in model tooling and MLOps breadth, where Red Hat’s OpenShift AI ecosystem and NVIDIA’s agent frameworks, available across Dell and HPE stacks alike, have a longer track record.
Final Thoughts
NAI 2.8 addresses a real and increasingly common problem, namely agentic AI deployments that have outpaced the governance and cost controls organizations apply to other production systems. The general availability of the MCP gateway, paired with agent-level consumption tracking, gives Nutanix customers tools that were previously assembled from point solutions or left unmanaged.
For enterprises already running Nutanix for virtualization, NAI 2.8 provides IT organizations with a credible path to bring agentic AI under the same governance and cost discipline applied to the rest of their infrastructure, without waiting for a separate AI platform initiative to mature elsewhere in the stack.
For those looking to migrate from competing solutions, Nutanix just gave you another reason to switch.



