HPE ProLiant DL394 Gen12 Puts NVIDIA Vera CPU at the Center of Agentic AI
NVIDIA's 88-core Vera CPU pairs with GPU clusters to tackle agentic AI orchestration at scale

Hewlett Packard Enterprise's ProLiant Compute DL394 Gen12, the first enterprise server built around NVIDIA's Vera CPU, is entering general availability this fall — arriving as evidence accumulates that the CPU, not the GPU, is the component most likely to throttle large-scale agentic AI deployments. The server pairs NVIDIA's 88-core Armv9.2 processor with HPE's existing GPU server lineup, targeting the orchestration layer that chatbot-era infrastructure was never designed to serve efficiently. The New York Stock Exchange, which processes 1.1 trillion messages per day, is exploring the DL394 Gen12 alongside Redpanda streaming and NVIDIA's AI stack for AI-ready market infrastructure — an early signal that the architecture is being taken seriously in environments where latency failures are financially consequential.
Why Agentic AI Runs Out of CPU Before It Runs Out of GPU
The bottleneck the DL394 Gen12 addresses is structural. In chatbot-era AI — large models responding to discrete queries in high-volume batches — GPUs dominate execution time, and one CPU per four to eight GPUs is a reasonable configuration. Agentic AI changes the execution profile fundamentally.
When an AI agent receives a task, it does not run one inference call. It runs inference, parses the output, routes to tools (web search, code execution, database queries), executes those tools, manages state across sub-agents, assembles context, and calls the model again. The GPU operates intensively during inference but waits during all surrounding steps. Those steps — the orchestration layer — are CPU-bound and scale with agent fleet size.
Market research firm TrendForce projects that CPU-to-GPU ratios in dense agentic deployments are shifting from the historical 1:4–1:8 range toward 1:1 or 1.2:1. That shift has direct capital expenditure consequences: infrastructure configured for chatbot inference is structurally under-provisioned for agent orchestration, and the gap cannot be closed through software optimization alone.
Independent research published in August 2026, focused on agentic AI resource dynamics, has found that CPU provisioning is now a first-class infrastructure decision in agent deployments — with properly matched CPU allocation capable of dramatically reducing latency for CPU-sensitive agent tasks compared with standard configurations.
Read more: NVIDIA Vera Rubin NVL72 Makes MLPerf Debut With 3.7x Throughput Lead Over Blackwell
NVIDIA Vera's Monolithic Architecture and the NUMA Difference
The NVIDIA Vera CPU uses NVIDIA's own "Olympus" core architecture, an Armv9.2 implementation with 88 physical cores built as a monolithic die rather than the chiplet assemblies used in AMD EPYC and Intel Xeon Scalable. The monolithic design matters for the specific work agents perform.
Chiplet-based processors introduce NUMA (non-uniform memory access) topology — some cores have longer paths to some memory regions. For GPU workloads like matrix multiplication, NUMA overhead is largely irrelevant. For the latency-sensitive state management and context assembly that agents execute across all cores simultaneously, NUMA latency compounds under concurrency. According to HPE's press release, Vera's single-die architecture gives all 88 cores uniform access to up to 1.5TB of LPDDR5X memory at 1.2 terabytes per second of bandwidth — enabling up to 14 gigabytes per second per core.
In the Vera Rubin platform, NVLink-C2C delivers 1.8 terabytes per second of coherent CPU-GPU interconnect bandwidth, enabling the CPU and GPU to share a unified memory address space without copying data across PCIe — reducing the serialization overhead between agent orchestration and model inference calls.
DeepInfra, a production AI inference platform that received early hardware access through NVIDIA's open AI ecosystem, published benchmark results on July 21, 2026, showing Vera delivering up to 2.2 times faster orchestration than x86 baselines and supporting up to 1.6 times more concurrent AI agents at the same quality of service. DeepInfra tested Vera against AMD Zen5 (Turin), Intel Granite Rapids, and Intel Sapphire Rapids using its own production AI agent and captured traffic, and reported Vera as the fastest of all four architectures tested in every workload category.
Phoronix, in early benchmarks on NVIDIA-provided pre-production hardware published May 26, 2026, described the Vera CPU as delivering the most competitive Arm server performance Phoronix had ever tested, with performance in several workloads competitive with current-generation AMD EPYC 9005 Turin processors despite a 40-core count disadvantage. A kernel compilation test saw 88-core Vera outperform NVIDIA's own 128-core Intel Xeon 6980P comparison. NVIDIA's own performance claims should be evaluated alongside these third-party results as the platform matures in production settings.
HPE's AI Factory Integration and the Enterprise Deployment Picture
HPE positions the DL394 Gen12 as the CPU layer of its AI Factory stack, which also includes the ProLiant DL380a and XD685 for GPU-dense compute and the XD700 for high-GPU-density racks, managed through HPE Compute Ops Management. The DL394 was added to HPE Private Cloud AI at HPE Discover 2026 in Las Vegas in June, enabling balanced CPU-GPU deployments under unified management.
The DL394 Gen12 also embeds security at every level. According to HPE, the server is among the first to meet NIST's quantum computing-resistant security requirements, featuring Silicon Root of Trust and HPE's iLO-enabled secure enclave — which the company says safeguards servers at every phase of their lifecycle. HPE Compute Ops Management provides AI-driven operations designed to reduce server management time.
HPE's Q3 FY2026 financial results, reported September 2, 2026, indicate substantial committed demand across its AI portfolio. The company posted record revenue of $12.2 billion (up 34% year-over-year), with server revenue of $6.8 billion (up 35.3%) within the Cloud & AI segment. AI systems orders reached $2.4 billion in the quarter, up more than 30% sequentially, and the AI systems backlog grew to $6.8 billion by quarter-end. Those figures reflect the full AI Factory portfolio; the DL394 Gen12 is the newest component of a lineup already generating significant revenue.
Competitive Landscape: Dell, Lenovo, and Supermicro Are Also Building Vera Servers
HPE is first to market with the DL394 Gen12, but that advantage is narrow and temporary. Dell, Lenovo, and Supermicro have each announced enterprise server platforms built around NVIDIA Vera CPU, creating a multi-OEM category rather than an HPE-exclusive product. HPE's defensible advantages are integration depth — the AI Factory software stack, iLO security with Silicon Root of Trust and quantum-resistant firmware, and enterprise support infrastructure — rather than exclusive hardware access.
The broader competition is AMD's response. AMD announced its 6th-generation EPYC processors with explicit agent orchestration positioning, arguing that modern x86 with optimized memory subsystems can match Vera's throughput while preserving compatibility with the enterprise x86 software stack. The x86 versus Arm question in enterprise AI orchestration has not yet been resolved by independent production evidence at scale.
Read more: CoreWeave Scales Vera Rubin to Multi-Rack Clusters as MLPerf v6.1 Confirms 3.7x Gain
Limitations: New Silicon, High Thermal Profile, and Limited Production History
The DL394 Gen12 represents a first-generation enterprise deployment for NVIDIA Vera CPU. The existing third-party benchmark evidence comes from pre-production hardware and a platform granted early access by NVIDIA — useful data points, but not yet sufficient to characterize performance across every agent workload type and configuration under true independent conditions. Vera's 450-watt TDP is higher than most comparable server CPUs, with data center cooling and power density implications. The LPDDR5X SOCAMM2 memory format is less broadly supported in enterprise supply chains than RDIMM DDR5, which may complicate serviceability logistics at scale. These are knowable constraints that infrastructure architects should weigh against the performance differential.
The Infrastructure Refresh That Agentic AI Is Forcing
The DL394 Gen12's significance exceeds any single benchmark. The working assumption of most enterprise AI infrastructure built between 2022 and 2025 — that heavily GPU-weighted configurations are sufficient — is being revised as agent deployments mature. If the CPU:GPU ratio shift that TrendForce projects proves accurate at scale, data centers optimized for chatbot inference will require structural refresh, not software patches. HPE, Dell, Lenovo, Supermicro, and AMD are each positioning for that refresh cycle. The outcome will depend less on any single server's benchmarks than on which vendor can deliver the integrated portfolio, management tooling, and enterprise support that large-scale agentic deployments demand. The DL394 Gen12 is HPE's opening argument that the right answer to agentic AI is not simply more GPUs — it is the right CPU paired with the right GPU, managed as a coherent system at production scale.