Nvidia’s Vera CPU represents a notable shift in the CPU landscape, directly challenging the dominance of Intel and AMD. This new chip is designed to cater to hyperscalers and cloud service providers, with companies like Alibaba, Meta, and Oracle already committed to deploying it.
Vera is built on the Armv9.2 architecture, featuring 88 custom cores that support 176 threads and can handle up to 1.5 TB of LPDDR5X memory. Importantly, it operates as a standalone platform, independent of Nvidia’s GPUs, which is a significant move for the company.
Architecture Overview
The Vera CPU’s architecture is designed to minimize execution bottlenecks, making it suitable for two primary workloads: managing GPUs in Nvidia’s Vera Rubin systems and hosting AI agents that do not rely on GPUs. The chip consists of a monolithic compute die that houses all 88 cores, fabricated using TSMC’s 3nm process.
Surrounding the compute die are various chiplets, including eight LPDDR5X memory controllers and two I/O dies for PCIe 6.4 and CXL 3.1 connectivity, as well as a dedicated NVLink interface. This design is somewhat similar to Amazon’s Graviton 4 CPUs, which also utilize a monolithic compute die.
Superchip Configuration
Nvidia’s Vera CPU can be configured in both single-socket and dual-socket setups, with the latter referred to as the Vera CPU Superchip. This configuration allows for a total of 176 cores and 352 threads, connected via NVLink at a bandwidth of 1.8 TB/s. The Superchip can support up to 3 TB of LPDDR5X memory, providing an impressive aggregate memory bandwidth of 2.4 TB/s.
Custom Olympus Cores
Vera’s cores, named Olympus, are Nvidia’s first fully custom CPU cores, moving away from previously used off-the-shelf Arm designs. Olympus features a 10-wide decoder, eight integer ALUs, and six vector/FP pipelines, along with a custom neural branch predictor that enhances performance by predicting two execution paths simultaneously.
Additionally, Nvidia has implemented several optimizations to reduce pipeline stalls and improve instruction throughput. The architecture supports simultaneous multithreading (SMT), branded as spatial multithreading, which allows for more efficient resource utilization by enabling two threads to operate on a single core.
Vera’s memory subsystem is equally advanced, featuring eight LPDDR5X memory controllers capable of supporting data rates up to 9600 MT/s. This results in a memory bandwidth of 1.2 TB/s per chip, positioning Vera among the fastest memory subsystems available today.
This article was produced by NeonPulse.today using human and AI-assisted editorial processes, based on publicly available information. Content may be edited for clarity and style.








