Nvidia’s Vera CPU: A Deep Dive into the Olympus Cores

Nvidia's Vera CPU marks a significant entry into the CPU market, featuring 88 custom cores and advanced memory capabilities aimed at cloud providers and AI workloads.

Nvidia’s Vera CPU represents a notable shift in the CPU landscape, directly challenging the dominance of Intel and AMD. This new chip is designed to cater to hyperscalers and cloud service providers, with companies like Alibaba, Meta, and Oracle already committed to deploying it.

Vera is built on the Armv9.2 architecture, featuring 88 custom cores that support 176 threads and can handle up to 1.5 TB of LPDDR5X memory. Importantly, it operates as a standalone platform, independent of Nvidia’s GPUs, which is a significant move for the company.

Architecture Overview

The Vera CPU’s architecture is designed to minimize execution bottlenecks, making it suitable for two primary workloads: managing GPUs in Nvidia’s Vera Rubin systems and hosting AI agents that do not rely on GPUs. The chip consists of a monolithic compute die that houses all 88 cores, fabricated using TSMC’s 3nm process.

Surrounding the compute die are various chiplets, including eight LPDDR5X memory controllers and two I/O dies for PCIe 6.4 and CXL 3.1 connectivity, as well as a dedicated NVLink interface. This design is somewhat similar to Amazon’s Graviton 4 CPUs, which also utilize a monolithic compute die.

Superchip Configuration

Nvidia’s Vera CPU can be configured in both single-socket and dual-socket setups, with the latter referred to as the Vera CPU Superchip. This configuration allows for a total of 176 cores and 352 threads, connected via NVLink at a bandwidth of 1.8 TB/s. The Superchip can support up to 3 TB of LPDDR5X memory, providing an impressive aggregate memory bandwidth of 2.4 TB/s.

Custom Olympus Cores

Vera’s cores, named Olympus, are Nvidia’s first fully custom CPU cores, moving away from previously used off-the-shelf Arm designs. Olympus features a 10-wide decoder, eight integer ALUs, and six vector/FP pipelines, along with a custom neural branch predictor that enhances performance by predicting two execution paths simultaneously.

Additionally, Nvidia has implemented several optimizations to reduce pipeline stalls and improve instruction throughput. The architecture supports simultaneous multithreading (SMT), branded as spatial multithreading, which allows for more efficient resource utilization by enabling two threads to operate on a single core.

Vera’s memory subsystem is equally advanced, featuring eight LPDDR5X memory controllers capable of supporting data rates up to 9600 MT/s. This results in a memory bandwidth of 1.2 TB/s per chip, positioning Vera among the fastest memory subsystems available today.

This article was produced by NeonPulse.today using human and AI-assisted editorial processes, based on publicly available information. Content may be edited for clarity and style.

Avatar photo
GEAR-5

A meticulous tech analyst obsessed with silicon, circuitry, and impossible benchmarks. GEAR-5 tracks every hardware and gadget launch like a sacred ritual. His geek-level curiosity is as sharp as his thick-framed glasses, and his mission is simple: dissect every device from the future to reveal what’s truly worth it — and what’s just marketing smoke.

Articles: 769