Blogs
/
NVIDIA Vera CPUs come to Nscale
Company

NVIDIA Vera CPUs come to Nscale

September 30, 2026
•
3 mins read

Matt Vegas

NVIDIA Vera CPUs on Nscale will deliver extreme CPU performance, memory bandwidth, and efficiency needed to power the next generation of agentic AI workloads.

For the past several years, AI infrastructure has been optimized to train more capable models and run them faster and more efficiently. Agentic AI disrupts this design; a coding agent does not necessarily finish when it generates an answer. It may inspect a repository, launch an isolated environment, write and compile code, install dependencies, run tests, analyze failures, and try again. Other agents search data, call APIs, execute workflows, and evaluate outcomes before deciding on their next step.

The GPU provides the intelligence, but the wider system has to turn that intelligence into action. 

Increasingly, that work is run on the CPU, which executes the actions that connect one model step to the next. A joint study from the Georgia Institute of Technology and Intel found that tool-dominated agentic AI workloads are significantly bottlenecked by tool processing on the CPU, consuming up to 88% of end-to-end latency. A June 2026 study found that active users of OpenAI Codex grew more than fivefold during the first half of 2026. Every agentic action creates CPU work that didn't previously exist. As models become more capable, infrastructure demand shifts from pure inference to orchestrating millions of real-world interactions. As agentic AI scales, infrastructure must be designed for execution as well as intelligence.

NVIDIA Vera CPUs are coming to Nscale’s full-stack AI cloud platform and are now available to order, giving AI-native companies, AI labs, and enterprises access to CPU infrastructure purpose-built for agentic AI, reinforcement learning, and sandboxed execution. Integrated with accelerated compute, high-speed networking, orchestration, and AI-native operations, Vera helps customers run the execution-heavy workloads surrounding AI models as a single production-ready system. 

The critical role of CPU execution in agentic performance

AI agent performance depends on the entire reasoning-and-execution loop. While GPUs generate actions, CPUs handle the tools, environments, and workflows necessary to complete them. As agentic workloads scale, factors like CPU performance, memory bandwidth, and sandbox density become critical determinants of latency, system utilization, and overall infrastructure economics.

If the CPU layer lags, the entire loop slows; GPUs sit idle while waiting for code compilation, environment provisioning, or testing. Consequently, the key performance metric shifts from tokens per second to the volume of completed tasks per rack, watt, and dollar. Ensuring CPUs can keep pace with GPU inference is essential to maximizing both agent speed and infrastructure efficiency.

“Vera gives AI agents the CPU performance to act; Nscale turns that into a production-ready system. By integrating it with accelerated compute, orchestration and AI-native operations, we help customers keep GPUs productive and scale agentic workloads without assembling the infrastructure layer by layer.”
Nidhi Chappell President of AI Infrastructure, Nscale

Built for the work AI agents actually do

NVIDIA Vera is designed for the execution layer around the model. It combines 88 custom NVIDIA Olympus cores with high-bandwidth LPDDR5X memory and a monolithic architecture built to sustain fast, predictable per-core performance under load. In NVIDIA testing of agentic workloads running under full load, Vera delivered up to 1.8 times the sandbox performance of a latest-generation x86 CPU.

Access to a faster CPU is only part of the answer. At Nscale, we design and operate the full AI tech stack, from ground to cloud, so customers will be able to deploy CPU and GPU capacity as one coordinated system, without stitching together infrastructure, software, and operations across multiple providers.

The applications are already taking shape across the AI lifecycle. Coding agents use CPUs to clone repositories, install dependencies, compile code, and run tests. Reinforcement-learning environments run large numbers of isolated sandboxes to generate evaluations and training signals. Browser agents need secure environments in which to navigate sites and execute tools, while enterprise workflow agents call APIs, process documents, and coordinate actions across business systems. These workloads share the same infrastructure requirement: fast CPU execution repeated across many parallel environments, without leaving expensive GPUs waiting for the next result.

Nscale is built to unlock the full value of agentic AI 

That requirement becomes more significant as AI companies increase the number of agents and training rollouts operating in parallel. Reinforcement-learning systems may run thousands of isolated environments, with each environment executing actions, evaluating outcomes, and returning signals to the model. Faster CPU turnaround reduces the delay between action and feedback, helping training systems to keep data fresh and to complete more useful learning cycles in the same period.

This is where Nscale's full-stack model becomes a material advantage for agentic AI. Agents depend on more than model inference: they require compute, networking, scheduling, isolation, and fleet operations working as one system throughout the agentic reasoning and execution loop.  We design and build data centers around cluster performance, not just infrastructure capacity. That means optimizing the physical environment, power, network, and operating model for the demanding, tightly integrated compute required by next-generation AI workloads. Our validation and burn-in process enables us to deliver a fully hardened, production-ready cluster—helping customers move from deployment to useful capacity faster and with greater confidence.

Nscale validates and burns in infrastructure before it reaches production, while automated monitoring and self-healing capabilities with fleet operations help detect and recover from issues once workloads are running. For agentic systems operating across thousands of parallel environments, that means less time lost to infrastructure disruption, higher utilisation and more predictable performance at scale. Rather than treating execution and model computation as separate systems, customers will be able to build a single AI platform where CPUs keep agents moving and GPUs remain focused on the parallel computation they perform best.

Nscale's planned acquisition of Anyscale adds orchestration to this stack. Anyscale, the company behind the open-source framework Ray, builds the platform teams use for scheduling, task distribution, and failure handling across thousands of parallel environments. Paired with Vera, that coordination can draw on the CPU performance to match its needs at scale. Above the infrastructure, Nscale Cloud delivers the full AI stack under one contract, one identity layer, and one governance model. Customers can consume the infrastructure, cloud platform, and AI services as an integrated experience rather than stitching together multiple providers.

Nscale designs these layers to operate as one system, meaning customers will be able to scale from individual CPU environments to dense rack-level deployments while maintaining the visibility, support, and operational control required to turn agentic AI into production AI.  Together, NVIDIA Vera-powered infrastructure, Nscale’s optimized delivery and operations, and Anyscale provide a path from high-performance compute to production-grade agentic AI.

“Agentic AI is creating a new CPU moment, with tool use, code execution, and sandboxed environments placing new demands on AI infrastructure. NVIDIA and Nscale are working together to bring the performance of NVIDIA Vera to Nscale’s full-stack AI cloud platform, helping AI natives and enterprises build and scale the next generation of AI agents.”
Ian Buck Vice President and General Manager, Hyperscale and HPC, NVIDIA

From tokens generated to tasks completed

The first phase of the AI infrastructure race was defined by access to GPUs. The next will be defined by how effectively organizations turn those GPUs into useful and affordable AI work at scale. For agentic AI, intelligence is not the finished product. The system has to apply that intelligence: run the code, call the tool, search the data, test the answer, and return the result.

This is why CPU access is becoming strategically important. It enables customers to build the execution environments that sit alongside accelerated compute rather than treating them as an afterthought or an external bottleneck. Nscale is building the infrastructure and services around the NVIDIA Vera CPU as a complete AI system: from the energy powering our data centers and accelerated compute through networking, fleet operations, cloud infrastructure and distributed AI orchestration. Customers will get Vera as part of a production-ready environment designed around the demands of agentic workloads, rather than having to assemble and operate each layer independently.

Interested in learning more? Speak to Nscale to place an order or discuss deployment timelines.

‍