NVIDIA says Vera Rubin NVL72 can deliver up to 30x more agentic AI throughput per megawatt than GB300, targeting long-context inference costs.

NVIDIA says its upcoming Vera Rubin NVL72 platform can deliver up to 30 times more agentic AI throughput per megawatt than the company’s GB300 NVL72 systems, offering a potential efficiency gain for data centers running increasingly long, tool-using AI workflows.
The claim comes from early NVIDIA measurements using the SemiAnalysis AgentX workload, which preserves recorded coding sessions, expanding context, tool calls and sub-agent activity. NVIDIA says the results are especially relevant as AI agents move beyond single-turn chat and begin performing multi-step research, coding, customer-service and decision-support tasks.
The figures are vendor-reported and have not yet been independently validated. NVIDIA says the Vera Rubin results are pending review by SemiAnalysis, and that the tests do not include Vera CPU performance for tool calling.
NVIDIA’s comparison focuses on throughput per megawatt rather than peak tokens per second from an isolated inference request. That distinction matters for agentic systems, where a model may repeatedly read accumulated context, call external tools, delegate subtasks and synthesize the results.
According to NVIDIA, AgentX represents those behaviors through real-world agentic coding trajectories. The company says the benchmark includes models such as Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro, allowing performance to be assessed across different model architectures and workloads.
On DeepSeek V4 Pro, NVIDIA reports that GB300 NVL72 delivers as much as 15 times the throughput per megawatt of its Hopper architecture. Vera Rubin is then claimed to raise that figure by up to another 30 times relative to GB300 NVL72. These are maximum results on the cited workload and model, not a guarantee that every deployment will see the same improvement.
NVIDIA also reports up to 35 times lower cost per million tokens compared with GB300 NVL72. The company links that measure to the economics of AI factories, where electricity capacity limits how many GPUs can operate and token cost affects the margin on inference services.
A conventional chat request may involve a relatively short input and output. NVIDIA describes typical sequences for chat or document summarization as ranging from roughly 1,000 to 8,000 tokens. Agent sessions can accumulate hundreds of thousands of input tokens as previous steps, tool results and sub-agent outputs remain available to later stages.
That creates a different systems problem. The infrastructure must process large amounts of context while maintaining response speed across many irregular requests. A workload can also alternate between context processing, known as prefill, and response generation, or decode. Treating both stages identically can leave some GPUs underused while others become bottlenecks.
NVIDIA’s answer combines hardware and software techniques. Disaggregated serving allows prefill and decode resources to scale separately, while rate matching attempts to keep the two stages synchronized. Distributed KV-cache systems keep previously processed context available across the GPU domain, and KV-aware routing can send a request to hardware that already holds relevant cached data.
The platform also relies on large-scale expert parallelism for mixture-of-experts models, CUDA kernels designed to combine computation with communication, and NVFP4 quantization to reduce the memory required for model weights. NVIDIA says fifth-generation Tensor Cores and its Transformer Engine support both major phases of inference.
The strongest performance claims in this story come from NVIDIA’s own infrastructure and developer blogs, rather than an independent benchmark report. The company attributes the workload design to SemiAnalysis AgentX, but says the Vera Rubin results are still awaiting SemiAnalysis review. That makes the figures useful as an early indication of NVIDIA’s target performance, not as a settled industry comparison.
The comparison is also limited in several ways. NVIDIA reports maximum gains, so the results may depend on model choice, context length, request mix, batching, software versions and power-management settings. The cited Vera Rubin measurements omit CPU work for tool calling, an important part of production agents. Actual applications may also spend substantial time waiting on databases, APIs or enterprise systems rather than generating tokens.
NVIDIA says its DSX MaxLPS technology can manage power at GPU, rack and workload levels and provision up to 40% more GPUs within the same megawatt budget. That is another company claim, and its practical value will depend on cooling, rack design, utilization and the behavior of deployed workloads.
The architecture’s economics are tied to NVIDIA’s broader software stack, including TensorRT LLM and NVIDIA Dynamo. The company argues that the NVL72 scale-up domain and sixth-generation NVLink technologies provide the communication bandwidth needed for distributed caching and expert parallelism. Those design choices could improve performance, but they also reinforce dependence on NVIDIA’s hardware, networking and software ecosystem.
For AI application teams, the announcement reinforces that agent cost cannot be evaluated only by measuring the price of a single model response. A research agent that repeatedly reuses a growing context may generate far more inference work than a chat assistant, even when the user sees only one final answer.
Infrastructure teams will therefore need to track metrics such as cost per completed task, energy per workflow and throughput under realistic context growth. A platform that performs well on short prompts may not be efficient when agents invoke tools, spawn sub-agents or revisit long histories.
For enterprise buyers, the immediate question is not simply whether Vera Rubin is faster than GB300. It is whether an organization has enough sustained agent traffic to justify a new platform, and whether its applications can exploit features such as caching, prefill-decode separation and model quantization. Teams with modest or unpredictable workloads may still favor shared cloud inference, while large AI service providers and power-constrained data centers have stronger incentives to optimize work per megawatt.
The announcement also raises a software portability issue. NVIDIA’s reported gains depend on codesign across GPUs, NVLink, CUDA, TensorRT LLM and NVIDIA Dynamo. Builders using that stack may gain access to more optimization, but they may face additional migration work if they later move models or serving workloads to other accelerator platforms.
The first signal will be the independent review of the SemiAnalysis AgentX results. It should clarify how the tests were configured, which measurements produced the highest gains and how Vera Rubin compares with GB300 across models and context lengths.
Buyers should also watch for production availability, system-level power measurements and results that include CPU-based tool orchestration. Those details will help distinguish token-generation efficiency from end-to-end agent efficiency.
Finally, the market will need deployment evidence. The most meaningful comparisons will measure completed agent tasks, latency, reliability, utilization and total operating cost across real workloads rather than headline throughput alone.
NVIDIA’s announcement identifies a real bottleneck in agent deployment: long-context, multi-step inference can multiply compute demand faster than user-facing request counts suggest. Measuring work per megawatt is therefore more relevant to large-scale agents than relying on conventional chat benchmarks.
But the 30x figure should be treated as an early, vendor-controlled performance claim. Its significance will depend on independent replication and on whether the improvement survives the messy parts of production agents—tool latency, orchestration, failed calls, uneven traffic and end-to-end cost. For builders, the practical lesson is to benchmark complete workflows now, before choosing infrastructure based on short-context token rates.