NVIDIA Claims Vera Rubin NVL72 Delivers Up to 30x More Agentic AI Work per Watt

NVIDIA says Vera Rubin NVL72 can deliver up to 30x more agentic AI throughput per megawatt than GB300, reshaping inference economics.

AI News

NVIDIA says its upcoming Vera Rubin NVL72 platform can deliver up to 30 times the agentic AI throughput per megawatt of its GB300 NVL72 predecessor. The claim comes from early results on SemiAnalysis AgentX, a benchmark designed to replay production-style coding-agent sessions rather than fixed-length chatbot prompts.

The result matters because AI agents generate far more inference traffic than conventional chat. NVIDIA cites OpenRouter data showing that an agentic request consumes roughly 15 times as many tokens as a simple chat interaction. Agents repeatedly reason, call tools, delegate tasks to sub-agents and carry an expanding context through each step, putting pressure on both computing capacity and power budgets.

NVIDIA’s results are vendor-reported and remain pending review by SemiAnalysis. They therefore indicate the company’s performance target for Rubin, not an independently confirmed industry ranking. Still, the benchmark choice points to a broader change in how infrastructure providers may need to measure AI systems as multi-step workloads move into production.

A benchmark built around real agent behavior

The AgentX workload, part of SemiAnalysis’s InferenceX benchmark suite, replays recorded Claude Code sessions with interleaved model reasoning and tool use. It preserves the original sequence lengths, context growth, reasoning intervals and tool-call delays, according to NVIDIA Developer Blog material.

That design differs from traditional inference tests built around a fixed input and output size. In an agent session, the prompt can expand from thousands of tokens to hundreds of thousands as the system accumulates research, code, tool results and delegated work. Previously processed context may also be reused through a key-value cache, while tool execution creates gaps between model calls.

AgentX varies concurrency and evaluates throughput against interactivity measures such as time to first token, end-to-end latency and tokens per second per user. NVIDIA says the headline Rubin result was measured at 160 tokens per second per user on the AgentX DeepSeek V4 Pro workload. At that operating point, the company reports up to 30 times higher throughput per megawatt than GB300 NVL72.

The comparison is not a general statement about every model or deployment. The cited result is tied to a particular platform configuration, workload and interactivity target. NVIDIA also says the measurement does not yet include Vera CPU performance for tool calling, leaving part of the end-to-end agent system outside the reported figure.

What NVIDIA is changing in the inference stack

NVIDIA attributes the result to a combination of hardware, networking and software rather than to the Rubin GPU alone. The NVL72 design creates a large scale-up domain in which GPUs share high-bandwidth, low-latency communication through sixth-generation NVLink and NVLink Switch systems.

That topology supports techniques aimed at long-context and mixture-of-experts workloads. NVIDIA highlights distributed KV caching, which keeps previously processed context available across GPUs, and KV-aware routing, which can direct a request toward hardware that already holds relevant cache data. Disaggregated serving separates prefill, where context is processed, from decode, where responses are generated, allowing each stage to scale independently.

The software layer includes NVIDIA TensorRT-LLM, NVIDIA Dynamo and CUDA kernels designed to combine computation with inter-GPU communication. The company also points to NVFP4 quantization and newer Tensor Cores and Transformer Engine capabilities as ways to reduce memory use and increase token throughput.

NVIDIA’s DSX MaxLPS power-management technology is presented as another lever. The company says it can provision up to 40% more GPUs within the same megawatt budget by managing power at the GPU, rack and workload levels. NVIDIA separately claims Rubin can reduce cost per million tokens by up to 35 times versus GB300 NVL72 on the cited workload.

Evidence, comparisons and limits

The same AgentX data gives a more gradual picture of NVIDIA’s generational gains. NVIDIA reports that GB300 NVL72 delivers up to 15 times the throughput per megawatt of H200 NVL8 for DeepSeek V4 Pro 1.6T, and up to 80 times for the larger Kimi K3 2.8T model. It also claims GB300 provides up to 10 times lower cost per million tokens than H200 in the cited comparison.

Those figures are also presented by NVIDIA using the SemiAnalysis workload. The underlying benchmark is intended to make comparisons more realistic by replaying identical traffic, but the Rubin measurements have not yet received the independent review referenced in NVIDIA’s posts. Readers should therefore distinguish between the benchmark methodology and the performance numbers reported from it.

The claims also do not establish that every agent workload will see a 30x gain. Results can vary with model architecture, context length, cache reuse, concurrency, tool latency, quantization settings and the required response speed. A coding agent that frequently waits for a compiler or external API may be constrained by software and orchestration overhead rather than raw accelerator throughput.

Why the result matters to AI builders and buyers

For AI infrastructure operators, performance per watt is becoming a capacity and unit-economics metric, not merely an environmental measure. If the reported gains hold in independent testing, a data-center operator could serve more concurrent agent sessions within a fixed power allocation, or provide the same capacity with fewer accelerators and lower operating cost.

That could affect how companies design AI factories for coding assistants, research agents, customer-service systems and internal automation. Teams may need to optimize the full workflow: context caching, prefill and decode scheduling, model routing, tool-call latency and the handling of sub-agent traffic. Selecting a faster GPU without changing those surrounding systems may leave a substantial portion of the claimed benefit unrealized.

The results also raise the competitive bar for infrastructure vendors. Fixed 8K-input and 1K-output tests remain useful for controlled comparisons, but they may say less about stateful agent traffic. Buyers evaluating an AI platform should ask whether its benchmarks preserve context growth, tool pauses, cache reuse and realistic user-experience targets.

For model developers, the emphasis on long-context serving and mixture-of-experts execution could make infrastructure-aware design more important. Quantization and caching can lower cost, but they also introduce quality, memory-management and operational trade-offs that must be tested on the actual agent workflow rather than inferred from a single throughput number.

What to watch next

The most immediate signal is SemiAnalysis’s review of the Vera Rubin NVL72 AgentX results. Independent publication of the test configuration, power accounting, model settings and interactivity measurements will help determine how reproducible the 30x figure is.

Benchmark results across additional models will also matter. NVIDIA cites DeepSeek V4 Pro and Kimi K3 prominently, but agent workloads differ substantially in context growth, expert routing and tool use. Results on other coding, research and enterprise models could show whether Rubin’s advantage is broad or concentrated in particular serving patterns.

Deployment evidence will be another test. NVIDIA’s claims concern preview measurements rather than disclosed customer production results. Buyers should look for validated cost per million tokens, sustained performance under cache pressure and complete end-to-end latency that includes tool calls and CPU-side orchestration.

Creati.ai perspective

NVIDIA’s announcement is most significant for reframing the infrastructure contest around useful agent work per unit of power. The headline 30x claim is not yet independent evidence, but the AgentX methodology addresses a real weakness in older inference benchmarks: they often model a single request instead of the expanding, interrupted workflow produced by an AI agent.

For builders and enterprise buyers, the practical lesson is to evaluate the entire serving system, not just accelerator specifications. Rubin’s reported advantage will matter only if caching, scheduling, model execution and tool orchestration can convert it into lower latency or lower token cost in production. Until the benchmark review and broader deployments arrive, NVIDIA’s figures should be treated as an important vendor claim and a signal of where AI infrastructure competition is heading.

Ads