NVIDIA says Vera Rubin NVL72 can deliver up to 30x more agentic inference throughput per megawatt than GB300, pending independent benchmark review.

NVIDIA says its upcoming Vera Rubin NVL72 platform can deliver up to 30 times more agentic-AI throughput per megawatt than the company’s GB300 NVL72 system. The claim comes from preview results on SemiAnalysis AgentX, a benchmark designed to replay production-style coding-agent sessions rather than fixed-length chatbot prompts.
The result matters because AI agents consume substantially more inference capacity than ordinary chat. NVIDIA’s sources cite OpenRouter data showing that a single agentic request uses about 15 times as many tokens as a simple chat request, as agents repeatedly reason, call tools, delegate work to sub-agents and carry an expanding context through a task.
The Vera Rubin comparison is not yet an independently validated result. NVIDIA measured the figures using the SemiAnalysis AgentX workload and says the results are pending SemiAnalysis review. That makes the announcement an early vendor-reported performance signal, not a confirmed industry benchmark.
AgentX is part of SemiAnalysis’s InferenceX benchmark suite. According to NVIDIA Developer Blog, it replays prerecorded Claude Code sessions through the AIPerf client, preserving the sequence of model calls, reasoning intervals, tool use, context growth and tool-call latency captured in the original trajectories.
That design attempts to address a weakness in conventional inference testing. Static tests often use predetermined input and output lengths, such as an 8,000-token prompt followed by a 1,000-token response. Real coding agents generate uneven traffic: a task may alternate between model reasoning, file operations, searches, compiler output and further model calls, while the accumulated context becomes much larger over time.
AgentX measures tokens per megawatt alongside user-experience indicators including time to first token, end-to-end latency and interactivity. NVIDIA’s headline Vera Rubin result was measured at 160 tokens per second per user on the AgentX DeepSeek V4 Pro workload. The company says Vera Rubin NVL72 reached up to 30 times the AI-factory throughput per megawatt of GB300 NVL72 at that operating point.
The benchmark also reports substantial gains for the current Blackwell generation. NVIDIA says GB300 NVL72 delivers up to 15 times the throughput per megawatt of H200 NVL8 on DeepSeek V4 Pro 1.6T, and up to 80 times the result on Kimi K3 2.8T. Those figures are also presented through NVIDIA’s reporting of the AgentX results and should be treated accordingly.
NVIDIA attributes the projected efficiency gains to a combination of hardware scale and software optimization rather than to the Rubin GPU alone. The NVL72 design connects 72 GPUs in a scale-up domain, allowing workloads to distribute model experts and maintain context across the system with high-bandwidth interconnects.
That matters for long-running agents because previously processed context can often be reused instead of being recomputed. NVIDIA describes a serving design that separates prefill, where the system processes incoming context, from decode, where it generates output. Rate matching between those stages is intended to prevent one side from sitting idle while the other becomes a bottleneck.
The company also points to distributed key-value caching, KV-aware request routing and cache offloading to host memory or storage. These techniques are designed to keep relevant portions of a growing agent session available across multiple requests. NVIDIA says sixth-generation NVLink and NVLink Switch technology provides higher packet rates and lower latency than conventional Ethernet alternatives for this scale-up communication.
The software layer includes NVIDIA TensorRT-LLM, NVIDIA Dynamo and CUDA kernels optimized for mixture-of-experts models. The developer-focused account also names SGLang and vLLM as serving runtimes used in the broader optimization stack, along with DeepGEMM kernels and lower-precision formats such as MXFP4 and MXFP8. NVIDIA says NVFP4 quantization on Rubin is intended to reduce memory use and increase throughput while preserving output quality, though the supplied evidence does not provide an independent quality assessment.
NVIDIA further says its DSX MaxLPS technologies can manage power at GPU, rack and workload levels, potentially allowing up to 40% more GPUs within the same megawatt budget. That is a capacity-planning claim from NVIDIA, not a demonstrated result established by the AgentX figures alone.
For teams building coding assistants, research agents or customer-service automation, the relevant metric is no longer simply how quickly one prompt produces an answer. A production agent may keep several model calls active, preserve a large context and pause while external tools return data. Infrastructure that performs well on short, isolated requests may therefore waste power or lose responsiveness on real workflows.
The Vera Rubin announcement frames power efficiency as both a deployment constraint and a pricing issue. NVIDIA says Vera Rubin NVL72 can reduce the cost per million tokens by up to 35 times relative to GB300 NVL72 on the cited workload. Its developer blog separately reports that GB300 can deliver up to 10 times lower token cost than H200 NVL8. These calculations depend on the benchmark configuration, energy assumptions and utilization levels, so enterprise buyers will need to test them against their own models and serving patterns.
The result could be most relevant to operators serving large mixture-of-experts models or agents with unusually long sessions. In those environments, memory movement, cache reuse and interconnect behavior can matter as much as raw accelerator arithmetic. Smaller teams, however, may not see the same economics if they use short-context models, low concurrency or hosted APIs that hide the underlying hardware.
There is also a reliability question. More aggressive caching, disaggregated serving and tool orchestration can improve utilization, but they add scheduling and state-management complexity. Builders will need to evaluate failure recovery, cache invalidation, tail latency and workload isolation rather than relying on peak tokens-per-second figures.
The first signal will be SemiAnalysis’s review of the Vera Rubin AgentX results. Independent publication of the test configuration, power measurements, model settings and interactivity targets will determine how comparable the 30x claim is across systems.
Buyers should also watch for results that include the Vera CPU’s role in tool calling. NVIDIA explicitly says the early figures do not yet reflect Vera CPU performance for that part of the workflow. Since agents spend time outside model generation, end-to-end measurements could differ from accelerator-only throughput.
Other useful follow-ups include availability and pricing for Vera Rubin NVL72, performance on models beyond DeepSeek V4 Pro, and results from operators running their own traces rather than prerecorded sessions. Changes to NVIDIA’s software stack could also affect the comparison, since the company says both Rubin and GB300 performance will improve through continuing optimization.
NVIDIA’s announcement is important less because it establishes a final 30x industry standard than because it highlights a changing definition of inference performance. Agent workloads expose bottlenecks that fixed prompt tests can miss: context reuse, tool-call gaps, cache capacity, concurrency and power delivery.
The practical takeaway for AI builders is to benchmark complete agent trajectories before committing to infrastructure. NVIDIA’s early results suggest that tightly integrated hardware and serving software may materially change the economics of long-running agents, but the scale of that advantage remains unconfirmed until the underlying AgentX measurements are independently reviewed and reproduced.