NVIDIA Puts Groq 3 LPX Into Production to Extend Vera Rubin for Agentic Inference

NVIDIA has moved Groq 3 LPX into full production alongside Vera Rubin, targeting faster, long-context inference for agents and AI cloud providers.

AI News

NVIDIA says its Groq 3 LPX inference system has entered full production and is being integrated with the Vera Rubin NVL72 rack-scale platform. The combination is designed to accelerate token generation for agentic applications, where models repeatedly reason, call tools and process increasingly long conversational or task histories.

The announcement expands NVIDIA’s Vera Rubin strategy beyond general-purpose accelerated computing. The company is positioning Rubin GPUs for large-scale context processing while Groq 3 LPX handles latency-sensitive decoding, the stage where a model generates output one token at a time. That division is aimed at AI cloud providers and enterprises seeking more responsive agents without giving up throughput across large deployments.

NVIDIA adds a low-latency layer to Vera Rubin

According to NVIDIA, Groq 3 LPX is an interactive AI inference accelerator designed to work alongside Vera Rubin NVL72 rather than replace it. The company describes the platform as a combination of GPUs and LPUs, or local processing units, with each component assigned to workloads where it is most effective.

The need for that separation comes from the behavior of agentic workloads. An agent may generate an answer, use an external tool, receive new information and begin another inference turn. Each turn can add to the session’s context, eventually requiring the system to process tens or hundreds of thousands of tokens while still producing a response quickly.

NVIDIA says Rubin GPUs can manage the large context-processing workload, while Groq 3 LPX is optimized for decode performance. The system can also be configured for prefill-decode disaggregation, attention-FFN disaggregation and external-drafter speculative decoding. These are infrastructure techniques that divide model-serving tasks or use predicted tokens to improve response speed.

A rack-scale deployment can include 256 LP30 accelerators connected through direct chip-to-chip links, NVIDIA says. The company’s developer documentation describes 128 GB of combined SRAM across those chips and a compiler that plans computation and communication before execution. That deterministic approach is intended to reduce the coordination overhead that can become significant when models are served with small batch sizes.

What the benchmark does—and does not—show

NVIDIA’s strongest performance claims come from an Artificial Analysis benchmark, making them third-party measurements cited by the vendor rather than independent evidence of broad production performance. On the Gemma 4 31B model with a 100,000-token context, Artificial Analysis measured a median speed of 3,431 output tokens per second on Groq 3 LPX, according to NVIDIA’s developer blog.

NVIDIA’s corporate announcement rounds that figure to 3,400 output tokens per second and says the result was four times faster than the nearest alternative platform. The source material does not identify that competing platform in the announcement, so the comparison cannot be independently assessed from the available evidence.

A second result comes from SPEED-Bench, a coding benchmark using Gemma 4. NVIDIA reports a median of 4,767 output tokens per second and an 80th-percentile result of 5,520 tokens per second. Those numbers describe a specific model, context and test methodology; they should not be treated as a universal measure of performance across models, batch sizes, context lengths or production traffic patterns.

The technical explanation offered by NVIDIA is that Groq 3 LPX uses compiler-scheduled chip-to-chip communication and overlaps data movement with computation. In conventional parallel inference, distributing work across processors can introduce synchronization and collective-communication costs. NVIDIA argues that preplanning transfers can remove the need for some real-time arbitration, particularly at the low batch sizes associated with highly interactive serving.

Neither source provides customer deployment volumes, pricing, power consumption, availability timelines for general developers or independently verified production utilization. The evidence establishes NVIDIA’s product and architecture claims, but it does not yet establish how the system performs economically across a broad range of commercial workloads.

Early cloud and infrastructure adopters

NVIDIA identifies three partners associated with the Vera Rubin expansion. Nebius is described as the first AI cloud to adopt Groq 3 LPX. NVIDIA says the hardware will be added to the company’s Token Factory so developers can build interactive agents, coding systems and other real-time applications at scale.

CoreWeave, meanwhile, has deployed Spectrum-X Multiplane in production, according to NVIDIA. The networking system connects Vera Rubin racks through multiple parallel switches, with the stated goal of providing high-bandwidth and lossless communication across the AI cloud infrastructure.

NVIDIA also says SpaceXAI plans to use Vera CPUs in its next-generation agentic AI architecture. The company links those CPUs to orchestration, tool use, code execution, data processing and simulation, including infrastructure spanning terrestrial data centers and orbital satellites. This is a stated plan, not evidence that such a deployment is already operating at scale.

These partner references indicate that NVIDIA is selling Vera Rubin as a complete AI factory stack: CPUs for orchestration and general-purpose work, GPUs for context and model computation, Groq accelerators for rapid decoding, and networking to connect the components. The commercial significance will depend on whether customers can deploy that stack more efficiently than separate accelerator pools optimized for different serving stages.

Why the design matters for AI builders and enterprises

For builders of coding assistants, research agents and customer-service systems, decode latency is experienced directly by users. A faster first response or faster intermediate output can make a multistep workflow feel more interactive, particularly when an agent must perform many sequential calls.

Long context also changes the infrastructure equation. A session that repeatedly reprocesses its history can consume substantial memory and compute even when the model itself is not unusually large. NVIDIA’s design targets that combination of large key-value caches, small-batch inference and frequent communication between processors.

The practical question for enterprises will be less about peak tokens per second than about service-level performance under real traffic. Buyers will need to evaluate latency at different concurrency levels, context sizes and model architectures, as well as the cost of keeping specialized inference capacity utilized. A system that is exceptionally fast for a benchmark workload may be less attractive if demand is variable or if applications require models and software that are not yet optimized for the platform.

The announcement also intensifies competition around inference specialization. NVIDIA is using its broader platform reach to pair a dedicated low-latency accelerator with its upcoming GPU, CPU and networking products. That could appeal to cloud providers building differentiated inference services, but it also increases software and operational complexity compared with a single accelerator architecture.

What to watch next

The clearest follow-up signal will be whether Nebius exposes Groq 3 LPX-backed services to developers and publishes production performance, pricing or capacity information. Those details would help distinguish a hardware announcement from a generally accessible serving option.

Customers should also watch for independent results beyond Gemma 4 31B and the reported coding benchmark. Results across larger and smaller models, varying context lengths, concurrent users and complete agent workflows would provide a stronger basis for evaluating the platform.

Deployment evidence from CoreWeave and SpaceXAI will be another important indicator. NVIDIA has confirmed CoreWeave’s production networking deployment and described SpaceXAI’s Vera CPU plans, but the sources do not provide scale, workload or utilization data. Power efficiency, software support and compatibility with popular inference frameworks will likewise determine whether the architecture gains adoption outside NVIDIA’s named partners.

Creati.ai perspective

NVIDIA’s news is best understood as a production and systems-integration milestone, not simply the launch of another inference chip. The company is responding to a specific serving problem: agents generate long chains of tokens while carrying large context, making both decode latency and interconnect overhead commercially important.

The reported results are promising but remain bounded by vendor-selected benchmarks and limited public deployment evidence. For AI teams, the meaningful test will be whether Groq 3 LPX and Vera Rubin deliver predictable latency and acceptable economics on their own agent workloads. If they do, NVIDIA’s integrated “token factory” approach could become a serious option for high-interactivity services; if not, specialized components may remain difficult to justify despite strong peak measurements.

Ads