A specialized stage in the inference pipeline

NVIDIA says Groq 3 LPX is now in full production as an extension to Vera Rubin NVL72. Instead of treating one processor as equally suited to every inference phase, the platform combines Rubin GPUs with LPUs aimed at deterministic, low-latency token generation.

A rack-scale deployment can include 256 LP30 accelerators connected through direct chip-to-chip links. NVIDIA presents the system as a codesigned platform, not a standalone replacement for the GPUs and networking that process context and coordinate the wider workload.

The reported performance

In a benchmark attributed to Artificial Analysis, NVIDIA reports 3,400 output tokens per second for Gemma 4 31B with a 100,000-token context, four times the nearest alternative platform in that test.

The result is relevant to interactive agents, but it is not a universal application benchmark. Model choice, batching, concurrency, input length and the division of work between accelerators can all change the experience. Buyers should seek the underlying configuration and reproduce representative traffic.

Why agents expose decode latency

An agent may call tools, receive results and generate another decision many times. Each turn creates a sequential dependency, so small per-token delays can accumulate across a long workflow even when the total output is modest.

Fast decode cannot remove external API delays, slow code execution or poor planning. The platform will create the most value when token generation is a measured bottleneck in a well-instrumented agent, rather than when teams add hardware before understanding where time is spent.

Adoption signals and open questions

NVIDIA names Nebius as the first AI cloud adopting Groq 3 LPX, says CoreWeave has deployed Spectrum-X Multiplane, and reports that SpaceXAI plans to use Vera CPUs for orchestration-heavy agent work.

Availability, pricing and independently reproduced economics will determine how quickly the architecture spreads. Platform teams should compare full request latency, throughput under concurrency, power, model coverage and operational complexity with alternative inference stacks.

Explore further

Follow the wider AI landscape from the AINewsInu homepage, where our editors connect product updates, reviews and practical analysis.

For first-party product information, Read NVIDIA's Groq 3 LPX announcement.

Sources & further reading

Social-media activity is treated as a signal of attention, not proof. Product claims are attributed to the linked publisher or announcement.