What NVIDIA measured

NVIDIA published early Vera Rubin NVL72 performance using SemiAnalysis AgentX, a workload built from recorded agentic coding sessions. The trajectories preserve growing context, tool calls and sub-agent activity instead of reducing the test to a short, fixed prompt.

The company reports up to 30 times higher throughput per megawatt and up to 35 times lower cost per million tokens than GB300 NVL72 on the stated agentic workload. Those are the headline claims, not independently established market-wide results.

Why the workload is different from chat

In a simple chat request, input and output lengths are often bounded. An agent repeatedly adds observations, files and tool results to its working context, then generates further actions. The serving system must handle changing sequence lengths and uneven work across many sessions.

A benchmark based on real trajectories is therefore more informative than a single sequence length. It still represents a chosen workload, model mix and scheduling policy, so readers should not transfer the multiplier unchanged to customer service, research or other agent patterns.

Important limitations in the announcement

NVIDIA says the results are pending SemiAnalysis review and do not yet incorporate Vera CPU performance for tool calling. The post also describes ongoing software optimization, meaning the comparison is an early platform snapshot rather than a settled independent benchmark.

The largest multiplier is reported for specific points on a throughput-and-interactivity curve. Capacity planners should examine the entire curve and the latency target they must meet instead of selecting the most favorable isolated ratio.

What the claim changes

Even with those limits, the metric choice is important. Power-constrained data centers increasingly care about completed work per megawatt, not peak chip operations. Agent workloads make that shift more urgent because they can consume many more tokens than one-shot requests.

Operators should reproduce the workload with their model, context distribution and concurrency, then include cooling, idle power and external tool time. Useful efficiency is the number of correct workflows completed inside an energy and latency budget.

Explore further

Follow the wider AI landscape from the AINewsInu homepage, where our editors connect product updates, reviews and practical analysis.

For first-party product information, Read NVIDIA's Vera Rubin results.

Sources & further reading

Social-media activity is treated as a signal of attention, not proof. Product claims are attributed to the linked publisher or announcement.