NVIDIA's Groq 3 LPX inference accelerator is in full production, delivering 3,400 output tokens per second — 4x faster than the nearest alternative.
NVIDIA's Groq 3 LPX inference accelerator is in full production, delivering 3,400 output tokens per second — 4x faster than the nearest alternative.

NVIDIA's Groq 3 LPX inference accelerator reached full production, hitting 3,400 output tokens per second on the Gemma 4 31B model with a 100,000-token context — 4x faster than the nearest alternative platform for agentic AI workloads. The benchmark, run through Artificial Analysis, represents the fastest performance ever recorded for the model.
"Inference is the growth engine of AI," Jensen Huang, founder and CEO of NVIDIA, said. "Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation."
The Groq 3 LPX extends the Vera Rubin NVL72 platform, pairing Rubin GPUs for large-scale context processing with LPUs for latency-sensitive decode workloads. A rack-scale deployment can include 256 LP30 accelerators connected through direct chip-to-chip links. Nebius is the first AI cloud to adopt the chip, integrating it into its Nebius Token Factory production inference platform. Groq, the purpose-built AI inference cloud, plans to be among the earliest adopters. NVIDIA did not disclose the specific platform used in the 4x comparison.
The production milestone comes as NVIDIA faces intensifying competition in AI inference from Cerebras, AMD, and custom silicon from hyperscalers. NVIDIA's data center business generated $81.6 billion in Q1 2027 revenue, up 85 percent year over year, and the company's ability to extend the Vera Rubin platform with specialized inference acceleration could determine whether it maintains its dominant position as AI workloads shift from training to inference.
Agentic AI systems generate massive volumes of tokens across hundreds or thousands of inference steps, making token generation speed critical for agents to reason, act, and complete complex tasks in real time. The Groq 3 LPX is purpose-built to address decode latency — the rate at which tokens are generated for an individual user — which determines how quickly an agent can complete each step of its work. Faster generation gives agents more time to inspect files, write and test code, call tools, verify results, and iterate while maintaining a responsive user experience.
The LPX architecture is part of NVIDIA's broader Vera Rubin platform, which spans seven chips and five purpose-built racks. These rack platforms feature BlueField-4 DPUs and work with Vera CPU racks, Vera BlueField-4 STX storage, and Spectrum-6 SPX Ethernet to optimize multi-agent systems for the highest throughput per watt and lowest-latency inference.
Nebius plans to bring Groq 3 LPX to its Nebius Token Factory, giving developers access to extreme token generation speed for highly responsive agentic AI applications. "Generation is the phase of inference that determines how responsive an AI system actually is, and that's exactly what NVIDIA Groq 3 LPX is built to accelerate," said Danila Shtan, chief technology officer of Nebius. "As the first AI cloud bringing it to production via Nebius Token Factory, we're making sure every step of an agent's loop feels instant — through the same API developers are already using, with no migration to a new stack."
The adoption confirms real-world demand for the new chip, but NVIDIA faces pressure from competitors in the inference market. Cerebras recently introduced its CS-4 wafer-scale engine, claiming it generates in one second what a GPU rack needs 30 seconds for. AMD continues to push its MI300X and MI400 series into AI data centers. Meanwhile, hyperscalers including Amazon and Google are developing custom inference silicon to reduce their dependence on NVIDIA.
NVIDIA shares have been supported by the company's dominant position in AI accelerators, with analysts setting a median price target of $308.50. The company's Q1 2027 revenue of $81.6 billion represented an 85 percent year-over-year increase. The Groq 3 LPX production milestone, combined with the Vera Rubin platform rollout, could help NVIDIA maintain its pricing power as inference workloads become the primary growth driver in AI infrastructure spending.
This article is for informational purposes only and does not constitute investment advice.