NVIDIA announced at Hot Chips that its Groq 3 LPX interactive AI inference accelerator is now in full production, marking a major milestone for the company's agentic AI strategy.
Groq 3 LPX extends the inference performance of NVIDIA's Vera Rubin NVL72 systems by dramatically increasing token generation rates. The distinction matters because agentic AI systems — software agents that reason, code, call tools, and iterate through complex tasks — can generate massive volumes of tokens across hundreds or thousands of inference steps. The faster each token arrives, the more time agents have to inspect results, test code, and verify outcomes.
In Artificial Analysis benchmarking, Groq 3 LPX delivered a record 3,400 output tokens per second running Gemma 4 31B, an open-source agentic model, with a 100,000-token context window — the fastest performance ever recorded for that model. NVIDIA says this translates to 4x faster responsiveness than the nearest alternative platform, enabling agentic coding tasks that previously took hours to complete in minutes.
"Inference is the growth engine of AI," said Jensen Huang, NVIDIA's founder and CEO. "Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation."
Nebius, an AI cloud provider, is the first to adopt Groq 3 LPX, bringing it to production through its Nebius Token Factory inference platform. Groq, the purpose-built AI inference cloud, plans to follow as one of the earliest adopters.
The announcement comes as AI companies are increasingly focused on inference speed as a competitive differentiator. While training has traditionally consumed the bulk of AI infrastructure spending, the rise of agentic systems — which spend far more time reasoning and generating output than on one-time model training — has shifted industry attention toward making inference as fast and efficient as possible.
Groq 3 LPX is available through cloud providers and OEM partners including AWS, Azure, and Oracle in the second half of 2026.




