NVIDIA Groq 3 LPX is the interactive AI inference accelerator for the NVIDIA Vera Rubin platform. At the core of the platform is NVIDIA Vera Rubin NVL72, the most versatile machine ever built, delivering high throughput and interactivity across the widest range of AI workloads—from small to large models, both open and closed. Groq 3 LPX, when paired with Vera Rubin NVL72, extends the platform’s ability to address the highest-interactivity serving tiers, expanding Vera Rubin’s ability to power user experiences as part of AI factories.
How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubin
NVIDIA Groq 3 LPX is the interactive AI inference accelerator for the NVIDIA Vera Rubin platform. At the core of the platform is NVIDIA Vera Rubin NVL72, the... NVIDIA Groq 3 LPX is the interactive AI inference accelerator for the NVIDIA Vera Rubin platform. At the core of the pl

In this post, we report the first third-party benchmark of performance on Groq 3 LPX systems: Artificial Analysis has run its 100K context benchmark on the Gemma 4 31B model on Groq 3 LPX, measuring a world-class interactivity of 3,431 output tokens/second. The technologies that power this performance unlock the ability of the Vera Rubin platform to serve multiagent systems powered by 2T+ parameter models with high interactivity and long context, through pairing Groq 3 LPX with Vera Rubin NVL72.
Why is long context at high interactivity important? [](#why_is_long_context_at_high_interactivity_important )
Agentic sessions are characterized by multiturn inference. At the end of each turn, the agent’s output is appended to the continually growing context that is fed into all subsequent turns.
Figure 1. Context carried into each turn rises steadily across a multiturn agentic session
Related stories
Samsung Galaxy S26 FE: Delivering the Latest Flagship Experience, Focused on What Matters Most
Samsung Newsroom
Samsung Electronics today announced Galaxy S26 FE, the newest addition to the Galaxy S26 family and the first in the lineup to launch with One UI 9 — bringing the latest premium Galaxy experiences to more users from day one. With enhanced camera capabilities and more context-awar

Anthropic pushes into physical world with new standard to help AI agents operate machines - CNBC
Anthropic News
Anthropic pushes into physical world with new standard to help AI agents operate machines CNBC
AWS Elastic Disaster Recovery introduces Recovery Plans for orchestrated application recovery
AWS What’s New
AWS Elastic Disaster Recovery (AWS DRS) now offers Recovery Plans, a capability that automates the sequential launch of multi-server applications during recovery and drills. Instead of launching servers one at a time and tracking dependencies manually, you define the recovery seq

Giga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rules
NVIDIA Developer Blog
The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales to span hundreds of thousands of GPUs,... The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales

How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit
NVIDIA Developer Blog
Atomistic simulation requires three things: knowledge of the science, compute-efficient implementation of simulations, and accessible interfaces to the... Atomistic simulation requires three things: knowledge of the science, compute-efficient implementation of simulations, and ac

How we saved 100 terabytes of memory by optimizing 1.1.1.1’s DNS cache
Sebastiaan Neuteboom
Big Pineapple , the platform behind 1.1.1.1 , Gateway DNS , DNS Firewall , AS112 , and several other Cloudflare DNS services, stores over 250 billion DNS cache entries at any given time. At that scale, wasting a single byte per entry costs more than 250 gigabytes of memory across
