Nvidia says Groq racks will be online this year after $20 billion deal

Nvidia has announced that its Groq 3 LPX racks, utilizing technology from its $20 billion acquisition, are in production and will be deployed to cloud providers later this year. These chips are designed to provide low-latency inference, a critical requirement for responsive AI agents and premium cloud services.
Why it matters
The commercialization of this technology marks a significant step in the AI hardware race, specifically targeting the demand for faster, more efficient token processing.
Nvidia announced Monday that its Groq 3 LPX rack is in full production, marking the commercialization of technology from the company's largest acquisition on record.
The Groq rack will be deployed alongside Vera central processors and Rubin graphics processors at neocloud Nebius , and will be online later this year, Nvidia senior director Dion Harris told reporters.
Nvidia's race to manufacture Groq's chip and make it available to customers highlights the growing importance of low-latency inference that's needed to make artificial intelligence agents feel responsive without long lags for users, especially for coding. Cloud companies can charge more for these kinds of tokens, Nvidia says.
"For folks who are serving tokens, it unlocks the ability to offer premium tiers of service for those users and those customers who actually demand the most latency-sensitive" service agreements, Harris said.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in