Nvidia Groq 3 LPX Inference System Enters Production, Promises 35x Throughput Per Megawatt
The AI compute build-out sees semiconductor revenue surging to $1.6 trillion in 2026, though rising memory costs and power infrastructure bottlenecks continue to challenge rapid expansion.
The story
This week, Nvidia announced its Groq 3 LPX inference system has entered full production, with deployments at cloud provider Nebius slated for later this year. The system, produced by Samsung Foundry on a 4nm process, integrates 256 language processing units and aims to accelerate AI inference workloads.
Benchmarks by Artificial Analysis indicate the LPX system can process 3,400 tokens per second using Google's Gemma 4 31B model, a speed Nvidia claims is four times faster than the nearest alternative platform. When paired with Nvidia's Vera Rubin processors, the combined system is estimated to deliver up to 35 times the throughput per megawatt.
This launch comes as Gartner forecasts global semiconductor revenue to reach $1.6 trillion in 2026, with AI data centers driving a significant portion of this growth. However, the industry is also grappling with escalating memory costs and the immense power requirements for these high-density AI clusters.
Silicon
Groq 3 LPX
Maker: Nvidia (produced by Samsung Foundry)
What: AI inference system with 256 language processing units, processes 3,400 tokens/second with Google Gemma 4 31B model, 4x faster than alternatives.
For Whom: Nebius and other cloud providers deploying AI inference workloads.
The build-out
Data center power infrastructure investment
Who: Nvidia, Cloverleaf Infrastructure
Scale: Hundreds of millions of dollars investment, securing over 10 GW of planned data center power capacity
Where: Various locations for large-scale data center projects
Supply & policy signals
Nvidia raising AI server prices by over 15% for Grace Blackwell and Vera Rubin systems
Implication: Reflects surging memory costs, with server DRAM doubling in Q1 2026, making memory a quarter of high-end AI server rack costs.
Gartner forecasts DRAM revenue to increase 246.6% and NAND flash revenue 371.9% in 2026
Implication: Indicates strong demand and sharply stronger pricing, suggesting more dollars are spent without necessarily proportional growth in deployed capacity.
Taiwan experiencing an AI-driven power crunch, impacting foundry capacity and chip prices
Implication: Foundry capacity is tightening, and chip prices are climbing into the second half of 2026, despite a nearly 90% year-over-year increase in Taiwan's July ICT export orders led by AI servers.
TSMC increased its 2026 capital expenditure forecast to $60-64 billion
Implication: Highlights ongoing massive investment to meet long-term capacity plans and technology development roadmap driven by strong demand for AI applications.
What we'll be watching
- Nvidia's earnings report on Wednesday, August 26, for insights into data center revenue and future guidance.
- Further details from SEMICON Taiwan 2026 and the National Data Centers Summit 2026, both occurring on August 26, for new announcements and industry trends.
- Deployment of Nvidia's Groq 3 LPX inference systems at Nebius later this year.
- Developments regarding new high-bandwidth memory (HBM) capacity, as Deloitte projects meaningful new supply won't arrive until 2029 or 2030, and Gartner expects the crunch to persist through H1 2027.
Reporting + analyst voices: grounded via Google Search at publish time.