Nvidia announced a new AI platform called Vera Rubin, which the company says delivers a substantial performance improvement over the previous Grace Blackwell generation. Nvidia gave an example: CoreWeave, using the new configuration, can process ten times more tokens per watt when running the DeepSeek R1 model, highlighting gains in power efficiency and throughput for large language model workloads.
Tokens are the measurement units used in AI systems to represent pieces of input and output text — words or subword fragments — and they are commonly used to quantify the amount of work a model performs.
Vera CPU and benchmark figures
Nvidia also emphasized the standalone capabilities of the Vera central processing unit (CPU). According to the company, certain benchmark tests show the Vera CPU delivering up to 1.9 times better agent AI performance compared with the AMD Epyc Turin processor. Nvidia's statements are based on its internal measurements.
Timing and context
The announcement arrived ahead of AMD’s “Advancing AI” conference, which begins on Wednesday in San Francisco; Lisa Su, Chief Executive Officer of AMD, is scheduled to give the opening keynote on Thursday. The timing underscores the intensifying competition among chip makers.
Market scale and rising competition
Nvidia remains the dominant player in the data-center market: the company reported revenue of $215.9 billion in its 2026 fiscal year, compared with $34.6 billion generated by AMD in the comparable period. Still, competition is increasing as more major customers and manufacturers develop their own AI accelerators.
Among Nvidia’s customers, Google and Amazon are building custom AI accelerators, and Meta has developed its own chips to reduce dependence on Nvidia; Microsoft has made similar moves. OpenAI and Broadcom have also recently announced their own AI processors. These developments indicate that Nvidia’s dominance could face growing pressure as new architectures and suppliers enter the market.
Why this matters
Improvements in AI accelerators and efficiency have direct implications for how large language models and other generative AI systems are deployed and at what cost. Higher energy efficiency — for example, processing more tokens per watt — can lower operating expenses and enable wider adoption of larger models. At the same time, the proliferation of alternative chips and in‑house accelerators means market dynamics are likely to remain fluid going forward.
(This article is based on reporting from Erstemarket.)



