
Nvidia Vera Rubin NVL72 on CoreWeave 10x More Tokens per Megawatt Than Blackwell
In early June, CoreWeave completed the industry's first bring-up and validation of NVIDIA Vera Rubin NVL72 , standing up the rack-scale system end-to-end, with power, cooling, networking, and compute all verified and running. That milestone showed the platform running at production scale. The next question everyone has asked: How does it perform? Today, we’re happy to share the first-ever measured silicon performance on Vera Rubin NVL72.
DeepSeek R1, one of the reasoning models driving the current surge in agentic AI, was run across both Vera Rubin NVL72 and NVIDIA Blackwell NVL72 . In this evaluation, token throughput per megawatt was measured against interactivity (as TPS/user), which indicates the responsiveness a user or agent experiences while a model reasons through a task. At a matched interactivity target, Vera Rubin NVL72 generated 10x tokens-per-second per megawatt compared to NVIDIA GB200 NVL72 on the same DeepSeek R1 workload.
This evaluation compares the performance on each platform with all major, state-of-the-art inference optimizations turned on. These include large-scale expert parallelism, NVFP4 precision, multi-token prediction, and disaggregated prefill and decode, all enabled using NVIDIA TensorRT-LLM and NVIDIA Dynamo. At similar interactivity targets, Vera Rubin NVL72 delivers an order-of-magnitude improvement in token throughput per megawatt, empowering CoreWeave customers to run far more reasoning-model traffic within the same power budget or run the same traffic on far less power. That means lower cost-per-token and more room to scale at the same infrastructure investment, which is beneficial for agentic workloads that call tools and reason across many steps.
Here's what this will look like in practice:
Reasoning models like DeepSeek R1 don't answer in a single pass. They plan, check their own work, and revise, generating far more tokens per query than a typical response. Every one of those tokens has to move through memory and across the NVLink domain, which makes interactivity, not raw FLOPs alone, the number that decides whether an agent feels instant or sluggish.
That's exactly the constraint NVIDIA Vera Rubin NVL72 was built to address. 72 NVIDIA Rubin GPUs and 36 NVIDIA Vera CPUs share a single rack connected by a 260 TB/s all-to-all NVLink 6 fabric, with native NVFP4 support built into the Transformer Engine.
We've invested heavily in the operational layer underneath it to make sure that these efficiency gains hold up once a customer's workload is running in production at scale. We've written before about how CoreWeave operates rack-scale AI as a cloud service , from the patent-pending liquid cooling in Valvey to the unified rack control in Racky, all coordinated by our Rack LifeCycle Controller.
Hacker News
news.ycombinator.com