Cognition says Nvidia's Vera Rubin rack gave it up to 4.8x the token throughput of GB200
Cognition, the lab behind the Devin coding agent, says Nvidia's Vera Rubin NVL72 rack delivered "up to a 4.8x increase in total token throughput for SWE-2 inference workloads over GB200 NVL72" in its early tests. The figure comes from Nvidia's post of September 30 about CoreWeave's deployment…
Sources
- T2Cognition says Nvidia's Vera Rubin rack gave it up to 4.8x the token throughput of GB200UploadVR / Road to VR / MIXED