Cognition says Nvidia's Vera Rubin rack gave it up to 4.8x the token throughput of GB200

Cognition, the lab behind the Devin coding agent, says Nvidia's Vera Rubin NVL72 rack delivered "up to a 4.8x increase in total token throughput for SWE-2 inference workloads over GB200 NVL72" in its early tests. The figure comes from Nvidia's post of September 30 about CoreWeave's deployment…

aihardwareiottelecomxr

Sources