CoreWeave has put seven Nvidia Vera Rubin NVL72 racks into live service across two regions, yielding 504 GPUs in total, as the AI cloud provider scales Nvidia’s newest platform beyond a single rack, TechTarget reported on 16 September. CoreWeave shares rose after the announcement, and the company is valued at $37.1 billion, up 11% so far this year, Barchart reported.
Key points
- Seven Vera Rubin NVL72 racks totaling 504 GPUs now in production across two regions
- CoreWeave shares rose after the announcement, with the company valued at $37.1 billion, up 11% this year
- Each Rubin GPU provides 1.6 Tbps of scale-out network connectivity via dual ConnectX-9 SuperNICs
- New AI Object Storage features include cross-region write acceleration and an Archive tier
- Blockfusion signed a 15-year anchor lease for CoreWeave’s Niagara Falls AI campus
Multi-rack Vera Rubin deployment
A single NVL72 rack combines 72 Rubin GPUs with 36 Vera CPUs, Nvidia NVLink 6, ConnectX-9 SuperNICs, BlueField-4 DPUs, high-speed storage and liquid cooling. The September deployment connects hundreds of Rubin GPUs as a single scale-out cluster, according to Chen Goldberg, executive vice president of product and engineering at CoreWeave, as reported by ROI-NJ on 16 September. Goldberg said the multi-rack system gives customers building agentic AI more scale, swifter iteration and better productivity as models and agents keep learning and improving.
Networking and operational complexity
CoreWeave senior vice president of product Corey Sanders told TechTarget that shifting from one NVL72 rack to several racks presented the toughest challenges in networking and operations, noting that operational complexity rises faster than a simple multiple. The company’s architecture uses two-tier, non-blocking connectivity with multiple rails and planes. Each Rubin GPU has two ConnectX-9 SuperNICs, providing up to 1.6 Tbps of scale-out network connectivity per GPU. CoreWeave says its architecture can scale to roughly 128,000 GPUs per rail, though this is described as an architectural-scale target rather than a deployed cluster. The company turns off quicker GPU-to-GPU links during tests and pushes traffic through the backend network to reveal issues with switches, cabling and software.
Stephen Sopko, practice lead for semiconductors and deep tech at HyperFrame Research, told TechTarget that the single rack is Nvidia’s product while the multi-rack domain is the operator’s product. Sameh Boujelbene, vice president and analyst at Dell’Oro Group, described the 1.6 Tbps per GPU number as a significant architectural milestone on the grounds that it mirrors the bandwidth needs of next-generation AI systems.
Storage and campus expansion
CoreWeave also announced new AI Object Storage capabilities, including cross-region write acceleration and an Archive tier, ROI-NJ reported. The company’s Local Object Transport Accelerator employs managed caching on each CoreWeave Kubernetes Service node to serve reads at local NVMe speeds, cutting latency by a factor of eight against a conventional storage cluster and yielding up to 7 GB/s of throughput per GPU. Cécile Robert-Michon, director of internal infrastructure at Cohere, said the unified dataset footprint across regions with reads cached locally means nothing waits on the network. Cross-region write acceleration allows data to be written to a local destination while copying to a second region behind the scenes. The Archive tier is intended for cheaper long-term retention of checkpoints, datasets and model versions, with no charge to fetch, remove ahead of schedule or access data held in the Archive tier.
Blockfusion has signed a 15-year anchor lease with CoreWeave for its Niagara Falls AI campus, Konsulteer reported on 17 September. The lease provides a long-duration infrastructure commitment around which additional capacity can be deployed, pairing specialised GPU deployments with dedicated data-centre environments capable of supporting high-density computing.