AI Affairs, home

Friday 2 October 2026

Corporate

CoreWeave makes Nvidia Vera Rubin available as Cognition runs production workloads

Cognition reported up to 4.8x higher token throughput against GB200 NVL72 in early tests. CoreWeave also introduced Forge and plans to offer Vera CPUs.

Michael Intrator speaking onstage at Web Summit 2024 against purple lighting
Photo: Web Summit, CC BY 2.0, via Flickr (cropped)

CoreWeave made Nvidia Vera Rubin NVL72 systems with Spectrum-X 102.4T Ethernet available on 30 September, with first production customer Cognition reporting up to a 4.8x increase in token throughput against GB200 NVL72 in early tests, according to Nvidia’s announcement. CoreWeave announced the availability at its Fully Connected conference in San Francisco.

Key points

  • Cognition is running production workloads on Vera Rubin NVL72 and tested its inference performance against GB200 NVL72.
  • CoreWeave plans to offer Vera CPUs for agent environments; its tests found more than 3x faster sandbox startups.
  • CoreWeave launched Forge, combining tools for model development, evaluation and post-training.
  • Chief executive Mike Intrator said orders had not fallen despite regulatory obstacles to data-centre development.

Cognition tests Vera Rubin against GB200

Cognition uses CoreWeave for training, reinforcement learning and production inference for Devin, its AI software engineer. Nvidia said Cognition reached thousands of GPUs on the cloud provider within nine months.

CoreWeave received its first Vera Rubin NVL72 production racks in September. Cognition then tested inference using software engineering tasks sampled from FrontierCode, with agents assigned to solve them. Nvidia said Cognition’s initial SWE-2 inference trials measured total token throughput on Vera Rubin NVL72 at up to 4.8x the GB200 NVL72 baseline.

Silas Alberti, a member of Cognition’s founding team, said long contexts, concurrent requests and token volumes made cost per token consequential for its coding agent. Nvidia said CoreWeave assembled a production Vera Rubin cluster for Cognition in days. Customers can operate capacity through services including CoreWeave Kubernetes Service, Sandboxes and Inference.

Nvidia vice-president Ian Buck said CoreWeave’s V100 GPUs still run customer workloads nearly a decade after the Volta generation launched. His account places the new deployment alongside older equipment still serving customers, rather than replacing it across CoreWeave’s cloud.

Vera CPU rack supports 11,000 environments

CoreWeave also plans to offer Nvidia Vera CPUs for agent workloads. Nvidia said a rack in CoreWeave’s deployment contains 128 CPUs and 11,264 cores, with capacity for more than 11,000 concurrent agent environments when each uses one core. Its Sandboxes service runs isolated environments for agent tool use, reinforcement learning and model evaluation alongside the training jobs they support.

CoreWeave’s tests on Vera CPUs produced agent sandbox startup times more than 3x faster, Nvidia said. CoreWeave also measured a 1.7x performance gain on Terminal-Bench across passing tasks. Those are test results for the CPU offering, while Cognition’s production workload runs on the Vera Rubin NVL72 systems.

CoreWeave Sandboxes is generally available and can provide an isolated environment for each tool call or evaluation, according to Nvidia. The company said the service can run on serverless infrastructure or on infrastructure customers already use for training.

Forge connects CoreWeave’s model-development tools

CoreWeave launched Forge as an environment linking Weights & Biases, OpenPipe’s post-training capabilities and the open-source marimo notebook project. Nvidia said Canva, Capital One and MasterClass are among the first companies building on it. Forge brings production-agent monitoring into the same environment as evaluation and further model training.

The offering includes CoreWeave Agent Lens, which Nvidia says improves failure detection by 20% and fixes issues at half the cost. Nvidia also says Forge’s serverless reinforcement learning trains 1.4x faster at 40% lower cost than a self-managed setup. CoreWeave ARIA and Sandboxes are generally available, while RL Rollouts, which loads new checkpoints into a running deployment, is in private preview.

Elsewhere on CoreWeave’s cloud, home-based care provider Ennoble Care has selected reserved Nvidia RTX PRO 6000 GPU capacity for clinical AI inference, Nvidia said. Ennoble Care serves about 50,000 high-need Medicare patients across 15 states and plans to use agents for documentation, decision support and administrative work.

Intrator sees infrastructure returns improving into 2028

CoreWeave has also raised capital: an earlier AI Affairs report covered its upsized $4.2 billion convertible bond offering at 2.875%. The new systems add to the infrastructure on which CoreWeave aims to earn returns.

Intrator said orders had not declined despite regulatory resistance to data-centre development, Stocktwits reported from a CNBC interview on 30 September. He argued that restrictions would affect where infrastructure is built rather than demand for it. Intrator also said margins on incremental infrastructure were expanding faster than interest rates were rising. He expects its economics to keep improving through 2027 and into early 2028.

Topics: Agents, Chips, Enterprise adoption, Inference