OpenAI made GPT-6 Astra Ultrafast available through its API on 1 October, running the model on Nvidia Blackwell GPUs, according to Nvidia’s announcement. Eligible ChatGPT Work and Codex users can also access it. Nvidia says Ultrafast generates tokens up to 8x faster than Astra Standard mode, a comparison of the speed at which responses are produced rather than the time taken by an entire task.
Key points
- GPT-6 Astra Ultrafast runs on Nvidia Blackwell GPUs and is available in the OpenAI API.
- Nvidia claims token generation is up to 8x faster than in Astra Standard mode.
- OpenAI says it used its own models to optimise the software that runs inference on Nvidia GPUs.
Nvidia’s 8x comparison for Astra Ultrafast
Token generation is the production of successive pieces of a response. Nvidia describes why speed matters when an application waits between actions: a coding assistant drafts a change, calls a tool, examines the outcome and picks a further action. The company says faster generation can shorten those cycles and the waits between tool calls, as well as make interactive applications more responsive.
Writing and checking code could involve less waiting between a change, a check of its result and the next change if responses arrive faster at each pass. That would depend on how much of the cycle is spent generating responses rather than carrying out the other parts Nvidia describes.
Nvidia’s up-to-8x figure compares token generation in Ultrafast with Astra Standard mode. The company presents the coding-cycle benefit as something faster generation can deliver, rather than giving an elapsed-time result for a complete edit-test-debug cycle.
Tillet and Ruddarraju describe the GPU work
OpenAI’s inference lead, Philippe Tillet, said its models were exceptionally good at programming Blackwell and Rubin GPUs, and that Astra could turn that knowledge into high-performance kernels. A kernel is a program that directs work on a GPU. Tillet said Astra could apply what the models had learnt about programming the hardware to write kernels for demanding inference work.
Uday Ruddarraju, OpenAI’s chief technology officer of compute, said the company used internal models to optimise inference on Nvidia GPUs. Nvidia describes that work as a continuing process of testing and improving the software used to serve the model. It also says its programmable platform allows computing resources to be reused across training, inference and reinforcement learning as models change.
Astra Ultrafast joins OpenAI’s GPT-6 releases
The Ultrafast release follows OpenAI’s GPT-6 Sol and Luna models, which AI Affairs reported had lower API prices and fewer factual errors. AI Affairs also reported the launch of GPT-6.1 Sol at one-fifth of Astra’s token price. Those releases concern different models or pricing, while Nvidia’s comparison for Ultrafast concerns generation speed against Astra Standard mode.
OpenAI separately cancelled a planned GPT-6.1 Astra release after internal safety concerns, AI Affairs reported. Nvidia directs developers seeking access, pricing and implementation details for the available Astra Ultrafast option to an Ultrafast guide.