OpenAI launches GPT-6 Astra Ultrafast on NVIDIA Blackwell GPUs
OpenAI has made GPT-6 Astra Ultrafast available in the OpenAI API and to eligible ChatGPT Work and Codex users, with the mode running on NVIDIA Blackwell GPUs. The company says inference optimizations designed for NVIDIA’s Blackwell architecture enable Ultrafast to deliver up to 8x faster token generation than Astra Standard mode.
The speedup is aimed at workflows where repeated response generation affects productivity, including coding agents that write code, use tools, test results and decide next steps. Faster generation can reduce delays during edit-test-debug cycles, shorten the time between tool calls and make interactive applications feel more responsive.
OpenAI is also using its own models to refine inference software on NVIDIA GPUs, taking advantage of the platform’s programmability to test and implement improvements after deployment. OpenAI executives said the collaboration with NVIDIA helped optimize inference and deliver the acceleration behind Astra Ultrafast.
NVIDIA’s programmable platform is positioned as a way for developers and researchers to reuse infrastructure across training, inference and reinforcement learning as models evolve. OpenAI says that flexibility can help teams repurpose compute resources as demand changes, improving utilization and reducing the need to overprovision for individual workloads.