News report
GPT-6 Astra Ultrafast Runs on NVIDIA Blackwell at Up to 8x Token Speed
OpenAI's GPT-6 Astra Ultrafast is live in the API, ChatGPT Work and Codex, with up to 8x faster token generation in Codex and Blackwell-backed inference.
On this page
Astra Ultrafast is now available beyond the DevDay announcement
OpenAI has made GPT-6 Astra Ultrafast available through the API and to eligible ChatGPT Work and Codex users. OpenAI describes Ultrafast as its premium speed tier for latency-sensitive work, with token generation reaching up to 8x Astra Standard speed in Codex at about 300 tokens per second and up to 6x in the API.
The speed figures describe token generation, not an 8x reduction in complete task time. Coding and agent workflows can still spend substantial time on reasoning, tool calls, builds, tests, network requests and other work between generated responses.
| Area | Confirmed position | Evidence boundary |
|---|---|---|
| Codex generation | Up to 8x faster than Astra Standard; OpenAI cites about 300 tokens/s | Maximum token-generation claim, not end-to-end task speed |
| API generation | Up to 6x faster | OpenAI claim; workload and request behavior still matter |
| Compute platform | NVIDIA Blackwell GPUs | NVIDIA disclosure about the deployed inference platform |
| Availability | API plus eligible ChatGPT Work and Codex users | Access depends on product and plan eligibility |
NVIDIA says Blackwell is the hardware behind the speed tier
NVIDIA says GPT-6 Astra Ultrafast runs on Blackwell GPUs and that OpenAI has continued optimizing inference around the architecture. The hardware disclosure adds an important implementation detail that OpenAI's speed-tier description alone does not establish.
NVIDIA quotes OpenAI inference and compute leaders describing model-assisted work on high-performance GPU kernels and ongoing inference optimization. Those statements explain the engineering direction, but they are not independent benchmark evidence comparing Blackwell with another accelerator.
Why the speed tier matters most for agent loops
The practical case for Ultrafast is strongest when generation latency appears repeatedly inside a loop: an agent writes code, invokes a tool, reads the result and decides what to do next. Faster decoding can shorten each model-response segment and make interactive coding or tool-use sessions feel more responsive.
It does not change the underlying distinction between model capability and serving speed. Astra Ultrafast is a faster serving tier for Astra rather than evidence of a new model with different reasoning quality, and users should evaluate whether generation latency is actually the bottleneck in their workload before treating the headline multiplier as a workflow multiplier.
Sources
Primary and technical sources
These sources support the reporting and analysis above. Current stories are updated when later evidence materially changes the facts.