News report

GPT-6 Astra Ultrafast Runs on NVIDIA Blackwell at Up to 8x Token Speed

OpenAI's GPT-6 Astra Ultrafast is live in the API, ChatGPT Work and Codex, with up to 8x faster token generation in Codex and Blackwell-backed inference.

On this page
  1. Astra Ultrafast is now available beyond the DevDay announcement
  2. NVIDIA says Blackwell is the hardware behind the speed tier
  3. Why the speed tier matters most for agent loops

Astra Ultrafast is now available beyond the DevDay announcement

OpenAI has made GPT-6 Astra Ultrafast available through the API and to eligible ChatGPT Work and Codex users. OpenAI describes Ultrafast as its premium speed tier for latency-sensitive work, with token generation reaching up to 8x Astra Standard speed in Codex at about 300 tokens per second and up to 6x in the API.

The speed figures describe token generation, not an 8x reduction in complete task time. Coding and agent workflows can still spend substantial time on reasoning, tool calls, builds, tests, network requests and other work between generated responses.

What OpenAI and NVIDIA currently claim for Astra Ultrafast
AreaConfirmed positionEvidence boundary
Codex generationUp to 8x faster than Astra Standard; OpenAI cites about 300 tokens/sMaximum token-generation claim, not end-to-end task speed
API generationUp to 6x fasterOpenAI claim; workload and request behavior still matter
Compute platformNVIDIA Blackwell GPUsNVIDIA disclosure about the deployed inference platform
AvailabilityAPI plus eligible ChatGPT Work and Codex usersAccess depends on product and plan eligibility

NVIDIA says Blackwell is the hardware behind the speed tier

NVIDIA says GPT-6 Astra Ultrafast runs on Blackwell GPUs and that OpenAI has continued optimizing inference around the architecture. The hardware disclosure adds an important implementation detail that OpenAI's speed-tier description alone does not establish.

NVIDIA quotes OpenAI inference and compute leaders describing model-assisted work on high-performance GPU kernels and ongoing inference optimization. Those statements explain the engineering direction, but they are not independent benchmark evidence comparing Blackwell with another accelerator.

Why the speed tier matters most for agent loops

The practical case for Ultrafast is strongest when generation latency appears repeatedly inside a loop: an agent writes code, invokes a tool, reads the result and decides what to do next. Faster decoding can shorten each model-response segment and make interactive coding or tool-use sessions feel more responsive.

It does not change the underlying distinction between model capability and serving speed. Astra Ultrafast is a faster serving tier for Astra rather than evidence of a new model with different reasoning quality, and users should evaluate whether generation latency is actually the bottleneck in their workload before treating the headline multiplier as a workflow multiplier.

Sources

Primary and technical sources

These sources support the reporting and analysis above. Current stories are updated when later evidence materially changes the facts.

  1. 01 OpenAI

    DevDay 2026 recap — Ultrafast availability and speed
  2. 02 NVIDIA

    How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast