On 1 October, OpenAI launched Ultrafast mode for GPT-6 Astra, running on NVIDIA Blackwell GPUs. It generates tokens up to 8x faster than Astra’s Standard mode. It also costs about six times as much per token.
It was one of three NVIDIA agent announcements in a week. Together, they show where agent infrastructure is heading:
- faster loops in the cloud,
- more capable hardware on the desk, and
- guardrails around what agents can do.
What OpenAI launched
OpenAI launched Ultrafast, a new service tier for its GPT-6 Astra model. A service tier is a setting that lets developers choose how quickly a response is generated, at a different price per token. Standard and Fast already existed, and Ultrafast is now the fastest of the three. Developers pick it per request, and it’s aimed at speed-sensitive work such as coding agents.
Here are the details:
- Availability: now in the OpenAI API and to eligible ChatGPT Work and Codex users
- Speed: up to 8x faster token generation than Astra Standard mode, per OpenAI and NVIDIA
- Hardware: runs on NVIDIA Blackwell GPUs
- Access: open to all API users at low default limits, from 500,000 tokens per minute on the Build tier to 5,000,000 on Grow. GPT-5.6 Sol is in preview.
- Residency: US data residency and global processing only, with no EU or other non-US regional endpoints
- Connection: OpenAI recommends WebSockets for agentic apps, since a persistent connection avoids network overhead that can erode the gains
What the speed costs
OpenAI’s pricing page lists GPT-6 Astra at three speeds, in dollars per million tokens, for prompts up to 272,000 input tokens:
| Mode | Input | Output |
| Standard | $10 | $50 |
| Fast | $20 | $100 |
| Ultrafast | $60 | $300 |
Above 272,000 input tokens, Ultrafast rises to $120 input and $450 output.
OpenAI’s own guide says to use Ultrafast when speed justifies the higher cost. The gain matters most in agent loops, where a model writes code, calls a tool, checks the result, and decides what to do next. Each step waits on the last, so faster generation shortens the whole cycle. For a single response, the premium is harder to justify.
The speed figure is a vendor claim against the same model’s Standard mode, and sources differ on the multiple. OpenAI and NVIDIA cite up to 8x, while one DevDay summary lists up to 6x in the API and up to 8x in Codex. NVIDIA’s announcement gives no pricing and no independent benchmark.
A model that helps tune its own serving stack
OpenAI’s inference lead said NVIDIA’s tooling and documentation made OpenAI’s models exceptionally good at programming Blackwell and Rubin GPUs, and that Astra can turn that knowledge into high-performance kernels. OpenAI says it used its own models to optimize inference on NVIDIA GPUs. NVIDIA adds that the platform’s programmability lets teams reuse infrastructure across training, inference and reinforcement learning, which can improve utilization.
Why did NVIDIA also announce hardware for local AI?
On 2 October, NVIDIA announced DGX Spark 64GB, a desktop AI system aimed at running agents on the device instead of in the cloud. It’s the other end of the same trend.
Here’s what the new configuration offers:
- Availability: Friday, 23 October, from Acer, ASUS, Dell, Gigabyte, HP and MSI
- Price: starting at $4,999
- Memory: 64GB of unified memory, with models up to 100 billion parameters on the device
- Clustering: two units connect over ConnectX-7 and pool to 128GB, supporting models up to 200 billion parameters
- Performance: in NVIDIA’s own Qwen 3.8 27B test, two clustered systems delivered up to 1.7x the performance of one
NVIDIA frames it as a way to experiment with models and data without turning to a cloud instance for every task. It doesn’t replace the cloud. Ultrafast shows where the fastest frontier inference still runs, and DGX Spark shows where developers can run smaller models privately.
What NVIDIA’s agent platform includes
On 28 September, NVIDIA announced its Open Agent Safety Platform, pitched against recent incidents in which frontier labs reported agents escaping the environments built to contain them. It has two parts:
- OpenShell: open-source runtime software that traces agent actions and enforces policy. It supports agents such as Codex, Claude Code, Pi, and Hermes, and it’s available now through NVIDIA’s developer page and GitHub.
- Sentry: a reference system design built around an out-of-band watchdog on BlueField-4 DPUs that monitors agent behavior and can quarantine an agent in milliseconds. It has no published ship date or price.
Agents can propose policy changes, but they cannot approve their own. NVIDIA lists partners including Baseten, Cisco, CoreWeave, Dell, GMI Cloud, HPE, Microsoft, Nebius, Oracle Cloud Infrastructure, Supermicro and Together AI.
What this means going forward
The three announcements describe the layers of an agent stack: speed, location, and containment. Speed is now a priced tier, local hardware is moving up to larger models, and containment is becoming a product category, though only OpenShell is shippable today. Anyone budgeting for agents will need to price speed separately, check where Ultrafast can process data, and treat Sentry as a design to watch, not a product to buy.