Developers can now run OpenAI's smartest model 14× faster via a new tier

By Emeka Briggs
Tweet image from @Nairametrics

OpenAI is previewing Ultrafast, a Cerebras-powered API tier that runs GPT-5.6 Sol at up to 750 tokens per second. It is open only to select customers.

Share

OpenAI began previewing Ultrafast, a new API tier that runs its GPT-5.6 Sol model up to 14× faster than standard processing, on 13 August 2026, according to the company's announcement. The tier is launching first in the OpenAI API to a select group of customers, with access expected to expand as capacity grows.

The speed jump is the headline number, but the more concrete figure for developers is throughput: Ultrafast can generate up to 750 output tokens per second from GPT-5.6 Sol. That is the metric OpenAI and its hardware partner, Cerebras, are both citing to explain what the tier is for.

The speed jump, in numbers

Ultrafast is powered by Cerebras hardware, and Cerebras confirmed in a press release that the tier runs GPT-5.6 Sol at up to 750 output tokens per second and up to 14× faster than the same model on standard processing. OpenAI frames the tier as a way to bring its most intelligent model into products and workflows where response time matters, from incident response to live experimentation.

The tier is separate from the existing Fast mode. Fast mode for GPT-5.6 Sol previously offered up to 2.5× faster speeds than standard processing at twice the price, per the OpenAI API changelog. Ultrafast's 14× claim puts it in a different performance class, and it runs only on Sol, the frontier-grade variant OpenAI describes as its most capable model in the GPT-5.6 family.

For Nigerian and wider African developer teams, that difference matters practically. A tier that cuts inference time from seconds to fractions of a second could change what feels possible inside a fintech fraud-detection loop, a customer-support copilot, or a real-time market analytics dashboard. But right now that possibility is mostly theoretical for most builders.

Limited access, no published price

Ultrafast is currently available only in limited preview to a select group of API customers. As of the latest coverage up to 16 August 2026, OpenAI has not published a rate card or price for the tier, and there is no committed general availability date. Analysts at Digital Applied and ExplainX both note that the speed multipliers are vendor-reported claims without disclosed baseline workloads or third-party verification. Developers should treat them as marketing ceilings, not guaranteed performance.

The access and pricing gaps are significant for African builders for a second reason. Ultrafast is hosted entirely within OpenAI's stack on Cerebras hardware, which means teams using it remain dependent on US-based AI infrastructure. That raises the same questions African startups already face around data residency, latency from Lagos or Nairobi, and cost structure that can shift without local input.

What changes for business-facing AI

The Nairametrics amplification of the news emphasises business workflows: financial research, incident response, customer support, commerce, and live experimentation. That maps directly onto sectors where Nigerian banks, brokers, telcos, and e-commerce platforms operate. If Ultrafast pricing eventually becomes viable for enterprise use, the tier could raise competitive pressure on local AI platforms and regional model providers, who would need to compete on localisation, regulation, or cost.

"Powered by Cerebras, Ultrafast generates up to 750 tokens per second, bringing our most intelligent model to products and workflows where every second counts," OpenAI said in its announcement.

There is no named African regulator or startup publicly tied to Ultrafast yet. The ecosystem impact is prospective, not documented. But for African AI researchers and platform teams making infrastructure decisions, the absence of independent verification matters as much as the speed claim.

What to watch next: OpenAI's capacity rollout, and whether Fast mode pricing becomes the reference point once Ultrafast gets a public rate card. Until then, the tier remains a preview that most builders can read about but not touch. The developer community thread on OpenAI's forum and coverage from Help Net Security are the best early signals on how real the 14× figure turns out to be.

Share this article

Help others discover this story

https://www.techblit.com/developers-can-now-run-openais-smartest-model-14-faster-via-a-new-tier