Accelerating GPT-5.6 Sol Ultrafast

Aug 14, 2026 01:10 AM - 2 hours ago 1

Today, Cerebras and OpenAI are sharing an early look astatine Ultrafast Mode, a caller work tier launching first successful the OpenAI API and powered by Cerebras. Ultrafast is disposable initially to a prime group of customers, pinch entree expanding complete time. Cerebras powers GPT-5.6 Sol connected Ultrafast mode, delivering up to 750 output tokens per 2nd and without immoderate value compromise, allowing Sol Ultrafast to accelerate your astir time-sensitive, mission-critical work.

Frontier Intelligence astatine Unprecedented Speed

AI builders person ever needed to take betwixt velocity and intelligence. As models standard up successful size and intelligence, they incur higher computational and information activity costs, slowing down consequence times. Users often request to hold for high-quality results aliases judge inferior results wrong a shorter timeframe.

GPT-5.6 Sol Ultrafast resolves this tradeoff, bringing frontier intelligence to products and workflows wherever each 2nd matters. Compared pinch output speeds reported by Artificial Analysis GPT-5.6 Sol connected Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 connected Fast mode.

At Cerebras, we put Ultrafast to the trial by moving it head-to-head pinch celebrated models connected Humanity's Last Exam. HLE is simply a challenging exemplary benchmark that consists of 2,500 questions typically answerable only by those holding PhDs successful fields specified arsenic chemistry, economics, and literature.

In our evaluations, GPT-5.6 Sol connected Ultrafast mode answered each 2,500 HLE questions successful 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, much than 3 days of continuous compute, to get astatine the aforesaid conclusions. In different words, Ultrafast worked done the frontier of quality knowledge successful a azygous moving day, achieving comparable accuracy astir 7× faster.

Benchmarking was performed by Cerebras utilizing GPT 5.6 Sol Ultrafast pinch Codex connected xhigh reasoning connected July 10 and Claude Fable 5 pinch Claude Code connected xhigh reasoning connected July 13-15.

As exemplary capabilities proceed to advance, the scope of applications for accelerated conclusion expands. GPT-5.6 Sol is OpenAI’s champion exemplary yet for ineligible briefs, financial models, and engineering reports. On GDP-Val, a benchmark for economically valuable knowledge activity tasks, Ultrafast delivered a 5.6x end-to-end speedup pinch nary value degradation, showing really faster conclusion tin accelerate economically valuable work.

Benchmarking was performed by Cerebras connected July 31 2026 utilizing GPT 5.6 Sol and GPT 5.6 Sol Ultrafast connected mean reasoning wrong Codex.

High-Speed Intelligence Powers High-Stakes Work

Faster intelligence changes what’s imaginable for individuals and organizations. With Ultrafast, you tin now put agents connected the captious way of problems wherever each 2nd counts.

"With GPT-5.6 Sol Ultrafast, Cerebras enables AI that keeps up pinch really you think, code, and collaborate. We’re excited to spot really workflows and applications are transformed by Ultrafast inference."

Rohan Varma

Product astatine OpenAI

Ultrafast is simply a persistent separator for organizations utilizing frontier AI to quickly respond to incoming information. Companies operating web services tin leverage Ultrafast to root-cause and reside accumulation outages, preserving customer trust, preventing mislaid revenue, and redeeming downtime minutes against their SLAs. And successful adversarial, precocious stakes cyberattacks, Ultrafast is an invaluable instrumentality for information teams who must quickly observe and respond to bad actors to incorporate catastrophic losses.

More broadly, Ultrafast enables wholly caller modes of moving pinch agents, it delivers real-time insights and updates, truthful you don’t person to context-switch crossed aggregate parallel sessions to get the astir retired of your agents.

"Whereas formerly I mightiness person to hold a mates minutes for a task to finish, it now finishes for maine earlier I moreover person the opportunity to context-switch. It makes maine measurement much productive."

Jeffrey Wang

OpenAI Researcher

With Ultrafast, researchers and engineers tin reserve their attraction for going heavy connected prime problems that matter most, while continuing to usage Standard processing for parallelizing commodity tasks. Cerebras is excited to powerfulness the adjacent activity of AI innovation, raising the ceiling for what individuals and organizations tin execute pinch responsive AI.

Breakneck Speed is Enabled by Breakthrough Innovation

GPT-5.6 Sol connected Ultrafast mode is powered by Cerebras’ revolutionary Wafer-Scale Engine architecture, purpose-built for frontier AI workloads. Fast frontier conclusion is simply a information activity problem: connected GPUs, conclusion connected ample models is bottlenecked by representation bandwidth, arsenic exemplary weights must beryllium many times transferred betwixt on-chip representation and off-chip retention to make successive tokens wrong a exemplary response.

Cerebras takes a contrarian attack to eliminating this inefficient information movement: we battalion 44 GB of SRAM connected each wafer-sized chip. Weights enactment on-chip, and tokens travel uninterrupted done exemplary layers pipelined crossed wafers. This method attack scales smoothly pinch exemplary size, paving the measurement for a continued velocity advantage connected early frontier models.

Ultrafast: Now successful Limited Preview

GPT-5.6 Sol on Ultrafast mode is disposable successful a constricted preview coming to a prime group of customers. Access will grow arsenic capacity grows. Sign up for updates.

More