OpenAI’s new Ultrafast mode runs GPT-5.6 Sol 14 times faster, on Cerebras chips

OpenAI is taking a significant step forward in its quest to make its most advanced model also its fastest, with the introduction of Ultrafast, a new tier of its API that can run the GPT-5.6 Sol model up to 14 times faster than before. This breakthrough is made possible by the company's partnership with Cerebras, a wafer-scale chipmaker that has provided the hardware necessary to achieve such impressive speeds. With Ultrafast, the GPT-5.6 Sol model can now produce around 750 output tokens per second, a significant increase from its previous performance.
NEW Tier Introduced
The key to this achievement lies not in the development of a new model, but rather in the innovative way that OpenAI is serving its existing one. By leveraging Cerebras's cutting-edge chips, the company is able to eliminate the latency that has long been a bottleneck in the development of frontier AI. This is a crucial step forward, as it enables the creation of AI systems that can respond in real-time, without sacrificing their intelligence or capabilities. On August 13, OpenAI opened a limited preview of Ultrafast to a small group of customers, with plans to expand access to the service as its capacity allows.
The introduction of Ultrafast represents a significant shift in the way that AI models are developed and deployed. Until now, developers have been forced to make a trade-off between speed and intelligence, with faster models often sacrificing some of their capabilities in order to achieve real-time responses. With Ultrafast, OpenAI is offering a solution that can deliver both frontier-grade reasoning and near-instant answers, without requiring developers to choose between these competing priorities. This is particularly important for the development of agentic software, which is a key area of focus for the entire AI industry.
Real-time Responses
The ability to create AI agents that can respond in real-time is crucial for the development of effective and engaging products. As OpenAI notes, an AI agent that takes thirty seconds to respond to every query is essentially a demo, rather than a fully functional product. In contrast, an agent that can respond in the time it takes to hold a conversation has the potential to be a truly revolutionary tool. With the introduction of Ultrafast, OpenAI is taking a significant step towards making this vision a reality, and it will be exciting to see how the company's customers and partners choose to utilize this powerful new technology.
The implications of Ultrafast are far-reaching, and have the potential to impact a wide range of industries and applications. By providing a platform that can support the development of highly advanced and highly responsive AI systems, OpenAI is helping to pave the way for a new generation of AI-powered products and services. As the company continues to expand access to Ultrafast, it will be interesting to see how developers and businesses choose to leverage this technology, and what kinds of innovative solutions they are able to create. With its powerful new API, OpenAI is once again demonstrating its commitment to pushing the boundaries of what is possible with AI, and to helping its customers and partners achieve their goals.
Source: thenextweb.com · 2026-08-14