Google and OpenAI have introduced ultra-fast AI models for working with code
8/14/2026, 06:52 AM • Евгения Слив

On August 13th, two major technology corporations simultaneously presented their new ultrafast solutions in the field of artificial intelligence. Google has announced the Gemini 3.7 Flash model, and OpenAI has launched a special Ultrafast mode for its GPT-5.6 Sol neural network. Both developments are focused on completing tasks related to programming and operation of autonomous AI agents as quickly as possible. Google positions its new product as a universal tool for writing code and automating complex business processes. The developers announced noticeable improvements in multi-stage planning, accurate following of instructions and high-quality generation of program code. At the same time, the cost of using the new model has decreased by exactly half compared to the starting price of its predecessor. The entire Flash line is traditionally focused on scenarios where not only the accuracy of the response is crucial, but also the speed of its receipt.
Back in July of this year, Google released the previous version of Gemini 3.6 Flash, which cost a dollar and a half for one million input tokens and seven and a half dollars for a million output tokens. According to the company, this model consumed seventeen percent fewer output tokens compared to version 3.5 of Flash. At the same time, the long-awaited flagship Gemini 3.5 Pro model has not yet entered the open market. Currently, the developers are actively testing this advanced product together with selected partners, and the exact timing of its full release remains unknown.
In response, OpenAI introduced Ultrafast mode, which is capable of operating up to fourteen times faster than the standard version of the neural network. The new modification generates up to seven hundred and fifty output tokens per second through the use of a specialized Cerebras infrastructure. At the start, this mode is available through the software interface only to a limited number of large clients. As computing power gradually increases, OpenAI plans to connect more and more ordinary users to the accelerated version. The company first announced the speed indicator of seven hundred and fifty tokens per second back in June when announcing the basic GPT-5 model.6. At that time, the developers talked about plans to launch Sol on Cerebras chips, but they immediately warned that not everyone would be able to use the accelerated version initially. Earlier in August, the SpaceXAI team also introduced its flagship Grok 4.6 model, focusing on long-term agency tasks and the creation of interactive applications.
