In recent years, we have become accustomed to a fairly accepted idea in the world of artificial intelligence: faster models tend to be less intelligent than those that take longer to respond. These “lightweight” options work very well in terms of cost and latency for many applications, but when reasoning is critical, usually go a notch higher.
Now, in the race for leadership in the development of artificial intelligence, something unusual has just happened. Gemini 3 Flash, Google’s new model, outperformed GPT-5.2 Extra High, OpenAI’s higher reasoning variant, in several performance tests. And it forces us to reconsider some of the rules we’ve taken for granted.
A fast model that also has reasons. Google’s new model comes with a very specific promise: to demonstrate that “speed and scalability don’t have to come at the expense of intelligence.” While it was designed with efficiency in mind, both in terms of cost and speed, Google insists that Gemini 3 Flash also excels at comprehension tasks.
According to the company, the model can adjust your thinking skills. It is capable of “thinking” longer if the use case requires it, but also uses an average of 30% fewer tokens than Gemini 2.5 Pro, measured with typical traffic, to perform a wide range of tasks with high accuracy and without a response time penalty.
The truth is in the benchmarks. Are the tests perfect? No. But they are still one of the most useful tools we have for comparing AI models, pitting them against each other, and determining in which scenarios they perform better or worse. And in this area Gemini 3 Flash shows itself well.
In SimpleQA Verified, a test that measures reliability in knowledge questions, the Gemini 3 Flash scores 68.7%, compared to 38.0% for the GPT-5.2 Extra High. In multimodal reasoning under MMMU-Pro, Google’s model scores 81.2%, compared to OpenAI’s 79.5%. Video-MMMU Flash achieves 86.9% compared to 85.9% for GPT-5.2 Extra High.

When we look at multilingual and cultural capabilities, Flash is again ahead, with 91.8% compared to 89.6% for GPT-5.2 Extra High. In the Common Sense Global PIQA across 100 languages, the difference remains: 92.8% for Flash vs. 91.2% for the OpenAI model. All indications are that Gemini 3 Flash is specifically optimized for capturing nuances beyond the English language and thinking more freely in a global context.
It is also great at using tools and agents. In Toolathlon, Flash scores 49.4% compared to GPT-5.2 Extra High’s 46.3%. In the FACTS Benchmark Suite, the difference is significant, but still in favor of Google: 61.9% vs. 61.4%. In long-running tasks, the Flash tool appears to be more consistent.
But he is not the king of pure reasoning. Now you should see the full photo. While Gemini 3 Flash outperforms OpenAI’s best model in several benchmarks, if you’re looking for “pure” reasoning, the balance shifts. In the most demanding tests in the field, the GPT-5.2 Extra High continues to set the benchmark.
The OpenAI model leads the visual puzzle-oriented ARC-AGI-2 with 52.9% compared to Flash’s 33.6%. In AIME 2025, it reaches 100% when executing the code, compared to 99.7%. And in SWE-bench Verified, aimed at software development, it scores 80.0%, compared to 78.0% for the Gemini 3 Flash.
What exactly is GPT-5.2 Extra High. Throughout the article, the name GPT-5.2 Extra High comes up several times, and it’s normal to wonder if it’s something new or little-known. It’s not really the model that the general public is usually told about.
Google uses this name in its comparison table to denote the maximum level of reasoning available in the OpenAI API for GPT-5.2 Thinking and Pro. It is listed as “xhigh” in the official OpenAI documentation.
Where you can use Gemini 3 Flash. Access to Gemini 3 Flash is country independent. If you have access to the Gemini program, you’re already using this model, which has become standard. It also reaches developers through API, AI Studio and Vertex AI. In the United States, the rollout goes even further, with Gemini 3 Flash becoming the standard model for Google’s search engine’s AI mode.

Cost of using Gemini 3 Flash. For those who want to integrate Gemini 3 Flash into their applications, the model costs $0.50 per million input tokens and $3 per million output tokens. This is a slight increase from Gemini Flash 2.5, which was $0.30 per million tokens entered and $2.50 per million tokens issued.
An increasingly tense race. Gone are the days when Google tried to stand up to ChatGPT and Bard, or when OpenAI seemed to be years ahead of the rest. Today, the distance between the big players in AI has shrunk dramatically. The competition is more direct, more technical and, above all, much closer.
Images | Google
In Xataka | Amazon Prepares 10 Billion Investment in OpenAI Because If You Can’t Beat the Enemy, It’s Best to Join Them

