Vigyata.AI
Is this your channel?

Google Proved Longer AI Thinking = Worse Answers

472 views· 11 likes· 13:41· Mar 3, 2026

🛍️ Products Mentioned (2)

Google's new research proves AI models that think longer actually get WORSE answers, not better. I break down the data, show you what "deep thinking" really means, and reveal how to get the same accuracy at half the cost. ☎️ Do you need any career or technical help? Book a call with me: https://calendly.com/mg_cafe Reference paper added to discord channel: https://discord.gg/2kcjQFMCr5 ****************** LET'S CONNECT! ******************* Join Discord Channel: https://discord.gg/2kcjQFMCr5 ✅ You can contact me at: LinkedIn: https://www.linkedin.com/in/mohammad-ghodratigohar/ Email: mo.ghodrati95@gmail.com Twitter: https://twitter.com/MG_cafe01 🔔 Subscribe for more cloud computing, data, and AI analytics videos by clicking on the subscribe button so you don't miss anything. #AIResearch #GoogleAI #deepthinking

About This Video

In this video I break down a pretty uncomfortable result from Google: with every extra dollar you spend on “reasoning” or long chain-of-thought tokens, you can actually make the model’s answers worse. And this isn’t my opinion—Google shows a negative correlation between answer length (token count) and accuracy across multiple models and scientific/math tasks. So the old mental model of “longer thinking = smarter” just doesn’t hold the same way for recent systems, because overthinking can amplify flawed heuristics and waste compute on uninformative generation. Then I switch the framing from “how long” to “how deep.” I explain what Google calls a deep thinking token: a token that needed more internal work across layers before the model’s prediction stabilized. They quantify this by comparing predictive distributions across layers using Jensen–Shannon divergence (JSD). From there, they define DTR (Deep Thinking Ratio): what fraction of the generated tokens are truly “deep.” The practical win is simple: sample multiple candidate answers, compute DTR early (like after the first ~50 tokens), reject low-DTR continuations, and finish the answer with the highest DTR. The takeaway for builders is clear: stop rewarding verbosity, measure mechanistic effort, and use early rejection to boost accuracy while cutting cost—often around 50%.

Frequently Asked Questions

🎬 More from MG