
Google's scale problem, your quality problem
Written by AI · Translated by AI · Read the Swedish original
A cheaper model being 95 percent as capable is an average that does not say where the missing percent goes. You notice it first when an error costs something.
When Demis Hassabis describes Gemini Flash he is not really describing a small model. He is describing a plumbing problem.
Google has AI running through Search, AI Overviews, AI Mode, Maps, YouTube, the Gemini app and more than a dozen products with user counts most companies never see. Every answer costs money. Every delay shows. Every extra second is multiplied across products used by billions.
So when Hassabis says the systems have to be delivered "extremely fast, extremely efficiently and cheaply and with low latency", the key word is not fast. The key word is delivered.
A model in Google's hands is not just an intelligence engine. It is something that has to fit inside a global system without blocking the flow.
Distillation is the attempt to carry behaviour over from a larger, stronger model to a smaller one that is cheaper and faster to run. Hassabis describes that compression into the Flash and Flash Lite models as one of Google's strengths.
The interviewer frames the trade as roughly 95 percent of the capability at a tenth of the price. Hassabis does not confirm the exact figures, but he accepts the shape of the argument: some capability can be traded for lower cost, lower latency and faster iteration.
For Google that trade can be reasonable even when the faster model gives up some capability on certain tasks. If the system is good enough to deliver across enormous products, the gains compound. A slightly weaker model that can be placed everywhere may create more product value than a stronger model that is too slow or too expensive to push through the pipes.
That is good engineering. It is not proof that the same model is right for your hardest workflow.
The problem with "95 percent as capable" is that it is an aggregate claim. It does not say where the missing capability shows up. It may vanish in places you do not care about. It may appear in a corner of the task that matters a great deal. The number alone does not answer that question.
A support tool is a simple example. A Flash model can summarise common tickets well, route most requests correctly and keep the queue moving. That is useful. But if your business depends on catching the unusual ticket where the customer describes the real fault in one vague sentence, the average score is not what you should trust. The question is whether the model catches that case often enough.
This is not about speed causing hallucinations. That is too simple, and the conversation with Hassabis does not establish it. The supported point is narrower: Google balances capability against cost, latency and scale. Flash is shaped by that balance.
Whoever is building has to ask a different question. Not "is this model smart enough in general?" Not "is it fast?" Not "did the demo look clean?" The question is where the model fails under the workload that matters.
That means testing the cases where being wrong costs something. Use the real edge cases. Use the messy inputs. Use the examples that caused refunds, support escalations or angry Slack threads last quarter. Compare the smaller model against the stronger one on those cases, not on the easy traffic where everything looks fine.
If it holds up there, use it. Fast and cheap are not small virtues. But do not borrow Google's optimisation target without checking whether it matches yours.
The model that solves Google's scale problem may not solve your quality problem.
Ask upplyst.ai
Why does it matter?
Flash is optimised for Google's problem: enormous volume, low latency, cheap to run. That is a different problem from catching the unusual case, handling the hard edge case, or keeping a product clean where mistakes have consequences. The model working for Google is not proof it works for your hardest workflow.
What is the background?
In the interview Demis Hassabis describes how Google's Flash models are shaped by a concrete requirement: delivering AI across a dozen products with billions of users, where every answer costs money and every delay shows. Distillation, compressing behaviour from a stronger model into a faster and cheaper one, is central to that strategy.
What is uncertain?
The interviewer frames it as 95 percent of the capability at a tenth of the price. Hassabis does not confirm the exact figures. And the figure itself does not answer where the missing capability actually shows up, which is the only question that matters to someone choosing a model for a specific workflow.