
Google's New AI Chip: A Deep Dive into Frozen v2
Google is building a chip designed for one job: running Gemini faster and cheaper. The chip, codenamed Frozen v2, was reported by The Information on Monday and gives Google a potential answer to a problem it can't spend its way out of fast enough: It is running out of capacity to serve the AI demand it has already generated.
In March, Google told Meta it couldn't fill the volume of Gemini compute Meta wanted to purchase. Meta had to instruct employees to ration their AI usage. Google—spending up to $190 billion on AI infrastructure this year—was turning away customers because it didn't have enough servers to serve them.
So now it's building a chip designed only for its AI models. This strategic pivot underscores an urgent reality: even the deepest corporate pockets have limits when it comes to scaling AI infrastructure. Google's capital expenditure on AI has skyrocketed from $32 billion in 2023 to an estimated $190 billion in 2026, yet demand continues to outstrip supply. The company's data centers are running near full capacity, and the energy costs of operating millions of GPUs are becoming a significant financial and environmental burden.
There’s not much information about this new chip, but by naming convention it’s not another upgrade to Google's Tensor Processing Units (TPUs)—the custom chips Google has been building since 2015 that power Gemini and its Cloud services for outside developers. Those Tensor chips run any AI model loaded onto them. Frozen v2 does something different. Per the reports, it bakes part of Gemini's architecture—the structural blueprint that determines how the model routes and processes information—directly into the hardware.
In machine learning, "freezing" means locking something permanently in place. Here, what gets frozen is the architecture, not the model's weights (the actual knowledge Gemini picks up through training, which stays updatable). By hardwiring this blueprint into the chip's circuits, the chip skips redundant calculations and stops shuttling data across memory on every query. Engineers project a six to ten times improvement in tokens—the small text chunks that make up each AI response—generated per watt of electricity consumed.
That's the difference between Google serving ten queries for the power cost of one. At Google's scale, a 6–10x efficiency gap isn't abstract—it's billions of dollars. The company currently operates hundreds of thousands of Nvidia GPUs, each consuming hundreds of watts. If Frozen v2 can deliver even a fraction of its projected efficiency gains, Google could significantly reduce its energy bills and carbon footprint while simultaneously increasing the number of AI queries it can handle.
If you use Gemini, Frozen v2 won't change how it feels to you. But it changes what it costs to run—and a cheaper-to-run Gemini competes harder against OpenAI, Anthropic, and Chinese labs that already account for up to 45% of U.S. company AI token usage, largely because they run 60–90% cheaper. You may not have cheaper AI, but Google will likely be more profitable.
Alphabet shares climbed roughly 3% during Monday's session on the news, touching $356 intraday. The company reports Q2 2026 earnings on Wednesday, July 22, and the pump receded in today’s session as investors wait for Google’s most recent results. The stock's performance reflects growing investor confidence in Google's ability to maintain its competitive edge in the AI arms race, despite headwinds from rising infrastructure costs and regulatory scrutiny.
This is yet another effort by a major AI company to kill its over-reliance on Nvidia hardware to develop its products. Nvidia controls roughly 85% of the GPU market for AI, and every major tech company wants out. Nvidia's dominance has created a single point of failure in the AI supply chain: if Nvidia faces production delays or allocates chips to competitors, companies like Google are left scrambling for alternatives.
Nvidia's hardware was originally built for video games, not language models—it works, just with overhead that purpose-built chips don't carry. At Google's scale, a 6–10x efficiency gap isn't abstract. It's billions of dollars. Meta, Amazon, Microsoft, and OpenAI all have custom silicon programs for exactly that reason. Meta's MTIA chips are designed for recommender systems and AI inference. Amazon's Trainium and Inferentia chips power its AWS cloud AI services. Microsoft is developing its own AI chips codenamed "Athena" to reduce dependence on Nvidia. OpenAI is working with Broadcom on custom AI accelerators.
As Decrypt reported in March, even AWS—which committed to deploying 1 million Nvidia GPUs through 2027—is building its own chips simultaneously to cut that long-term exposure. The trend is clear: the era of one-size-fits-all AI hardware is ending. Custom silicon is becoming a competitive necessity, not just a cost-saving measure.
Frozen v2 is still exploratory. Key design decisions aren't finalized, Google hasn't confirmed the project exists, and the chip won't be offered to outside Cloud customers—hardware hardwired for one model can't run anyone else's. Deployment is targeted for 2028 at the earliest, according to reports. The long timeline reflects the immense complexity of designing and manufacturing a custom ASIC (Application-Specific Integrated Circuit) from scratch. Google will need to verify the chip's performance across a wide range of workloads, ensure compatibility with its existing software stack, and scale production to meet the demands of its global data centers.
In the meantime, Google is paying SpaceX $920 million a month to rent 110,000 Nvidia GPUs from xAI's data centers as a bridge. This interim solution highlights the extraordinary lengths Google is willing to go to secure compute capacity. The deal, reportedly struck in early 2026, gives Google access to some of the most advanced Nvidia H200 and B200 GPUs installed at xAI's massive data center cluster in Memphis, Tennessee. While expensive, the arrangement allows Google to continue serving Gemini customers while Frozen v2 development proceeds.
The broader implications of Google's chip strategy extend beyond finance. By developing a chip optimized for a single architecture, Google may be able to achieve performance levels that general-purpose GPUs cannot match. For example, Google's earlier TPU designs already outperformed commodity GPUs on certain machine learning tasks by a factor of 1.5 to 2. If Frozen v2 hits its targets, that advantage could grow to an order of magnitude, giving Google a significant moat in AI inference.
Moreover, the chip could have knock-on effects for the AI ecosystem. If Google can offer Gemini at lower prices thanks to reduced compute costs, it could pressure competitors to slash their own pricing, potentially accelerating adoption of AI across industries. On the flip side, a highly specialized chip may lock Google into a specific model architecture, making it harder to pivot to new AI paradigms that emerge before 2028. The bet is that Gemini's architecture will remain relevant for years to come.
From a technical standpoint, Frozen v2 represents a shift from general-purpose AI accelerators to domain-specific processors. Similar to how Google's Pixel Visual Core chip optimized for camera processing in smartphones, Frozen v2 is designed from the ground up to handle the unique mathematical operations of transformer models—the backbone of modern large language models. This includes specialized matrix multiplication units, efficient memory hierarchies, and dedicated hardware for attention mechanisms. By eliminating the overhead of supporting arbitrary models, Frozen v2 can devote more silicon area to the operations that matter most for Gemini.
The development also raises questions about the future of Google's Cloud business. Currently, Google Cloud offers TPUs as a service to external customers for training and inference. If Frozen v2 is restricted to internal use only, Google may lose some external revenue, but the savings from running its own flagship product could more than compensate. Alternatively, Google could eventually offer a limited version of Frozen v2 to select partners under tight NDA, though no such plans have been announced.
In the long run, Google's investment in custom silicon could reshape the semiconductor landscape. As more companies design their own chips, the traditional merchant semiconductor model may face disruption. Companies like Nvidia, AMD, and Intel could see their dominant positions eroded by a wave of customized alternatives. For now, however, Nvidia's lead in AI remains formidable, and any challenger will need years to catch up.
Google's Frozen v2 project is a high-stakes bet that custom hardware is the key to winning the AI race. With a 2028 target and billions in potential savings on the line, the company is going all in on a chip that could redefine how AI is served at scale. Whether it succeeds will depend not only on engineering execution but also on the unpredictable evolution of AI models themselves.
Source:Decrypt News
