Google is reportedly developing a new AI server chip designed specifically to run its Gemini family of models more efficiently, marking another major step in the race to optimize artificial intelligence infrastructure.
According to Reuters, citing The Information, the new chip is internally referred to as Frozen v2 and could significantly reduce the computational cost of serving Gemini while easing the company’s growing AI capacity constraints.
TL;DR
- Google is reportedly developing a Gemini-focused AI server chip codenamed Frozen v2.
- The chip could deliver 6x to 10x better token-per-watt efficiency than Google’s latest TPUs.
- It is expected to complement, not replace, Google’s Tensor Processing Units.
- The project aims to reduce AI inference costs and address internal compute shortages.
- Deployment is reportedly targeted for 2028, although the design remains under development.
Google’s latest reported hardware initiative reflects a broader industry shift toward tightly integrating AI software with custom silicon.
Rather than building another general-purpose accelerator, Frozen v2 is said to embed portions of Gemini’s architecture directly into the chip, allowing Google’s infrastructure to execute AI inference workloads with significantly greater efficiency.
According to Tom's Hardware, engineers working on the project estimate that the processor could deliver between six and ten times more AI tokens per unit of power than Google’s newest custom AI chips.
Such an improvement could substantially lower operational costs while enabling Google to serve a larger number of Gemini requests using the same amount of computing infrastructure.
The reported project arrives as Google faces increasing pressure to scale its AI infrastructure.
Reuters reported that shortages in AI computing capacity have reportedly created internal friction and even forced Google Cloud to decline certain customer opportunities because of limited available compute resources.
A Google Cloud spokesperson did not confirm the existence of Frozen v2 but acknowledged the company’s continued investment in hardware and software co-design.
“Our teams are constantly researching and experimenting with new innovations. By co-designing our hardware and software from the ground up, we ensure our systems are integrated and highly optimized,” the spokesperson told Reuters.
Importantly, Frozen v2 is not expected to replace Google’s Tensor Processing Unit roadmap.
Instead, reports indicate it would become a separate family of specialized processors dedicated to Gemini inference, while TPUs continue handling broader AI training and inference workloads across Google Cloud.
The development also follows Google’s rollout of its Ironwood TPU generation and its growing investment in vertically integrated AI infrastructure.
By designing both the AI models and the silicon they run on, Google aims to reduce dependence on third-party hardware while improving performance and energy efficiency across its AI services.
The report comes shortly after Reuters reported that Google delayed the release of Gemini 3.5 Pro because the model had not yet met internal quality targets, particularly in coding performance.
Topics for more insights:
Together, the reported model improvements and specialized hardware investments indicate that Google is pursuing both software and infrastructure enhancements to strengthen its position in the increasingly competitive AI market.
While Frozen v2 reportedly remains several years away from deployment, with current expectations pointing to 2028, the initiative highlights how the AI race is rapidly expanding beyond model capabilities into purpose-built hardware designed specifically for frontier AI systems.





















