Google embeds Gemini architecture directly into silicon: Frozen v2 chip announced

Google has officially begun developing a new generation specialized processor — Frozen v2. This is not just another tensor accelerator, but a fundamentally different approach: key elements of the Gemini language model architecture will be integrated directly at the silicon substrate level. This will drastically reduce the volume of data transferred and the number of computational operations when processing user requests.
The project is positioned as a complement to the existing line of tensor processing units (TPUs), not as a replacement. Internal sources confirm that one of the main goals of Frozen v2 is to alleviate the acute shortage of computing power in Google Cloud. Due to resource constraints, the cloud division previously had to turn down a number of contracts with major external clients.
According to preliminary engineering estimates, the new chip will be able to process six to ten times more tokens per unit of energy consumed compared to the company's latest AI accelerators. However, there is an important limitation: Frozen v2 will only be compatible with future versions of Gemini if Google maintains the basic architecture of the model. For now, the project is considered experimental — production volumes will not match those of mass-market TPUs. The chip is scheduled to be put into operation in 2028.
Notably, after news of Frozen v2 emerged, Alphabet (GOOG) shares on the Nasdaq rose by 1.5%.
Meanwhile, Google's AI division is going through a difficult period. The launch of the next version, Gemini 3.5 Pro, is delayed, and the company has lost four leading researchers who moved to competitors — Anthropic and OpenAI. Against this backdrop, Chinese models continue to strengthen their positions: according to the latest data, they account for up to 46% of all tokens processed by American companies.
Expert opinion: Google's decision to embed the Gemini architecture directly into the chip is a logical but extremely risky move. On one hand, this approach could yield a tremendous efficiency boost and reduce dependence on cloud computing power. On the other hand, it tightly ties the hardware to a specific model, which in a rapidly changing paradigm (especially amid the successes of Chinese and open-source competitors) could lead to a technological dead end. Betting on 2028 is a long-term play, and success here will depend not only on engineers but also on whether the Gemini architecture itself remains relevant by that time.