Google Building Frozen v2 Chip to Hardwire Gemini
Google is building a server chip called Frozen v2 that hardwires part of Gemini’s architecture into silicon and could raise tokens-per-watt efficiency 6–10×; shares rose about 3%.
Google is developing a server chip codenamed Frozen v2 that embeds part of the Gemini model’s architecture into physical circuits, a design engineers say could improve tokens-per-watt efficiency by six to ten times. Alphabet shares rose about 3% after reports of the project.
Frozen v2 is being designed specifically to run Gemini faster and at lower cost by locking the model’s structural blueprint into hardware while keeping the model’s weights updatable. In machine learning, “freezing” an architecture means fixing the routing and processing steps in place; the chip would skip repeated calculations and reduce data movement across memory during each query.
The design differs from Google’s existing Tensor Processing Units, which are general-purpose accelerators that can run many models. The new chip would embed the model’s processing layout in silicon, which engineers estimate would allow many more tokens per unit of power compared with general accelerators.
Work on Frozen v2 is exploratory. Key design choices are not finalized and Google has not publicly confirmed the project. Deployment is targeted as early as 2028, and the company does not plan to offer the hardware to external cloud customers because a chip optimized for one model cannot run other models effectively.
The effort responds to sharply higher demand for AI compute that has strained Google’s server capacity. The company has limited some sales of compute to customers and is spending heavily on AI infrastructure this year, including short-term contracts for external GPUs while custom chips are developed. Reports indicate Google is paying hundreds of millions of dollars per month to rent large blocks of Nvidia GPUs from outside data centers as a bridge until internal hardware is available.
Engineers involved with the project say a six- to tenfold improvement at scale would reduce power and data-center costs for running Gemini. The company’s plan would keep Gemini’s user experience unchanged while lowering the operational cost of serving queries.
Nvidia currently supplies the majority of GPUs used for AI workloads, and several large technology companies are pursuing internal silicon to lower long-term reliance on third-party GPUs. Frozen v2, if realized, would represent a chip specialized for a single model architecture: the routing and processing blueprint would be fixed in circuits while model weights can still be updated.
As development continues, timelines and performance estimates remain provisional. Investors awaited Alphabet’s quarterly results on July 22 as market reaction to the reports eased after an initial share rise.
The material on GNcrypto is intended solely for informational use and must not be regarded as financial advice. We make every effort to keep the content accurate and current, but we cannot warrant its precision, completeness, or reliability. GNcrypto does not take responsibility for any mistakes, omissions, or financial losses resulting from reliance on this information. Any actions you take based on this content are done at your own risk. Always conduct independent research and seek guidance from a qualified specialist. For further details, please review our Terms, Privacy Policy and Disclaimers.







