Gemini 3.5 Pro, the never-ending delay: why Google’s flagship model still hasn’t shipped

On May 19, 2026, on the Google I/O stage, Sundar Pichai unveiled Gemini 3.5 Pro and asked the audience to be patient until the following month for general availability. Almost two months later, the release date has slipped more than once, and as of mid-July Google’s flagship model is still stuck in limited preview on Vertex AI. It’s one of the most closely followed stories of summer 2026 — not just because of the delay itself, but for what it reveals about the real difficulty of shipping state-of-the-art agentic models.

A timeline of a release that keeps slipping

The initial target was June. By the end of the month, though, prediction markets like Polymarket had already priced the odds of a June 30 launch down to marginal levels, a clear sign that consensus had shifted toward a delay. Google subsequently confirmed the slip to July, citing the need to gather more feedback from early enterprise testers. As of this writing, the model remains accessible only through an allowlist for select enterprise accounts on Vertex AI, still carrying “preview” in its model name. Some sources point to July 17 as the new target date, but that’s a figure reported by third parties, not officially confirmed by Google through a model card, pricing, or API documentation.

What’s behind the delay

There isn’t a single cause behind the slowdown. Google’s stated reason centers mainly on token consumption in agentic workflows: a model built to reason over very long contexts and chain together multiple tool calls can burn through tokens fast enough to make the cost per task hard to justify for many enterprise use cases, and Google evidently prefers to fine-tune this before launch rather than ship a product with an unpredictable bill attached.

There’s also a human factor that’s drawn a lot of attention: during the week of June 21-27, several senior researchers from DeepMind’s coding team left Google, with four of them moving to Anthropic within just six days, adding to earlier departures toward Meta and OpenAI. That talent drain has compounded the technical challenge of closing the performance gap with rival models on agentic and long-context benchmarks, at a moment when every week of delay translated into more media attention for the competition.

The market fallout

The combination of these two signals — the flagship delay and the talent exodus — didn’t go unnoticed by investors. On June 22, 2026, Alphabet shares dropped roughly 5% in a single session, wiping out an estimated $225 billion in market value — the stock’s heaviest intraday drop in more than a year. It wasn’t triggered by any single negative announcement, but by the market reading a technical delay and a signal of internal instability within the team behind the company’s flagship product as one combined story.

What the model promises, once it arrives

Delays aside, the specs that have circulated so far remain notable. Gemini 3.5 Pro is built around a 2-million-token context window, double that of Claude Opus 4.8 and most currently shipping competitors — the largest announced so far by any frontier model. The most anticipated feature is Deep Think, a deep-reasoning mode that, based on information gathered so far, will be reserved for the $250-a-month Ultra plan, while the standard Pro plan at roughly $22 a month will likely offer only the base feature set. On developer pricing, only unofficial estimates are circulating, in the range of $12-15 per million input tokens and $36-60 per million output tokens, based on previous generations rather than confirmed figures.

Meanwhile, Google hasn’t stood still

The flagship delay hasn’t erased Google’s other achievements during the same window. On June 22 the company launched Gemini 2.5 Pro with Deep Think, which posts solid results on reasoning benchmarks — 82.4% on GPQA Diamond and 89.8% on MMLU-Pro — giving the company a positive narrative to pair with the flagship’s delay. Meanwhile Gemini 3.5 Flash, the lighter, faster variant launched in May, is already generally available with a 1-million-token context window and performance that, according to several independent analyses, already surpasses the previous Pro generation on a number of coding and tool-use tasks — making the gap left by the flagship model even more obvious.

What this means if you’re evaluating Gemini for your product

The Gemini 3.5 Pro saga offers a practical lesson for anyone planning an architecture built around a single vendor: a few weeks’ delay on a frontier model is manageable, but once that model becomes critical infrastructure for a product, every week of waiting turns into a real risk for the roadmap. The safer approach in a situation like this is to build your product’s architecture around a stable interface rather than a specific model, so you can switch providers once general availability lands without having to rewrite the application — and in the meantime, use already-available alternatives, like Gemini 3.5 Flash or competing models from Anthropic and OpenAI, to keep development moving.

On ModelHive we’re continuing to track Gemini 3.5 Pro’s actual release date, and we’ll update this piece as soon as Google publishes an official model card with confirmed pricing and benchmarks.

Leave a Reply

Your email address will not be published. Required fields are marked *