For about sixty minutes on August 1, people running blind prompt battles in LMSYS Chatbot Arena were unknowingly testing a model Google hasn't released yet. It showed up in the active rotation under the identifier gemini-3.5-pro, took real prompts from real testers, and then disappeared. The automated bots that track Arena's model roster caught the removal roughly an hour after the first sighting.

Asked to explain what it was, the model gave the kind of non-answer that staged models usually give: a large language model built by Google. No version boasting, no capability rundown. Just enough to confirm whose infrastructure it was running on.

The One-Hour Window in Arena's Testing Rotation

Arena's blind battles are the point here. Testers don't choose which models they're comparing — they submit a prompt, get two anonymous responses, and vote. That setup makes an unannounced model slipping into the pool genuinely interesting, because whoever put it there wasn't looking for hype. They were looking for data.

What makes the timing harder to write off as a glitch is what happened alongside it. Brief micro-outages and API disruptions rippled across Google's developer infrastructure during the same window. Taken together, the pattern reads like backend staging work — new weights being wired into serving infrastructure, tested against live traffic, then rolled back before anyone official had to comment on it.

A Launch Timeline That Keeps Slipping

The sighting lands after months of missed targets, which is why a one-hour appearance got as much attention as it did.

 

Milestone

 

 

What happened

 

 

I/O 2026 (May)

 

 

Google announces Gemini 3.5 Pro; Sundar Pichai points to the following month

 

 

June

 

 

The stated target passes with no release

 

 

July

 

 

The revised window also passes with no release

 

 

July 20

 

 

Logan Kilpatrick says the model is in partner testing and the team hopes to land it soon

 

 

August 1

 

 

The model appears in Arena's rotation and is pulled within the hour

 

Pichai's original framing at I/O was about as concrete as a launch signal gets from a CEO on stage — next month, not next quarter, not later this year. That June date slipped to July. July came and went. By July 20, the message from Logan Kilpatrick, product lead at Google DeepMind, had shifted to partner testing and a hope that it would land soon, which is a noticeably softer commitment than a month with a number attached to it.

Three Flash Models Shipped, No Pro

In the meantime, Google did ship. Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber all arrived while 3.5 Pro stayed in limbo. Tulsee Doshi, senior director of product management in Google's Gemini group, positioned those releases around efficiency and power together — targeted models built for running AI agents rather than for winning benchmark headlines.

That's a defensible strategy on paper. It also isn't what the developers making the most noise are waiting for.

What Developers Actually Want From 3.5 Pro

The expected upgrades over the current Gemini 3.1 Pro cluster in three areas: agentic coding, reasoning, and multimodal performance. Each one maps to work the Flash tier struggles with.

  • Agentic coding. Multi-step tasks where the model plans, executes, checks its own output, and adjusts — the workload that exposes weak reasoning fastest.
  • Reasoning depth. Complex software engineering problems that need the model to hold a lot of context and logic at once instead of pattern-matching to a plausible answer.
  • Multimodal handling. Work that mixes code, images, and documents in a single task.

The frustration developers keep voicing is specific rather than general: Flash models are fast and cheap, and they hold up fine on straightforward tasks, but they fall short once complex software engineering enters the picture. Cost efficiency stops mattering when the output needs rewriting. There's also the 2-million-token context window attached to the Pro model, which is the kind of spec you can't approximate with a lighter tier no matter how efficient it is.

The Gemini 4 Problem

Google has confirmed it's already begun pre-training Gemini 4. That detail changes the pressure on 3.5 Pro in an awkward way. Every month the model sits unreleased, it moves closer to arriving as something that already feels like last generation — announced at an event most of a year earlier, benchmarked against a landscape that kept moving, and shipped into a cycle where its successor is already in training.

A delayed model can still be a great model. But it has to clear a higher bar to feel like a release rather than a backlog item being closed out.

Why an Arena Sighting Is Worth Paying Attention To

Arena testing has historically functioned as one of the last stages before a public rollout. Models get staged there, collect blind head-to-head votes against the current field, and then launch. It isn't a guarantee — plenty gets tested that never ships on the timeline observers expect — but it's a meaningfully stronger signal than a leak or a config string found in an app teardown.

Combined with the infrastructure disruptions on the same day and Kilpatrick's partner-testing comment two weeks earlier, the sequence points in one direction. Whether that means days or another slipped month is the part nobody outside Google DeepMind can answer right now, which is exactly why announcements from that team are getting read closely.