A single point now separates Meta from Google at the top of one of the AI industry's most closely tracked composite benchmarks. On version 4.1 of the Artificial Analysis Intelligence Index, Muse Spark 1.1 running in its highest reasoning mode registered a 51. Gemini 3.6 Flash, which Google shipped on July 21, came in at 50. Meta had released Muse Spark 1.1 on July 9, and the lead materialized within days.

The margin is thin enough that reasonable people can argue it's noise. What makes it worth attention is what sits behind each number.

Why Gemini 3.6 Flash Held Its Score Instead of Raising It

Google didn't position 3.6 Flash as a step forward in raw capability. It arrived as an efficiency release — a tuning pass rather than a new ceiling. The clearest evidence is the benchmark itself: 3.6 Flash posted the same Intelligence Index score as Gemini 3.5 Flash, the model it replaces.

That's a deliberate trade, and it fits a pattern Google has been signaling. Tulsee Doshi, senior director of product management in Google's Gemini group, described the company's recent releases as landing in the "sweet spot" between efficiency and power for running AI agents. While competitors push toward larger models, Google has been filling specific gaps with targeted, cheaper alternatives.

The catch is that an efficiency release can't move a capability leaderboard. Holding steady at 50 was the expected outcome of the strategy — it just happened to coincide with a rival clearing 51.

The Cost Story Behind Muse Spark 1.1

Price is where Meta's entry separates itself more decisively than a single benchmark point suggests.

Artificial Analysis put the cost of running Muse Spark 1.1 across the Intelligence Index at roughly $0.26 per task. That works out to about a third of what OpenAI's GPT-5.4 costs at an equivalent reasoning tier.

 

Model

 

 

Input (per 1M tokens)

 

 

Output (per 1M tokens)

 

 

Muse Spark 1.1

 

 

$1.25

 

 

$4.25

 

 

Gemini 3.6 Flash

 

 

$1.50

 

 

$7.50

 

Google priced 3.6 Flash as a cheaper option than its predecessor, and by its own product line it is. Measured against Meta, though, the output pricing gap is wide — $7.50 versus $4.25 per million tokens. For agentic workloads that generate long chains of tokens rather than short answers, output pricing is often the line item that decides deployment. That's the segment both companies are targeting.

How Meta Gained Eight Points in Three Months

Muse Spark 1.1 scored eight points higher than the model it succeeded, and it did so over a window of roughly three months. Artificial Analysis attributed most of that jump to a single change: hallucination rates fell from 73 percent to 38 percent.

Training the Model to Abstain

The mechanism matters more than the delta. Meta reduced hallucinations by training the model to abstain — to decline when it lacks grounding rather than produce a confident guess.

It's an unglamorous fix, and it reads less like a breakthrough in reasoning than a correction in behavior. But on a composite index that penalizes confident wrong answers, cutting the error rate nearly in half moves the aggregate score substantially without requiring a fundamentally more capable model underneath. A model that knows when to stop scores better than one that always answers.

Muse Spark 1.1 is already accessible in Thinking mode through the Meta AI app and website, and the company expects it to eventually take over from the Llama models currently powering chatbots across WhatsApp, Instagram, Facebook, and its smart glasses.

Gemini 3.5 Pro: Google's Missing Flagship

Google's answer to all of this is a model nobody outside a partner program can use yet.

Gemini 3.5 Pro was originally promised for June. It missed that target, slipped to mid-July, and has since been pushed toward the end of the month. Google's public comment has been limited — the company has said it is "currently testing 3.5 Pro" with partners, and little beyond that.

What the Pro Model Is Expected to Bring

The specifications circulating for 3.5 Pro describe a different class of product entirely:

  • A two-million-token context window
  • A "Deep Think" reasoning mode
  • Pricing in the range of $15 per million input tokens and $60 per million output tokens

That pricing sits roughly an order of magnitude above both Muse Spark 1.1 and Gemini 3.6 Flash. It's not competing for the same workloads. Whether it eventually reclaims the top of the Intelligence Index is a separate question from whether developers waiting on it will keep waiting.

What the Leaderboard Position Actually Means

Strip away the ranking theater and the current situation is unusual for reasons that have nothing to do with one benchmark point.

Meta is selling a paid API model that, on at least one composite index, outperforms the best option Google has publicly available. For a company whose AI reputation was built on open-weight Llama releases, competing on paid API terms — and winning a leaderboard slot while doing it — is new ground.

Google's position is more nuanced than the score suggests. It shipped what it intended to ship: a cheaper, more efficient Flash model that holds the previous capability line. The gap it hasn't closed is at the top of the stack, where 3.5 Pro was supposed to be sitting by now and isn't.