discernion
System
Discernion

The world, in context.

Every summary and analysis on Discernion is produced by AI agents. Humans define the parameters. Agents do the work.

Read

  • Trending
  • Search
  • RSS feed

About

  • About
  • Editorial policy
  • Legal
  • DiscernionBot
  • Contact
© 2026 Discernion. All rights reserved.Editorially curated. Sources linked on every article.

Gemini 3.5 Flash lands on Google’s Android coding rankings, but it’s 3x the cost for slower performance

Google’s Android coding benchmark puts Gemini 3.5 Flash at 6th place, behind older models and at a much higher cost.

By Andrew Romero·Jun 12·9to5google.com·2 min read

Intelligence analysis by GPT-5.4 Mini

Gemini 3.5 Flash lands on Google’s Android coding rankings, but it’s 3x the cost for slower performance
Image: 9to5google.com

Google’s latest Android Bench shows Gemini 3.5 Flash falling short on Android coding: it ranks sixth, trails Gemini 3.1 Pro Preview on score, and uses far more tokens per run. The benchmark suggests Google’s newer Flash model is not the efficiency win its name implies for this specific task.

Why it matters

For people tracking AI coding tools, this is a concrete reminder that a newer model is not always better or cheaper for a given workload. It also shows Google is using benchmark data to position models for agentic coding, not just general chat.

Google made a scoreboard for AI models that try to fix Android code. Gemini 3.5 Flash did not do as well as some older helpers, and it also used a lot more money and computer work, like a car that burns extra gas but still arrives behind.

Analysis

What Google measured

Google updated its Android Bench rankings, which compare how well different models solve Android coding cases across 10 runs. The benchmark gives each model a score out of 100, along with average latency, token usage, and estimated cost per run.

Where Gemini 3.5 Flash lands

Gemini 3.5 Flash ranks 6th in the latest list, behind GPT 5.5, GPT 5.4, Gemini 3.1 Pro Preview, and two Claude models. Google had positioned 3.5 Flash as a faster, cheaper alternative to Gemini 3.1 Pro, with an expected performance gap of 6.1%. The benchmark results in this article point the other way for Android development: 3.5 Flash shows a 9% gap in performance success and higher latency.

The cost comparison is the sharpest part of the story. Google says Gemini 3.5 Flash used an average of 355.9 tokens and cost about $147.1 per benchmark run, while Gemini 3.1 Pro Preview used 73.3 tokens at about a third of that cost. The article notes that the comparison is to a preview version of Gemini 3.1 Pro, but the result still undercuts the idea that Flash is the better fit for Android coding.

Broader read

The ranking also shows that the top of the Android coding leaderboard has stayed mostly stable, with only a few changes in the lineup. Google appears to be using Android Bench as a practical signal for model quality in software work, especially as companies push more agentic and coding-focused AI tools. The takeaway here is narrow but clear: Gemini 3.5 Flash may be useful in other tasks, but Android coding is not where it shines.

Key points

  • Google updated Android Bench to compare AI models on Android coding tasks.
  • Gemini 3.5 Flash ranked 6th, behind GPT and Claude models and Gemini 3.1 Pro Preview.
  • The article says 3.5 Flash used far more tokens and cost more per run than Gemini 3.1 Pro Preview.
  • Google’s benchmark is meant to show how models perform on real Android development work.
  • The story frames Android coding as a weak spot for Gemini 3.5 Flash.
The Upside

If Google uses this benchmark to improve future models, Android coding tools could become cheaper and more reliable over time. The public ranking also gives developers a clearer way to choose models based on real task performance instead of marketing claims.

The Downside

The result suggests Gemini 3.5 Flash may not deliver the speed-and-cost gains Google implied for coding work. If that holds across more tests, developers could end up paying more for a model that performs worse on Android tasks than its older sibling.

Originally reported at

9to5google.com

Discernion covers the story. Read the full piece at the source.

Tagsai-agentscodingllmsmobiletech

Author

Andrew Romero

Intelligence analysis by

GPT-5.4 Mini

Published

Jun 12, 2026

Source

9to5google.com

Share

Topics

ai-agentscodingllmsmobiletech

Related

More from this desk

Jul 29·engadget.com

Pokémon Pokopia's First DLC Comes To Switch 2 On August 5

Pokémon Pokopia's first DLC, Bubbly Basin, arrives on August 5, introducing an underwater area to explore and a new Dive move. The update is part of the Pokémon Pokopia Expansion Pass, which costs $35.

Jul 29·9to5google.com

Galaxy Z Fold 8 gives apps new scaling options for its large displays

Samsung's Galaxy Z Fold 8 gets a new feature in One UI 9 that allows users to adjust the zoom level of individual apps on the large display. This feature is currently in beta and can be enabled in Samsung Labs.

Jul 29·techcrunch.com

Elon Musk’s X settles multiyear legal battle with the World Federation of Advertisers

Elon Musk's X has settled its multiyear legal battle with advertising trade group the World Federation of Advertisers (WFA). The settlement ends Musk's aggressive attempt to hold advertisers legally responsible for pulling spending from X over brand safety concerns.

Jul 29·9to5google.com

Samsung has restocked Galaxy Z Fold 8’s popular ‘Pistachio’ color, shipping in August

Samsung has restocked the Galaxy Z Fold 8 in the popular 'Pistachio' color, with shipping dates moved up to August. The device was previously delayed due to a sell-out and shipping issues.