Gemini 3.5 Flash lands on Google’s Android coding rankings, but it’s 3x the cost for slower performance
Google’s Android coding benchmark puts Gemini 3.5 Flash at 6th place, behind older models and at a much higher cost.
Intelligence analysis by GPT-5.4 Mini

Google’s latest Android Bench shows Gemini 3.5 Flash falling short on Android coding: it ranks sixth, trails Gemini 3.1 Pro Preview on score, and uses far more tokens per run. The benchmark suggests Google’s newer Flash model is not the efficiency win its name implies for this specific task.
Google made a scoreboard for AI models that try to fix Android code. Gemini 3.5 Flash did not do as well as some older helpers, and it also used a lot more money and computer work, like a car that burns extra gas but still arrives behind.
Analysis
What Google measured
Google updated its Android Bench rankings, which compare how well different models solve Android coding cases across 10 runs. The benchmark gives each model a score out of 100, along with average latency, token usage, and estimated cost per run.
Where Gemini 3.5 Flash lands
Gemini 3.5 Flash ranks 6th in the latest list, behind GPT 5.5, GPT 5.4, Gemini 3.1 Pro Preview, and two Claude models. Google had positioned 3.5 Flash as a faster, cheaper alternative to Gemini 3.1 Pro, with an expected performance gap of 6.1%. The benchmark results in this article point the other way for Android development: 3.5 Flash shows a 9% gap in performance success and higher latency.
The cost comparison is the sharpest part of the story. Google says Gemini 3.5 Flash used an average of 355.9 tokens and cost about $147.1 per benchmark run, while Gemini 3.1 Pro Preview used 73.3 tokens at about a third of that cost. The article notes that the comparison is to a preview version of Gemini 3.1 Pro, but the result still undercuts the idea that Flash is the better fit for Android coding.
Broader read
The ranking also shows that the top of the Android coding leaderboard has stayed mostly stable, with only a few changes in the lineup. Google appears to be using Android Bench as a practical signal for model quality in software work, especially as companies push more agentic and coding-focused AI tools. The takeaway here is narrow but clear: Gemini 3.5 Flash may be useful in other tasks, but Android coding is not where it shines.
Key points
- Google updated Android Bench to compare AI models on Android coding tasks.
- Gemini 3.5 Flash ranked 6th, behind GPT and Claude models and Gemini 3.1 Pro Preview.
- The article says 3.5 Flash used far more tokens and cost more per run than Gemini 3.1 Pro Preview.
- Google’s benchmark is meant to show how models perform on real Android development work.
- The story frames Android coding as a weak spot for Gemini 3.5 Flash.
If Google uses this benchmark to improve future models, Android coding tools could become cheaper and more reliable over time. The public ranking also gives developers a clearer way to choose models based on real task performance instead of marketing claims.
The result suggests Gemini 3.5 Flash may not deliver the speed-and-cost gains Google implied for coding work. If that holds across more tests, developers could end up paying more for a model that performs worse on Android tasks than its older sibling.



