Skip to content

Gemini 3.5 Flash (medium)

Available

Google · 2026-05-19 · 1,000,000 tokens

An AI model from Google, suited to a broad range of AI workloads.

Supported modalities:textimagevideocode

Quick Overview

Text Generation5/10
Code Generation6/10
Reasoning6/10
Multimodal4/10

Benchmark Results

Scores from leading benchmark suites.

artificial analysis intelligence46.7

Performance Metrics

Latency and throughput performance.

P50 Latency
256.042tokens/sec

Dive Deeper

AI model analysis

Gemini 3.5 Flash (medium) Review: Fast, Capable, and Hard to Verify

Gemini 3.5 Flash (medium) Review: Fast, Capable, and Hard to Verify
Summary

- **Where it stands:** Gemini 3.5 Flash (medium) ranks 37 of 578 on the Artificial Analysis Intelligence Index at 45.4 - **Price:** $3.375 per 1M blended tokens - **Speed:** 276.619 output tokens per second, 0.3s to first token - **Pick it when:** You need very fast production responses and already depend on Google Search or Maps grounding - **Watch out:** The `medium` variant is not listed as a separate official Google API model, so its endpoint mapping remains uncertain

01

Gemini 3.5 Flash (medium) in brief

Gemini 3.5 Flash (medium) looks attractive for developers who value response speed, but its exact official identity needs verification before production adoption. The data snapshot places the model at 37 of 578 on the Artificial Analysis Intelligence Index, with a score of 45.4. That position makes it a strong general-purpose model, although it does not establish a clear lead over nearby alternatives.

Google lists gemini-3.5-flash as a Stable model for sustained frontier performance, with emphasis on agentic and coding tasks in the Gemini API model documentation. However, the same official model list does not separately identify Gemini 3.5 Flash (medium) or gemini-3-5-flash-medium. The product name in this review therefore appears to describe a benchmark variant whose direct API mapping is not confirmed.

The practical profile is clear enough to guide an initial shortlist. Gemini 3.5 Flash (medium) combines a high output rate with a moderate blended price. The unresolved model identity is the main adoption risk, not an observed performance failure.

02

The main tradeoff for developers

Gemini 3.5 Flash (medium) offers a compelling speed profile, while nearby models expose a sharper cost or coding-performance advantage. The benchmark position suggests that Gemini 3.5 Flash (medium) belongs in serious evaluation, but not as an automatic default.

Decision factor Gemini 3.5 Flash (medium) Nearby reference point What the difference means
General intelligence Strong position at 37 of 578 Qwen3.7 Max scores 46, while MiniMax-M3 scores 44.4 The model sits close to both stronger and weaker general benchmarks
Coding evidence No coding index value in the snapshot Qwen3.7 Max scores 66, GPT-5.6 Terra (medium) scores 64.7 Coding selection should depend on task testing, because the available evidence is incomplete
Speed 276.619 output tokens per second Qwen3.7 Max reaches 204.156 Gemini has the clearest throughput advantage among these reference points
Commercial fit Google API, grounding, and caching options Other providers offer lower listed blended prices Google ecosystem requirements may justify the premium for some workloads

Google describes gemini-3.5-flash as balancing speed, intelligence, Google Search, and grounding on the Gemini API pricing page. That positioning supports use cases involving Google services, but it does not prove better answer quality for every developer workflow. The available benchmark and official evidence point to a fast, credible candidate with an unusually important verification step.

03

What the ranking and speed mean in practice

Gemini 3.5 Flash (medium) is best understood as a high-throughput general model whose benchmark rank supports broad use, not universal superiority. A position of 37 of 578 indicates substantial competitiveness across the evaluated intelligence index. It does not show that the model will win every coding, extraction, planning, or tool-use test.

The strongest operational signal is output speed. The recorded median is 276.619 output tokens per second, with 0.3s to first token. That combination should benefit interactive developer tools, streaming assistants, code explanation, structured drafting, and applications where users judge quality while the response is still arriving. Fast generation can also reduce the perceived cost of waiting, even when the token price is not the lowest available.

The conclusion changes when the workload is coding-heavy. The data snapshot provides coding index scores for several nearby models, including 66 for Qwen3.7 Max and 64.7 for GPT-5.6 Terra (medium), but it provides no coding index value for Gemini 3.5 Flash (medium). Google’s official documentation highlights agentic and coding tasks, yet the model documentation does not provide an official benchmark score for this specific variant. Developers should therefore run repository-level tests before selecting it for code generation, debugging, or autonomous changes.

The evidence is also insufficient for claims about context capacity, maximum output, multimodal behavior, or failure patterns. Google’s verified pages do not specify those details for this model entry. Teams should treat those capabilities as open validation items rather than assume that the benchmark label covers them.

04

When the price is justified

Gemini 3.5 Flash (medium) is reasonably priced for fast interactive workloads, but it becomes difficult to justify when latency is not a product requirement. The blended price is $3.375 per 1M tokens, with input priced at $1.5 and output priced at $9. Those values create a meaningful premium over the lower-cost nearby references in the data snapshot.

The premium makes sense when output speed, Google integration, or operational simplicity matters more than minimizing token spend. It is especially relevant for applications that can use Google Search or Google Maps grounding. Google lists a shared monthly free allowance of 5,000 grounding queries, then charges $14 per 1,000 queries, according to the Gemini API pricing documentation. The free API tier does not provide Google Search or Google Maps grounding, so production grounding requires a paid tier and quota planning.

Batch and Flex pricing reduce the listed input and output rates to $0.75 and $4.5 per 1M tokens. Priority pricing increases them to $2.7 and $16.2. These modes make the model’s economics depend heavily on workload urgency. Offline processing can improve the case for Gemini 3.5 Flash (medium), while high-priority output can make the premium more pronounced.

Context caching is supported on paid access, but the official material does not provide enough model-specific detail to estimate savings for a particular application. Developers should measure cached and uncached traffic separately. A cheaper model with slower output may be the better choice for large asynchronous jobs, especially when users do not wait for tokens to stream.

05

Recommendation for model selection

Gemini 3.5 Flash (medium) deserves a production trial for latency-sensitive Google-centric applications, but teams should verify its API mapping before committing. The combination of a 45.4 intelligence score and 276.619 median output tokens per second makes it a credible option for assistants, agent interfaces, and responsive developer products.

Choose Gemini 3.5 Flash (medium) when the application benefits from fast streaming, Google Search or Maps grounding, paid-tier caching, or an existing Google API deployment. The model’s position near the top of the evaluated population supports a broad evaluation rather than a narrow niche test.

Prefer a nearby alternative when coding quality is the primary acceptance criterion and a coding benchmark is required for the decision. Qwen3.7 Max and GPT-5.6 Terra (medium) have coding index values in the data snapshot, while Gemini 3.5 Flash (medium) does not. Prefer a lower-cost option when the workload is asynchronous, output-heavy, and independent of Google grounding.

The most important launch gate is identity verification. Google officially lists gemini-3.5-flash as Stable, but does not list gemini-3-5-flash-medium as a separate endpoint in its model catalog. Confirm that the provider’s internal slug maps to the intended Google model, then test representative prompts, tool calls, structured outputs, and safety behavior.

06

Questions to answer before adoption

Gemini 3.5 Flash (medium) should enter evaluation with explicit checks for endpoint identity, coding quality, quotas, grounding access, and workload economics. The benchmark data answers where the model ranks and how quickly it generates, but it does not answer every implementation question.

Developers should confirm the exact API model name, test real application prompts, and validate output limits before production rollout. The absence of official context and output specifications is material evidence uncertainty. It should be recorded as a deployment condition, not silently filled with assumptions.

Frequently asked questions

Is Gemini 3.5 Flash (medium) a good default model for developers?

Gemini 3.5 Flash (medium) is a strong candidate for fast general-purpose applications, but it should not be adopted as a default until the provider confirms that the medium slug maps to Google’s official gemini-3.5-flash endpoint and passes task-specific tests.

Is Gemini 3.5 Flash (medium) suitable for coding tasks?

Gemini 3.5 Flash (medium) may suit coding workflows because Google positions the official Flash model for coding and agentic tasks, but the available snapshot lacks a coding index for this variant, so repository-level evaluation remains necessary.

Why choose Gemini 3.5 Flash (medium) over a cheaper model?

Gemini 3.5 Flash (medium) earns its premium when fast streaming, Google Search or Maps grounding, or Google API integration directly improves the product, while asynchronous workloads may receive better economics from cheaper alternatives.

Does Gemini 3.5 Flash (medium) have a confirmed context window?

Gemini 3.5 Flash (medium) does not have a confirmed context window in the verified official material used here, so developers should test the actual endpoint and document the observed limit before accepting large-context workloads.

Can developers use Google grounding on the free tier?

Developers cannot rely on Google Search or Google Maps grounding through the free tier, because Google’s pricing documentation states that production grounding requires paid access and remains subject to quotas and query charges.

Sources

  1. Gemini API model documentationOfficial model name, Stable status, API alias, positioning, and the absence of a separately listed medium variant.
  2. Gemini API pricingOfficial pricing, free tier limits, Batch and Flex pricing, Priority pricing, context caching, and Google Search or Maps grounding rules.
  3. Artificial AnalysisBenchmark score, ranking, price snapshot, latency, output speed, and adjacent-model comparison data.

Published: