Gemini 3.5 Flash (minimal)
AvailableGoogle · 2026-05-19 · 1,000,000 tokens
An AI model from Google, suited to a broad range of AI workloads.
Quick Overview
Benchmark Results
Scores from leading benchmark suites.
Performance Metrics
Latency and throughput performance.
Dive Deeper
AI model analysis
Gemini 3.5 Flash (minimal) Review: Fast, Competitive, and Difficult to Verify

- **Where it stands:** Gemini 3.5 Flash (minimal) ranks 99 of 578 on the Artificial Analysis Intelligence Index at 34.9 - **Price:** $3.375 per 1M blended tokens - **Speed:** 248.998 output tokens per second, 0.3s to first token - **Pick it when:** You need fast, general-purpose responses at a price close to GPT-5 and can validate the exact serving endpoint yourself - **Watch out:** Google’s official documentation does not list the minimal variant, so its API identity, limits, and behavior remain unconfirmed
Gemini 3.5 Flash (minimal) is a fast, high-ranking model with an unresolved public-availability question.
Gemini 3.5 Flash (minimal) combines a strong benchmark position with unusually high measured output speed, but developers should verify the model’s public API identity before building around it.
The data brief places Gemini 3.5 Flash (minimal) at position 99 of 578 on the Artificial Analysis Intelligence Index, with a score of 34.9. That position makes it a serious general-purpose candidate, even though the score does not establish superiority for coding, mathematics, tool use, or long-context work.
The measured latency is 0.3 seconds to the first token, while median output speed reaches 248.998 tokens per second. Those figures make the model attractive for interactive applications, response drafting, routing, and other workloads where users notice waiting time.
The main qualification is documentation. Google’s official model directory lists Gemini 3.5 Flash as a stable model with the API alias gemini-3.5-flash, but it does not list Gemini 3.5 Flash (minimal) or gemini-3-5-flash-minimal. The official pricing page also does not provide separate pricing for the minimal variant.
That mismatch means this review evaluates the measured dataset entry while treating the deployment details as unverified.
The model’s best case is fast general inference at a near-mainstream price, not a proven specialist advantage.
Gemini 3.5 Flash (minimal) offers a balanced selection profile, with speed and broad intelligence evidence stronger than the available documentation evidence.
For developers, the practical choice is less about whether the model looks competitive and more about whether the benchmark result maps to a service they can reliably access. The dataset shows a score of 34.9, which is close to several adjacent models. MiMo-V2-Omni records 35, while Claude Opus 4.5 (Non-reasoning), GPT-5 (high), and GPT-5.1 Codex (high) each record 34.7. Kimi K2.6 (Non-reasoning) records 34.6.
That narrow neighborhood suggests Gemini 3.5 Flash (minimal) belongs in a competitive shortlist. It does not show a decisive intelligence lead over the nearby alternatives. Its clearer differentiator is measured speed, since the data brief reports no output-speed value for those adjacent models.
| Decision factor | Gemini 3.5 Flash (minimal) | What the adjacent data suggests |
|---|---|---|
| General intelligence | Competitive | Nearby models sit between 34.6 and 35 |
| Interactive speed | Strong measured result | Comparable output-speed values are not provided |
| Cost position | Mid-market | Kimi K2.6 is cheaper; several others are more expensive |
| Public verification | Unclear | Google documents gemini-3.5-flash, not the minimal variant |
The score is useful for screening. It is not enough to approve production use without endpoint, quota, context, and output-limit checks.
Gemini 3.5 Flash (minimal) is most compelling for latency-sensitive tasks, while specialist capability remains unproven.
Gemini 3.5 Flash (minimal) should be tested first in interactive workloads where fast visible output matters more than a demonstrated specialist score.
A 0.3-second time to first token supports responsive user experiences. A median output speed of 248.998 tokens per second supports quick completion after generation begins. In practice, that profile fits chat interfaces, coding assistants that stream explanations, support triage, structured extraction, and agent steps that must keep a user or another service moving.
The benchmark position gives the result useful context. A rank of 99 of 578 shows that the model is not merely fast. It also sits within a competitive portion of the evaluated field. However, the Artificial Analysis Intelligence Index is a broad signal. It cannot answer whether the model is best for repository-level coding, mathematical reasoning, JSON reliability, visual inputs, tool calling, or long prompts.
The data brief provides no coding-index score for Gemini 3.5 Flash (minimal), while GPT-5 (high) records 37.8 on the coding index. GPT-5 (high) also records 94.3 on the math index, and GPT-5.1 Codex (high) records 95.7. Those figures support a specialist alternative for coding or mathematics, but they do not prove that Gemini performs poorly in either area.
Google’s model documentation does not disclose a confirmed context window, maximum output length, multimodal range, or detailed API parameters for the minimal variant. Evidence is therefore insufficient for workloads that depend on those limits.
Gemini 3.5 Flash (minimal) is reasonably priced for fast general use, but output-heavy workloads can change the value equation.
Gemini 3.5 Flash (minimal) is attractive at $3.375 per 1M blended tokens, provided the workload does not produce unusually expensive output and the measured endpoint is genuinely available.
The input price is $1.5 per 1M tokens, and the output price is $9 per 1M tokens. That split matters because applications that generate long answers, code patches, transcripts, or multi-step agent messages will feel output pricing more strongly than short classification or extraction tasks.
The model’s blended price is close to GPT-5 (high) and GPT-5.1 Codex (high), both listed at $3.4375 per 1M blended tokens. Kimi K2.6 (Non-reasoning) is listed at $1.7125000000000001, so it is the clearer cost reference for teams whose primary objective is lower token spend. MiMo-V2-Omni and Claude Opus 4.5 (Non-reasoning) are listed at $15 and $10 per 1M blended tokens, respectively.
| Workload pattern | Likely cost reading |
|---|---|
| Short inputs and short outputs | Speed may matter more than small price differences |
| Long generated answers | The $9 output rate deserves close testing |
| High-volume budget routing | Kimi K2.6 provides a lower-cost adjacent reference |
| Specialist coding or math | A higher-priced model may be justified if quality reduces retries |
The official Google pricing page confirms prices for gemini-3.5-flash, including Standard, Batch, Flex, and Priority tiers. It does not confirm a separate price for the minimal variant. The dataset price should therefore be treated as an evaluation reference, not as a verified procurement quote.
Choose Gemini 3.5 Flash (minimal) for fast general tasks after verifying access, and avoid making it the sole specialist model without task tests.
Gemini 3.5 Flash (minimal) is a good candidate for latency-sensitive general workloads, but its unresolved documentation status makes staged validation essential.
The strongest case is an application that benefits from quick first output and quick completion, has moderate general reasoning needs, and can tolerate provider or endpoint uncertainty during evaluation. Examples include conversational product features, support classification, lightweight coding assistance, content transformation, and agent routing. The benchmark score of 34.9 and the measured speed of 248.998 output tokens per second support this starting position.
The case weakens when the workload depends on confirmed context limits, maximum output size, supported modalities, or stable API naming. Google’s model directory identifies gemini-3.5-flash as the stable API alias, while the minimal name is absent. The pricing documentation likewise lacks a minimal-specific entry. Developers should not infer that the documented model and dataset entry are identical.
| Recommendation | Reason |
|---|---|
| Start a controlled trial | The rank and speed justify evaluation |
| Keep a fallback model | Public identity and limits are not confirmed |
| Test coding and math separately | Adjacent models have stronger specialist evidence |
| Confirm billing before launch | Minimal-specific pricing is not documented |
The final decision should come from representative prompts, structured-output checks, tool-call tests, failure analysis, and a verified production endpoint. The current evidence supports shortlisting Gemini 3.5 Flash (minimal), not unconditional adoption.
What developers still need to verify before adoption
Gemini 3.5 Flash (minimal) requires endpoint and capability verification because the available official documentation names a different stable alias.
The benchmark data answers questions about relative position, measured latency, throughput, and reference pricing. It does not answer whether the minimal variant is a publicly documented Google API model, whether it has the same capabilities as gemini-3.5-flash, or whether its limits remain stable over time.
That evidence gap should shape the evaluation plan. First, confirm that the exact model identifier can be called through the intended provider. Next, record the returned model metadata and supported parameters. Then test the prompts that represent the real product, including malformed inputs, long outputs, structured responses, tool calls, and retry behavior.
Developers should also separate model quality from service quality. A fast benchmark result can lose its practical value if quotas, regional availability, rate limits, or output truncation create production failures. None of those deployment details are established in the supplied brief.
The safest conclusion is conditional: Gemini 3.5 Flash (minimal) is worth testing for fast, general-purpose inference, but the current public evidence is insufficient to treat it as a fully verified production target.
Frequently asked questions
Is Gemini 3.5 Flash (minimal) an officially documented Google API model?
Gemini 3.5 Flash (minimal) is not confirmed as an officially documented Google API model because Google’s model directory lists gemini-3.5-flash but does not list the minimal name or alias.
Is Gemini 3.5 Flash (minimal) fast enough for interactive applications?
Gemini 3.5 Flash (minimal) appears well suited to interactive applications because the supplied data reports 0.3 seconds to first token and 248.998 median output tokens per second.
Is Gemini 3.5 Flash (minimal) cheaper than nearby alternatives?
Gemini 3.5 Flash (minimal) is cheaper than MiMo-V2-Omni and Claude Opus 4.5 on the supplied blended-price references, but Kimi K2.6 is listed at a lower blended price.
Should developers use Gemini 3.5 Flash (minimal) for coding?
Gemini 3.5 Flash (minimal) deserves coding tests because Google describes the documented Flash family for agent and coding tasks, but the supplied data does not provide a coding score for the minimal variant.
What is the main adoption risk?
The main adoption risk is identity and capability uncertainty, since the supplied official sources do not confirm the minimal API alias, context window, output limit, modalities, parameters, or separate price.
Sources
- Gemini API ModelsVerifying Google’s documented model name, stable status, API alias, and the absence of a confirmed minimal variant and capability details.
- Gemini API PricingVerifying Google’s documented Gemini 3.5 Flash pricing, product description, and the absence of separate minimal-variant pricing.
- Artificial AnalysisAttribution for the supplied benchmark rank, intelligence score, latency, output speed, blended price, and adjacent-model data.
Published: