Gemini 3.5 Flash (medium) vs GPT-5 mini (high): The Ultimate Performance & Pricing Comparison
Deep dive into reasoning, benchmarks, and latency insights.
The Final Verdict in the Gemini 3.5 Flash (medium) vs GPT-5 mini (high) Showdown
The current catalog does not contain complete performance evidence for both models, so this page does not declare an overall winner. Use the available fields as comparison signals and validate the models on your own workload.
Model Snapshot
Key decision metrics at a glance.
Machine-readable comparison data
| Model | Metric | Value | Unit | Source / snapshot |
|---|---|---|---|---|
| Gemini 3.5 Flash (medium) | Reasoning | 6.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5 mini (high) | Reasoning | 9.0 | benchmark or capability score | Artificial Analysis · current catalog |
| Gemini 3.5 Flash (medium) | Coding | 6.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5 mini (high) | Coding | 2.0 | benchmark or capability score | Artificial Analysis · current catalog |
| Gemini 3.5 Flash (medium) | Multimodal | 4.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5 mini (high) | Multimodal | 2.0 | benchmark or capability score | Artificial Analysis · current catalog |
| Gemini 3.5 Flash (medium) | Long Context | 6.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5 mini (high) | Long Context | 3.0 | benchmark or capability score | Artificial Analysis · current catalog |
| Gemini 3.5 Flash (medium) | Blended Price / 1M tokens | $3.375 | USD per 1M tokens | Artificial Analysis · current catalog |
| GPT-5 mini (high) | Blended Price / 1M tokens | $0.688 | USD per 1M tokens | Artificial Analysis · current catalog |
| Gemini 3.5 Flash (medium) | P95 Latency | — | milliseconds | Artificial Analysis · current catalog |
| GPT-5 mini (high) | P95 Latency | — | milliseconds | Artificial Analysis · current catalog |
| Gemini 3.5 Flash (medium) | Tokens per second | 276.619 | tokens per second | Artificial Analysis · current catalog |
| GPT-5 mini (high) | Tokens per second | — | tokens per second | Artificial Analysis · current catalog |
Data provided by Artificial Analysis; live values use the current catalog.
Overall Capabilities
This radar chart visually maps the core capabilities (reasoning, coding, math proxy, multimodal, long context) of `Gemini 3.5 Flash (medium)` vs `GPT-5 mini (high)`.
Benchmark Breakdown
This grouped bar chart provides a side-by-side comparison for each benchmark metric.
Speed & Latency
Lower time to first token is better; higher tokens per second is better.
The Economics of Gemini 3.5 Flash (medium) vs GPT-5 mini (high)
Pricing Breakdown
Compare input and output pricing in USD per 1M tokens.
Real-World Cost Scenario
Per run: 1M input tokens + 250k output tokensGemini 3.5 Flash (medium)$3.75
GPT-5 mini (high)$0.75
GPT-5 mini (high) costs $3 less per run
Gemini 3.5 Flash (medium) vs GPT-5 mini (high): Which Model Should Developers Choose?
This article is a dated snapshot published on 2026-08-07. Live cards above use the current catalog; missing live fields are not inferred.

- Winner overall: Gemini 3.5 Flash (medium), with an Artificial Analysis Intelligence Index of 45.4 vs 25.3
- Cheaper: GPT-5 mini (high) at $0.6875 vs $3.375 per 1M blended tokens
- Faster: Gemini 3.5 Flash (medium) at 276.619 median output tokens per second
- Pick Gemini 3.5 Flash (medium) when: You need the stronger measured general intelligence score and Google Search or Maps grounding
- Watch out: The named variants do not map cleanly to current official model catalogs, and several capability claims remain unverified
Gemini 3.5 Flash (medium) vs GPT-5 mini (high)
Gemini 3.5 Flash (medium) is the stronger measured general-purpose option, while GPT-5 mini (high) is the safer cost choice only if its identity and access are confirmed. The data snapshot gives Gemini an Artificial Analysis Intelligence Index of 45.4, compared with 25.3 for GPT-5 mini (high). Data provided by https://artificialanalysis.ai/ reports the comparison values, but the research evidence does not establish a complete capability comparison.
The larger issue is model identity. Google’s official catalog lists gemini-3.5-flash as Stable, with the stable API alias gemini-3.5-flash, but it does not separately list Gemini 3.5 Flash (medium) or gemini-3-5-flash-medium. Gemini API model documentation confirms the official name and status while leaving the exact mapping unresolved.
OpenAI’s current model catalog does not list gpt-5-mini or the GPT-5 mini (high) display name. OpenAI Models therefore cannot confirm the model’s current API identity, supported parameters, context window, or availability. Developers should treat this comparison as a decision aid for a labeled evaluation entry, not as proof that the two named variants are directly callable production endpoints.
Executive summary for developers
Gemini 3.5 Flash (medium) wins the available quality signal, while GPT-5 mini (high) wins the available price comparison. Gemini scores 45.4 on the Artificial Analysis Intelligence Index, while GPT-5 mini (high) scores 25.3. Data provided by https://artificialanalysis.ai/ supplies these measurements.
The evidence does not show which model is better at coding or mathematics in a direct comparison. The snapshot contains a GPT-5 mini (high) coding index of 15.6 and math index of 90.7, but it contains no corresponding Gemini values. That absence prevents a defensible coding or mathematics winner. It also means developers should not infer that Gemini’s higher intelligence score predicts better repository editing, debugging, or mathematical reliability.
Gemini has a clearer official product story. Google describes gemini-3.5-flash as a model balancing speed and intelligence, with Google Search and grounding support. Gemini API pricing documents that positioning and the related billing rules. GPT-5 mini (high) has no comparable verified current pricing or positioning in the supplied OpenAI pages. OpenAI Pricing does not list it.
For a production shortlist, Gemini is the better-supported candidate. For a cost-sensitive experiment, GPT-5 mini (high) deserves testing because its supplied blended price is lower, but only after the endpoint, pricing, and access path are verified.
Performance: what the measured gap means
Gemini 3.5 Flash (medium) has the only reported output-speed measurement, at 276.619 median output tokens per second. Data provided by https://artificialanalysis.ai/ reports no corresponding GPT-5 mini (high) value, so the evidence supports a Gemini speed measurement, not a complete speed ranking.
That distinction matters for developer workflows. A measured output rate can improve interactive coding assistance, streamed explanations, and agent traces where users see tokens arrive progressively. It does not prove that Gemini completes a task sooner. End-to-end time also depends on prompt size, tool calls, queueing, retries, structured-output validation, and application-side processing. The supplied snapshot reports 0.3 seconds of latency for each model, so the available latency signal does not separate them.
GPT-5 mini (high) may still perform better on a specific task, especially because the snapshot reports a math index of 90.7 and a coding index of 15.6 for that model. Gemini has no matching values in the supplied data. Those missing comparison points are a hard evidence limit, not evidence of poor Gemini performance.
The practical test should therefore measure task completion time and correction rate, not token streaming alone. Use representative repository changes, tool-use loops, and structured responses. Keep the same prompts, tools, stopping rules, and acceptance tests. The supplied research found no reliable Reddit, Hacker News, or X posts that could validate coding feel, speed perception, or recurring behavioral quirks.
Cost: the cheaper model can still cost more
GPT-5 mini (high) is substantially cheaper on the supplied blended-token measure, but Gemini 3.5 Flash (medium) may be economically preferable when quality reduces retries or tool calls. Data provided by https://artificialanalysis.ai/ lists GPT-5 mini (high) at $0.6875 per 1M blended tokens and Gemini at $3.375 per 1M blended tokens.
The price gap favors GPT-5 mini (high) for high-volume generation, classification, summarization, and other workloads where output quality is already sufficient. Its supplied input price is $0.25 per 1M tokens and its output price is $2 per 1M tokens. Gemini’s corresponding prices are $1.50 and $9.00. Output-heavy applications feel this difference more strongly because generated tokens carry the higher rate for both models.
Unit price is not the same as workflow cost. A less suitable model can require extra validation calls, human review, retries, or fallback routing. The research material does not provide failure rates, task success rates, or token consumption by workflow, so it cannot establish a total-cost winner.
Gemini also supports Google Search and Google Maps grounding under documented pricing rules. The pricing page states that each has a shared free allowance of 5,000 monthly queries, followed by a charge of $14 per 1,000 queries. Gemini API pricing provides those rules. The free tier does not provide that grounding, so a seemingly low-cost prototype can change economics when current web or map information becomes necessary.
GPT-5 mini (high) leads on 3 of 3 metrics
Recommendation by developer scenario
Gemini 3.5 Flash (medium) should be the default shortlist choice when verified Google API access and general task quality matter more than token price. The measured Intelligence Index is 45.4, and Google officially describes gemini-3.5-flash as Stable. Gemini API model documentation supports that availability claim, while Data provided by https://artificialanalysis.ai/ supports the measured score.
Choose Gemini for applications that need Google Search or Google Maps grounding, provided the project can operate on the paid tier. Choose it also for interactive systems where the reported 276.619 median output tokens per second is relevant to perceived responsiveness. Neither point proves superior task completion, so acceptance tests remain necessary.
Choose GPT-5 mini (high) for a cost-first benchmark or a workload where the supplied math signal of 90.7 matches the product’s actual needs. Its blended price of $0.6875 per 1M tokens is materially lower than Gemini’s $3.375. That recommendation is conditional because OpenAI Pricing does not list the model, and OpenAI Models does not confirm the display name or API mapping.
Do not make a final procurement decision from this snapshot alone. First verify each endpoint, then run the same task set across coding, mathematics, tool use, structured output, and long-context behavior. The supplied research provides no verified context-window values, output limits, detailed parameters, or model-specific failure cases for either named variant.
Questions to resolve before production
Developers should resolve model identity and evaluation coverage before treating either label as a production contract. Google documents gemini-3.5-flash as Stable, but the supplied research does not confirm that the (medium) label maps to that endpoint. Gemini API model documentation is the authoritative starting point for that check.
OpenAI’s supplied catalog and pricing page do not list gpt-5-mini or GPT-5 mini (high). OpenAI Models and OpenAI Pricing should be checked again before implementation. The research also found no reliable community posts with disclosed test methods, so anecdotal coding or speed claims should not fill the evidence gap.
The most important unresolved question is task-specific quality. Gemini has the higher general intelligence score, while GPT-5 mini (high) has the only supplied coding and math scores. That asymmetric evidence makes broad capability claims unsafe. A small, reproducible bake-off is required before selecting a default model.
Sources
- Gemini API model documentationVerifying Gemini’s official model name, Stable status, API alias, and the absence of a separately listed medium variant.
- Gemini API pricingVerifying Gemini’s positioning, token prices, grounding rules, free-tier limitations, and pricing modes.
- OpenAI ModelsChecking the current OpenAI model catalog and the absence of a separately listed gpt-5-mini or GPT-5 mini (high) entry.
- OpenAI PricingChecking the current OpenAI pricing catalog and the absence of supplied pricing for gpt-5-mini.
- Artificial AnalysisAttributing the supplied intelligence, coding, mathematics, speed, latency, and pricing comparison data.
Your Questions about the Gemini 3.5 Flash (medium) vs GPT-5 mini (high) Comparison
Is Gemini 3.5 Flash (medium) better than GPT-5 mini (high) overall?
Gemini 3.5 Flash (medium) leads on the available general intelligence measurement, scoring 45.4 versus 25.3, but the evidence does not establish an overall winner for coding, mathematics, context handling, or tool use.
Which model is cheaper for API workloads?
GPT-5 mini (high) is cheaper on the supplied blended-token comparison at $0.6875 per 1M tokens versus Gemini 3.5 Flash (medium) at $3.375, although retries and grounding can change total workflow cost.
Which model should I use for coding agents?
Gemini 3.5 Flash (medium) is the better-supported shortlist candidate, but the supplied evidence cannot prove a coding winner because only GPT-5 mini (high) has a reported coding index of 15.6.
Can I call Gemini 3.5 Flash (medium) directly with that exact name?
Developers should verify the endpoint first because Google’s official catalog lists gemini-3.5-flash as Stable but does not separately list Gemini 3.5 Flash (medium) or gemini-3-5-flash-medium.
Can I assume GPT-5 mini (high) is currently available from OpenAI?
Developers should not assume availability because the supplied OpenAI model catalog and pricing page do not list gpt-5-mini or confirm whether high is an official model identifier or parameter value.