GPT-5 (high) vs GPT-5.4 mini (medium): The Ultimate Performance & Pricing Comparison
Deep dive into reasoning, benchmarks, and latency insights.
The Final Verdict in the GPT-5 (high) vs GPT-5.4 mini (medium) Showdown
The current catalog does not contain complete performance evidence for both models, so this page does not declare an overall winner. Use the available fields as comparison signals and validate the models on your own workload.
Model Snapshot
Key decision metrics at a glance.
Machine-readable comparison data
| Model | Metric | Value | Unit | Source / snapshot |
|---|---|---|---|---|
| GPT-5 (high) | Reasoning | 9.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5.4 mini (medium) | Reasoning | 6.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5 (high) | Coding | 4.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5.4 mini (medium) | Coding | 6.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5 (high) | Multimodal | 3.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5.4 mini (medium) | Multimodal | 2.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5 (high) | Long Context | 4.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5.4 mini (medium) | Long Context | 4.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5 (high) | Blended Price / 1M tokens | $3.438 | USD per 1M tokens | Artificial Analysis · current catalog |
| GPT-5.4 mini (medium) | Blended Price / 1M tokens | $1.688 | USD per 1M tokens | Artificial Analysis · current catalog |
| GPT-5 (high) | P95 Latency | — | milliseconds | Artificial Analysis · current catalog |
| GPT-5.4 mini (medium) | P95 Latency | — | milliseconds | Artificial Analysis · current catalog |
| GPT-5 (high) | Tokens per second | — | tokens per second | Artificial Analysis · current catalog |
| GPT-5.4 mini (medium) | Tokens per second | — | tokens per second | Artificial Analysis · current catalog |
Data provided by Artificial Analysis; live values use the current catalog.
Overall Capabilities
This radar chart visually maps the core capabilities (reasoning, coding, math proxy, multimodal, long context) of `GPT-5 (high)` vs `GPT-5.4 mini (medium)`.
Benchmark Breakdown
This grouped bar chart provides a side-by-side comparison for each benchmark metric.
Speed & Latency
Lower time to first token is better; higher tokens per second is better.
The Economics of GPT-5 (high) vs GPT-5.4 mini (medium)
Pricing Breakdown
Compare input and output pricing in USD per 1M tokens.
Real-World Cost Scenario
Per run: 1M input tokens + 250k output tokensGPT-5 (high)$3.75
GPT-5.4 mini (medium)$1.875
GPT-5.4 mini (medium) costs $1.875 less per run
GPT-5 vs GPT-5.4 mini (medium): Which Model Should Developers Choose?
This article is a dated snapshot published on 2026-08-07. Live cards above use the current catalog; missing live fields are not inferred.

- Winner overall: GPT-5, with a 34.7 Artificial Analysis Intelligence Index score versus 29.8 for GPT-5.4 mini (medium)
- Cheaper: GPT-5.4 mini (medium) at $1.6875 vs $3.4375 per 1M blended tokens
- Faster: GPT-5 and GPT-5.4 mini (medium) tie at 0.3 seconds (latency)
- Pick GPT-5.4 mini (medium) when: lower token cost matters more than verified coding, math, and API capability evidence
- Watch out: GPT-5.4 mini (medium) is not confirmed as an official API model ID, and no direct coding or speed comparison is available
GPT-5 vs GPT-5.4 mini (medium) at a glance
GPT-5 is the safer documented choice, while GPT-5.4 mini (medium) is the cheaper but less verifiable option for production selection. OpenAI positions GPT-5 for coding, reasoning, and agentic tasks in its developer announcement. The official GPT-5 model documentation lists a stable gpt-5 alias, a fixed snapshot named gpt-5-2025-08-07, a 400,000-token context window, and a 128,000-token maximum output.
The comparison is less symmetrical than the names suggest. The data snapshot gives GPT-5 an Artificial Analysis Intelligence Index score of 34.7, compared with 29.8 for GPT-5.4 mini (medium). It also reports GPT-5 at 3.4375 dollars per 1M blended tokens and GPT-5.4 mini (medium) at 1.6875 dollars. Both models show 0.3 seconds of latency, while neither has a reported median output speed.
GPT-5.4 mini (medium) has a major identification problem. The official OpenAI Models page does not list the exact gpt-5-4-mini-medium slug. The official OpenAI API Pricing page lists gpt-5.4-mini instead. Developers should therefore treat the comparison label as an evaluation slug until an API mapping is confirmed.
Data provided by https://artificialanalysis.ai/
Summary: verified capability versus lower operating cost
GPT-5 leads the verified capability comparison, but GPT-5.4 mini (medium) leads the price comparison by a wide margin. The Artificial Analysis Intelligence Index reports 34.7 for GPT-5 and 29.8 for GPT-5.4 mini (medium). GPT-5 also has reported scores of 37.8 on the Artificial Analysis Coding Index and 94.3 on the Artificial Analysis Math Index. The data snapshot does not provide corresponding coding or math scores for GPT-5.4 mini (medium), so those dimensions remain unproven rather than lost.
GPT-5 has a clearer API contract. OpenAI documents text and image input, text output, function calling, structured outputs, streaming, custom tools, reasoning_effort, and verbosity in the GPT-5 developer documentation and GPT-5 model documentation. The same documentation states that audio and video input and output are unsupported, while fine-tuning and predicted outputs are unavailable.
GPT-5.4 mini (medium) has less model-specific evidence. The OpenAI Models page provides general information about current OpenAI models, including text and image input, text output, multilingual capability, and visual capability. It does not establish that every statement applies specifically to the evaluation slug. It also does not document that slug's context window, maximum output, reasoning parameters, or official benchmark results.
The practical conclusion is conditional. Choose GPT-5 when auditability, documented behavior, or complex reasoning matters. Choose GPT-5.4 mini (medium) for cost-sensitive experiments only after confirming the callable model ID and testing the target workload.
Performance: what the available evidence means in real applications
GPT-5 is the only model in this comparison with direct evidence for coding and mathematical performance. The data snapshot reports 37.8 on the Artificial Analysis Coding Index and 94.3 on the Artificial Analysis Math Index for GPT-5. GPT-5.4 mini (medium) has no corresponding values in the snapshot. This prevents a clean claim that GPT-5 is better at coding or math, but it does establish that GPT-5 has measured evidence where the mini variant does not.
For software teams, the missing coding score is more important than the raw intelligence gap. A general intelligence index can help rank models, but it does not answer whether a model will preserve interfaces, modify tests correctly, or navigate a large repository. A developer choosing GPT-5.4 mini (medium) for code changes would need a task-specific evaluation before trusting the cheaper model with autonomous edits.
GPT-5's official benchmark evidence is broad but has important conditions. OpenAI reports 74.9% on SWE-bench Verified, 88% on Aider polyglot, 96.7% on τ²-bench telecom, and 69.6% on Scale MultiChallenge in its developer announcement. OpenAI states that the SWE-bench result excluded 23 of 500 problems that could not be passed reliably on its infrastructure. The Aider result used high reasoning effort. These details make the results useful signals, not universal production guarantees.
Latency does not separate the models in the supplied data. Both are listed at 0.3 seconds, and neither has a median output-tokens-per-second value. The evidence therefore cannot answer which model feels faster during long generations, streaming responses, or tool-heavy workflows. Developers should measure time to first token, completion time, retries, and tool-call overhead in their own stack.
Community evidence adds a narrow warning for GPT-5. One Reddit author reported that GPT-5 handled small bug fixes quickly, while complete applications and UI generation could be too concise and lack design detail. The same Reddit discussion includes claims about hallucinations and incorrect edits in complex existing codebases. The post was an uncontrolled personal test, so it cannot establish a stable failure rate. No comparable community evidence was found for GPT-5.4 mini (medium).
Cost: the cheaper model changes the default economics
GPT-5.4 mini (medium) is the clear cost winner, but its lower price does not by itself make it cheaper for a complete developer workflow. The supplied data lists 1.6875 dollars per 1M blended tokens for GPT-5.4 mini (medium), versus 3.4375 dollars for GPT-5. GPT-5.4 mini (medium) also has lower standard input and output prices, at 0.75 dollars and 4.5 dollars per 1M tokens, compared with GPT-5 at 1.25 dollars and 10 dollars.
The price advantage matters most for high-volume tasks with predictable quality requirements. Classification, extraction, summarization, routing, and short code assistance can benefit when the cheaper model passes a representative acceptance set. The savings become less meaningful if the model needs repeated retries, longer repair prompts, human review, or escalation to GPT-5 after failed outputs.
Output-heavy workloads deserve special attention. GPT-5's listed output price is 10 dollars per 1M tokens, while GPT-5.4 mini (medium) is listed at 4.5 dollars. A workflow that generates long patches, explanations, or agent traces can therefore amplify the gap. A workflow dominated by cached prompts may behave differently because GPT-5's documentation lists cached input at 0.125 dollars per 1M tokens, while the OpenAI API Pricing page lists 0.075 dollars for GPT-5.4 mini. These figures still do not resolve the model identity issue.
GPT-5.4 mini (medium) also has additional listed service modes. The pricing brief records Batch and Flex prices of 0.375 dollars for input, 0.0375 dollars for cached input, and 2.25 dollars for output per 1M tokens. It records Fast mode at 1.50 dollars for input, 0.15 dollars for cached input, and 9.00 dollars for output. The official pricing page does not identify a separate gpt-5-4-mini-medium product, so developers must verify that these modes apply to the endpoint they can actually call.
The correct cost test is cost per accepted result, not cost per token. GPT-5.4 mini (medium) wins the displayed token economics. GPT-5 may still win a narrow workflow if its stronger verified evidence reduces retries, review, or failed tool actions.
GPT-5.4 mini (medium) leads on 3 of 3 metrics
Recommendation: select by risk tolerance and verification burden
GPT-5 is the recommended default for production systems that need documented reasoning, coding evidence, or a stable API identity. OpenAI documents gpt-5 as a callable alias and supports Chat Completions, Responses, and Batch endpoints in the GPT-5 model documentation. The documentation also marks the fixed snapshot gpt-5-2025-08-07 as Deprecated and recommends GPT-5.6, which creates migration risk even when the alias remains available.
GPT-5.4 mini (medium) is the better first candidate for cost-sensitive workloads, but only after an API identity check. The OpenAI API Pricing page names gpt-5.4-mini, not gpt-5-4-mini-medium. A deployment plan that copies the evaluation slug directly into an API request may fail before model quality is even tested. The first gate should be a successful call using the exact official model ID, followed by a workload-specific quality test.
Use GPT-5 for repository-wide changes, difficult debugging, mathematical reasoning, multi-step agents, and tasks where errors are expensive. GPT-5's documented reasoning_effort values include minimal, low, medium, and high, and its official benchmark results include coding and agent-relevant evaluations in the developer announcement. These facts support a stronger evidence case, although they do not guarantee success on a particular codebase.
Use GPT-5.4 mini (medium) for high-volume, lower-risk operations when its lower price has a measurable business benefit. Good candidates include structured extraction, simple transformations, routine support drafting, and bounded code suggestions. The evidence does not establish its context capacity, output ceiling, reasoning controls, or coding behavior. Those unknowns must remain explicit in the decision record.
A sensible rollout is staged. First, confirm gpt-5.4-mini is the intended endpoint. Second, run both models on the same representative requests. Third, compare accepted-result cost, correction rate, latency distribution, tool-call success, and human review time. Fourth, keep GPT-5 as an escalation path if the mini model fails quality gates. This approach preserves the price advantage without treating missing evidence as proof of equivalence.
The strongest overall recommendation is therefore not a universal winner. GPT-5 is the evidence-backed choice. GPT-5.4 mini (medium) is the economics-backed experiment. The supplied materials do not provide enough direct evidence to conclude that the mini model matches GPT-5 for coding, mathematics, long-context work, or production reliability.
Before you choose: unresolved questions
GPT-5 is easier to evaluate today because its official API identity and capability documentation are available. The GPT-5 model documentation supplies the clearest reference point for implementation checks, while the OpenAI Models page does not verify the exact GPT-5.4 mini (medium) evaluation slug.
GPT-5.4 mini (medium) remains attractive because its listed blended price is 1.6875 dollars per 1M tokens, less than GPT-5's 3.4375 dollars. That advantage should be treated as conditional until the official endpoint, service mode, context behavior, and acceptance rate are confirmed.
The comparison also has a version asymmetry. GPT-5 has a fixed snapshot marked Deprecated, while the supplied research does not find official evidence that GPT-5.4 mini (medium) is an independent API variant or that it has been replaced. Teams should record the exact model ID, snapshot behavior, and migration path during evaluation.
Sources
- GPT-5 for developersGPT-5 positioning, reasoning parameters, tool support, and official benchmark context
- GPT-5 model documentationGPT-5 model ID, context and output limits, modalities, pricing, endpoints, unsupported features, and deprecation status
- Tried GPT-5 Here Are My First ImpressionsUncontrolled community observations about GPT-5 debugging, application generation, and complex codebase risks
- OpenAI ModelsVerification of general model catalog information and absence of a dedicated gpt-5-4-mini-medium entry
- OpenAI API PricingVerification of the gpt-5.4-mini model ID, pricing modes, and absence of a dedicated gpt-5-4-mini-medium entry
- Artificial AnalysisData attribution for the supplied comparison snapshot, evaluation scores, latency, and token pricing
Your Questions about the GPT-5 (high) vs GPT-5.4 mini (medium) Comparison
Is GPT-5 better than GPT-5.4 mini (medium) for coding?
GPT-5 has stronger coding evidence because the data snapshot reports a 37.8 Artificial Analysis Coding Index score and OpenAI reports coding benchmarks, while GPT-5.4 mini (medium) has no corresponding coding score. That evidence supports GPT-5 as the safer coding choice, but it does not prove superiority on every repository or programming language.
Which model is cheaper for API usage?
GPT-5.4 mini (medium) is cheaper in the supplied pricing data, at $1.6875 per 1M blended tokens compared with $3.4375 for GPT-5. Its listed input price is $0.75 and output price is $4.5 per 1M tokens, versus $1.25 and $10 for GPT-5. Actual cost per accepted result can reverse if the cheaper model causes more retries, repairs, or human review.
Which model is faster?
Neither model is faster according to the supplied latency data because GPT-5 and GPT-5.4 mini (medium) are both listed at 0.3 seconds. The dataset does not provide median output tokens per second for either model, so it cannot establish which model streams longer answers faster or finishes tool-heavy tasks sooner. A production benchmark should measure time to first token, completion time, and retry overhead.
Can I call gpt-5-4-mini-medium directly through the OpenAI API?
The available official evidence does not confirm that gpt-5-4-mini-medium is a callable OpenAI API model ID. The official pricing page lists gpt-5.4-mini, while the exact evaluation slug is absent. Developers should verify the official model identifier and make a successful test request before integrating it into application code.
Does GPT-5 support audio and video input?
GPT-5 does not support audio or video input and output according to the GPT-5 model documentation. The documented API modality is text and image input with text output, so applications requiring direct audio or video processing need another model or an additional preprocessing service.
Should developers use the GPT-5 fixed snapshot in production?
Developers should avoid treating the fixed GPT-5 snapshot as a long-term risk-free dependency because the official model documentation marks gpt-5-2025-08-07 as Deprecated. The stable gpt-5 alias remains documented, but teams should monitor migration guidance, test replacements, and record the exact version used by production.