GPT-5.4 Pro (xhigh) vs GPT-5 mini (high): The Ultimate Performance & Pricing Comparison
Deep dive into reasoning, benchmarks, and latency insights.
The Final Verdict in the GPT-5.4 Pro (xhigh) vs GPT-5 mini (high) Showdown
The current catalog does not contain complete performance evidence for both models, so this page does not declare an overall winner. Use the available fields as comparison signals and validate the models on your own workload.
Model Snapshot
Key decision metrics at a glance.
Machine-readable comparison data
| Model | Metric | Value | Unit | Source / snapshot |
|---|---|---|---|---|
| GPT-5.4 Pro (xhigh) | Reasoning | 6.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5 mini (high) | Reasoning | 9.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5.4 Pro (xhigh) | Coding | 6.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5 mini (high) | Coding | 2.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5.4 Pro (xhigh) | Multimodal | 5.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5 mini (high) | Multimodal | 2.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5.4 Pro (xhigh) | Long Context | 8.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5 mini (high) | Long Context | 3.0 | benchmark or capability score | Artificial Analysis · current catalog |
| GPT-5.4 Pro (xhigh) | Blended Price / 1M tokens | $67.5 | USD per 1M tokens | Artificial Analysis · current catalog |
| GPT-5 mini (high) | Blended Price / 1M tokens | $0.688 | USD per 1M tokens | Artificial Analysis · current catalog |
| GPT-5.4 Pro (xhigh) | P95 Latency | — | milliseconds | Artificial Analysis · current catalog |
| GPT-5 mini (high) | P95 Latency | — | milliseconds | Artificial Analysis · current catalog |
| GPT-5.4 Pro (xhigh) | Tokens per second | 0 | tokens per second | Artificial Analysis · current catalog |
| GPT-5 mini (high) | Tokens per second | 0 | tokens per second | Artificial Analysis · current catalog |
Data provided by Artificial Analysis; live values use the current catalog.
Overall Capabilities
This radar chart visually maps the core capabilities (reasoning, coding, math proxy, multimodal, long context) of `GPT-5.4 Pro (xhigh)` vs `GPT-5 mini (high)`.
Benchmark Breakdown
This grouped bar chart provides a side-by-side comparison for each benchmark metric.
Speed & Latency
Lower time to first token is better; higher tokens per second is better.
The Economics of GPT-5.4 Pro (xhigh) vs GPT-5 mini (high)
Pricing Breakdown
Compare input and output pricing in USD per 1M tokens.
Real-World Cost Scenario
Per run: 1M input tokens + 250k output tokensGPT-5.4 Pro (xhigh)$75
GPT-5 mini (high)$0.75
GPT-5 mini (high) costs $74.25 less per run
GPT-5.4 Pro (xhigh) vs GPT-5 mini (high): Which Model Should Developers Choose?
This article is a dated snapshot published on 2026-08-16. Live cards above use the current catalog; missing live fields are not inferred.

- Winner overall: GPT-5 mini (high), the only model with reported evaluation results, including a 25.8 intelligence index and 15.6 coding index
- Cheaper: GPT-5 mini (high) at $0.688 vs $67.5 per 1M blended tokens
- Faster: Neither model, both at 0 median output tokens per second in the supplied snapshot
- Pick GPT-5 mini (high) when: you need a currently measurable, low-cost default for coding, structured tasks, and high-volume production traffic
- Watch out: GPT-5.4 Pro (xhigh) has no reported benchmark, latency, context, or official catalog entry in the supplied evidence
GPT-5.4 Pro (xhigh) vs GPT-5 mini (high)
GPT-5 mini (high) is the safer developer default because it is the only model with reported benchmark results and a listed blended price in the supplied data. Artificial Analysis reports a 25.8 intelligence index, a 15.6 coding index, and a $0.688 price per 1M blended tokens for GPT-5 mini (high). The same snapshot reports no evaluation results, no measured speed, and a $67.5 blended price for GPT-5.4 Pro (xhigh). Data provided by https://artificialanalysis.ai/
GPT-5.4 Pro (xhigh) may still be attractive for teams seeking a more expensive, potentially higher-effort reasoning configuration. However, the supplied evidence cannot confirm that interpretation. The current OpenAI Models directory does not list gpt-5-4-pro or “GPT-5.4 Pro (xhigh)”. The same directory also does not list gpt-5-mini or “GPT-5 mini (high)” as independent entries. That creates an availability problem for both names, but it creates a larger decision risk for Pro because no public benchmark or official model specification offsets the uncertainty.
This comparison therefore answers a practical question: which model can a developer choose with evidence available today? GPT-5 mini (high) wins that narrower decision. It does not prove that GPT-5 mini (high) is intrinsically more capable. It proves that its reported behavior is easier to inspect, price, and test before production adoption.
Executive summary for developers
GPT-5 mini (high) is the evidence-backed choice, while GPT-5.4 Pro (xhigh) remains an unverified option with a much higher reported cost. The comparison is asymmetric because the two model records do not contain comparable capability evidence.
| Decision area | GPT-5.4 Pro (xhigh) | GPT-5 mini (high) |
|---|---|---|
| Reported intelligence index | Not reported | 25.8 |
| Reported coding index | Not reported | 15.6 |
| Reported math index | Not reported | 90.7 |
| Blended price per 1M tokens | $67.5 | $0.688 |
| Input price per 1M tokens | $30 | $0.25 |
| Output price per 1M tokens | $180 | $2 |
| Median output speed | 0 | 0 |
| Latency | 0 | 0 |
GPT-5 mini (high) has usable evidence across coding, mathematics, general intelligence, instruction following, long-context retrieval, tool use, and terminal tasks. Its reported results include 0.838 on LiveCodeBench, 0.906666666666667 on AIME 25, 0.754421768707483 on IFBench, and 0.71 on LCR. These values do not establish universal production quality, but they give a developer something concrete to validate.
GPT-5.4 Pro (xhigh) has no reported score in any supplied evaluation. The absence includes the Artificial Analysis intelligence, coding, and math indexes, plus every listed benchmark. This means the comparison cannot establish a capability winner for difficult coding, reasoning, tool use, or mathematical work.
The official evidence is also incomplete for GPT-5 mini (high). OpenAI’s model directory does not confirm its context window, output limit, API parameters, tools, or whether “high” is a model ID or a configuration. Developers should treat the display names as records requiring endpoint-level verification, not as confirmed public API contracts.
Performance: the measured advantage belongs to GPT-5 mini (high)
GPT-5 mini (high) is the only model with measured task performance in the supplied snapshot, so developers can validate its fit even though its limits remain unclear. GPT-5.4 Pro (xhigh) has no reported score for any benchmark, which prevents a fair claim that it is better at software engineering, reasoning, mathematics, or agent work.
The available GPT-5 mini (high) profile is uneven in a useful way. Its reported math index is 90.7, while its coding index is 15.6 and its intelligence index is 25.8. Those figures suggest that benchmark strength may vary sharply by task family. A team building a mathematical verifier should not assume the coding score predicts the same result. A team building a coding agent should test repository navigation, patch quality, test repair, and tool-call reliability directly.
The benchmark pattern also exposes a boundary that a summary score would hide. GPT-5 mini (high) records 0.838 on LiveCodeBench but 0.333333333333333 on TerminalBench Hard and 0.0374531835205993 on TerminalBench v2.1. Those results do not describe every terminal workflow, yet they warn against equating code-generation ability with reliable autonomous execution. Production tasks that require shell commands, environment inspection, or multi-step repair deserve a separate pilot.
GPT-5 mini (high) also records 0.684210526315789 on tau2 and 0.154639175257732 on tau banking. That spread matters for developers building customer-service or transaction workflows. A model can look adequate for broad tool-use testing while remaining weak in a specialized domain. The supplied data cannot show whether GPT-5.4 Pro (xhigh) avoids those weaknesses because its corresponding values are missing.
Both models show 0 median output tokens per second and 0 latency seconds in the supplied snapshot. These values should be treated as unavailable or non-measured rather than as proof of instant responses. No reliable speed winner can be selected. The research brief also found no verifiable Reddit, Hacker News, or X posts describing either model’s coding feel or response speed. OpenAI’s current model documentation provides general statements about current OpenAI models, including text and image input, text output, multilingual capability, and vision, but does not confirm that those claims apply to either named record.
Cost: GPT-5 mini (high) changes the economics of experimentation
GPT-5 mini (high) is the clear cost choice, but GPT-5.4 Pro (xhigh) could still be cheaper overall if it prevents enough retries, reviews, or failed tool runs. The supplied pricing snapshot lists GPT-5.4 Pro (xhigh) at $67.5 per 1M blended tokens, compared with $0.688 for GPT-5 mini (high). It lists input pricing at $30 versus $0.25 and output pricing at $180 versus $2. Data provided by https://artificialanalysis.ai/
The chart below the article already shows the price gap. The practical meaning is more important: output-heavy workloads face a particularly strong penalty with Pro because its reported output price is $180 per 1M tokens. Long answers, repeated code patches, agent traces, and generated documentation can turn output volume into the main budget driver. GPT-5 mini (high) gives teams much more room to run evaluations, compare prompts, and retry failed tasks.
Cheap tokens do not guarantee cheap completed work. If GPT-5 mini (high) requires multiple attempts, manual correction, or an additional model for verification, the workflow cost rises beyond the listed token price. The current data does not report pass rates, retry counts, human review time, or end-to-end task cost for either model. Developers should therefore compare cost per accepted result, not only cost per 1M tokens.
The reverse case is also possible. GPT-5.4 Pro (xhigh) might justify its price for a small number of high-value tasks if it produces materially better first-pass outputs. The supplied evidence cannot verify that benefit. No Pro benchmark, latency measurement, context limit, or failure analysis is available. OpenAI’s pricing page also does not list GPT-5.4 Pro (xhigh) or GPT-5 mini (high), so the displayed prices should be validated against the actual billing route before procurement.
For most teams, GPT-5 mini (high) is the rational starting point because its low reported price lowers the cost of learning. The decision should change only after a controlled task sample shows that Pro reduces total work enough to offset its much higher token rates.
GPT-5 mini (high) leads on 3 of 3 metrics
Recommendation: choose by evidence threshold, not model name
GPT-5 mini (high) should be the default pilot model for most developer products because it combines measurable task results with the lower reported token prices. This recommendation applies to code assistants, structured generation, mathematical workflows, retrieval-heavy prompts, and high-volume automation where teams need affordable testing.
Choose GPT-5 mini (high) when:
- The product needs a low-cost baseline before deeper model procurement.
- The workload includes many requests, retries, or prompt experiments.
- The team can add tests and human review around terminal or tool-use actions.
- The task benefits from measurable benchmark evidence, even if the benchmarks are imperfect.
Consider GPT-5.4 Pro (xhigh) only after availability and behavior are verified. The research brief found no current official catalog entry for gpt-5-4-pro, no official price listing, no official context specification, and no reliable community test. That does not make Pro unusable. It means a team should not commit production architecture, budgets, or service-level expectations around the name alone.
A sensible evaluation should use the team’s real tasks. Include representative coding tickets, tests that must pass, terminal operations, structured outputs, long-context retrieval, and the human review needed to accept an answer. Record accepted-result rate, retries, review time, and total token use. The supplied materials do not provide those measurements, so no article-level claim can replace this pilot.
The most important decision rule is simple. If GPT-5 mini (high) meets the acceptance bar, its $0.688 blended price makes it the stronger operating choice. If it fails the bar, test Pro through a confirmed endpoint and compare total completed-task cost. Do not infer Pro’s superiority from the word “Pro”, the “xhigh” label, or the $67.5 price. The evidence does not support those shortcuts.
Availability is a separate gate. OpenAI’s model directory and pricing documentation do not confirm either supplied display name as a current independent model entry. Developers should verify the exact API model ID, access permissions, pricing route, and retirement policy before writing integration code.
Questions developers should answer before adoption
GPT-5 mini (high) is easier to evaluate, but neither model has a complete official specification in the supplied sources. The following questions identify the gaps most likely to affect implementation and procurement.
Sources
- OpenAI ModelsVerifying current model catalog coverage, general capability statements, API availability, and the absence of dedicated entries or specifications for the supplied model names.
- OpenAI PricingChecking current official pricing coverage, model listing status, and the absence of dedicated official prices or alias mappings for the supplied model names.
- Artificial AnalysisAttributing the supplied benchmark, pricing, speed, and comparison snapshot data.
Your Questions about the GPT-5.4 Pro (xhigh) vs GPT-5 mini (high) Comparison
Which model should a developer choose for a new production application?
GPT-5 mini (high) is the safer starting choice because it has reported benchmark results and a $0.688 blended price per 1M tokens. GPT-5.4 Pro (xhigh) lacks comparable measurements and is not listed in the supplied official model directory or pricing page. Teams should still validate the exact API model ID, task acceptance rate, and production behavior before launch.
Is GPT-5.4 Pro (xhigh) more capable than GPT-5 mini (high)?
GPT-5.4 Pro (xhigh) cannot be shown to be more capable from the supplied evidence because every listed evaluation is missing for Pro. GPT-5 mini (high) has reported results such as a 15.6 coding index, a 90.7 math index, and a 25.8 intelligence index. Those measurements establish evidence, not universal superiority.
Why is GPT-5 mini (high) much cheaper?
GPT-5 mini (high) is listed at $0.688 per 1M blended tokens, while GPT-5.4 Pro (xhigh) is listed at $67.5. The supplied materials do not explain the pricing rationale, and the official pricing page does not list either display name. Developers should verify billing before relying on these figures for a contract or forecast.
Which model is faster for coding or agent tasks?
Neither model is faster in the supplied snapshot because both report 0 median output tokens per second and 0 latency seconds. Those values are best treated as missing or non-measured data, not as instant-response evidence. The research brief also found no verifiable community tests that establish a speed difference or a consistent coding experience.
Can developers rely on the context window and tool support for either model?
Developers cannot rely on a confirmed context window, output limit, API parameter set, or tool-support profile for either supplied model name. The current OpenAI model directory does not list gpt-5-4-pro or gpt-5-mini as independent entries. Teams should confirm these details through the actual API documentation and endpoint behavior before implementation.
When could GPT-5.4 Pro (xhigh) be worth its higher price?
GPT-5.4 Pro (xhigh) could be worth testing when a high-value task benefits from fewer retries, less human review, or better first-pass completion. The supplied data does not report those outcomes, so the claim remains unverified. A real-task pilot must demonstrate lower cost per accepted result before the $67.5 blended price is justified.