GPT-5 (medium)
AvailableOpenAI · 2025-08-07 · 400,000 tokens
An AI model from OpenAI, strongest at reasoning, suited to a broad range of AI workloads.
Quick Overview
Benchmark Results
Scores from leading benchmark suites.
Performance Metrics
Latency and throughput performance.
Dive Deeper
AI model analysis
GPT-5 (medium) Review: Strong Math, Unclear Product Fit

- **Where it stands:** GPT-5 (medium) ranks 109 of 578 on the Artificial Analysis Intelligence Index at 33.7 - **Price:** $3.4375 per 1M blended tokens - **Speed:** output tokens per second is not reported, 0.3s to first token - **Pick it when:** your workload rewards mathematical reasoning and does not depend on documented product guarantees - **Watch out:** official OpenAI pages do not currently document GPT-5 (medium) as a listed model or confirm its limits Data provided by https://artificialanalysis.ai/
GPT-5 (medium) review: a math-first choice with a documentation problem
GPT-5 (medium) looks most compelling for developers who value mathematical reasoning more than broad benchmark leadership or clear product documentation. The model ranks 18 of 265 on the Artificial Analysis Math Index at 91.7, while its Intelligence Index position is 109 of 578 at 33.7. Those results describe a meaningful split: GPT-5 (medium) is a strong candidate for structured quantitative work, but the available evidence does not support calling it a universal best choice.
The larger concern is operational. The current OpenAI Models documentation does not list gpt-5-medium or provide model-specific details for its context window, output limit, API parameters, or benchmark performance. The current OpenAI API Pricing page also does not list the model. That absence does not prove that the model cannot be accessed, but it does make availability, alias stability, and long-term planning harder to verify.
For a developer evaluating this model, the practical answer is conditional. GPT-5 (medium) deserves consideration for mathematical analysis, symbolic work, and applications where a strong math ranking can reduce verification effort. It needs more caution for production systems that require documented limits, predictable naming, or a clearly supported commercial position.
Executive summary for developers
GPT-5 (medium) offers its clearest advantage in mathematical reasoning, while its broad positioning remains difficult to justify from the available evidence. The Intelligence Index result places it alongside several nearby models with the same reported score of 33.7, so the model’s distinct case comes from its math result and its price, not from a unique general-intelligence lead.
| Decision area | GPT-5 (medium) | What the nearby models suggest |
|---|---|---|
| Broad capability signal | A mid-table result on the Intelligence Index | Claude 4.1 Opus (Reasoning), GLM-4.7 (Reasoning), KAT Coder Pro V2, MiniMax-M2.5, and Qwen3.5 397B A17B (Reasoning) share the same reported Intelligence Index score of 33.7 |
| Mathematical work | A strong position on the Math Index | GLM-4.7 (Reasoning) reports 95, while Claude 4.1 Opus (Reasoning) reports 80.3 |
| Cost position | More expensive than several lower-priced nearby references | GLM-4.7 (Reasoning) is listed at $1 per 1M blended tokens; KAT Coder Pro V2 and MiniMax-M2.5 are listed at $0.525 |
| Evidence quality | Benchmark data exists, product documentation is incomplete | The official OpenAI pages do not document this model specifically |
This comparison is a selection aid, not a claim that the neighboring models are interchangeable. The data shows where GPT-5 (medium) sits, but it does not reveal coding behavior, context limits, tool reliability, or failure recovery.
Performance: what the ranking means in real workloads
GPT-5 (medium) is best interpreted as a specialist-leaning model for quantitative reasoning, not as a proven general-purpose leader. A rank of 18 of 265 on the Artificial Analysis Math Index at 91.7 places mathematical performance near the front of that evaluated group. In practical terms, this supports testing GPT-5 (medium) on tasks such as equation transformation, numerical reasoning, constraint checking, and multi-step analytical workflows.
The result should still be treated as a screening signal. A benchmark ranking does not establish reliability on your exact prompts, domain terminology, tool calls, or output format. It also does not show whether the model reaches the correct answer consistently, explains it in a reviewable way, or fails safely when the problem is underspecified. The brief contains no model-specific failure cases, community tests, coding reports, or official limitations. Those gaps mean that the mathematical result supports a pilot, not an automatic production decision.
The broader Intelligence Index ranking changes the interpretation. GPT-5 (medium) ranks 109 of 578 at 33.7, which is materially less persuasive as evidence of broad superiority. Developers should therefore separate task families during evaluation. Use math-heavy test sets for the model’s strongest apparent area, then run independent tests for coding, retrieval, instruction following, multilingual output, structured generation, and agentic workflows. Do not infer those capabilities from the math ranking alone.
OpenAI’s general documentation says that its latest models support text and image input, text output, multilingual capabilities, and vision through the Responses API and official Client SDKs. However, the OpenAI Models page does not confirm that those general statements apply specifically to GPT-5 (medium).
Cost: reasonable for high-value reasoning, weak for undifferentiated volume
GPT-5 (medium) is economically defensible when stronger mathematical reasoning prevents expensive downstream review, but its price is harder to justify for routine generation. The blended rate is $3.4375 per 1M tokens, with input priced at $1.25 and output at $10. The high output rate matters because reasoning-heavy tasks often produce longer responses, intermediate explanations, or repeated correction attempts.
This price creates a clear condition for adoption. GPT-5 (medium) can make sense when each request carries meaningful business or engineering value, and when its math performance reduces human checking or multi-model retries. It is less attractive for summarization, classification, simple extraction, boilerplate writing, or other workloads where a cheaper model can meet the acceptance threshold. Nearby references in the data brief include GLM-4.7 (Reasoning) at $1 per 1M blended tokens, and KAT Coder Pro V2 and MiniMax-M2.5 at $0.525. Those figures do not prove that the alternatives deliver equivalent quality, but they raise the burden of proof for GPT-5 (medium) on cost-sensitive workloads.
Latency is reported at 0.3 seconds to first token, but median output throughput is not reported. That makes end-to-end cost and user experience harder to estimate for long answers. Developers should measure completion length, retry rate, correction rate, and human review time in a controlled pilot. The OpenAI API Pricing page does not currently provide a listed price for GPT-5 (medium), so the data brief’s price should be verified against the actual endpoint before procurement.
Data provided by https://artificialanalysis.ai/ supplies the benchmark and pricing snapshot used in this review.
Recommendation: test GPT-5 (medium) selectively before committing
GPT-5 (medium) is worth testing for mathematics-heavy developer workflows, but the evidence is insufficient for a broad production recommendation. The strongest case includes quantitative analysis, formula-heavy transformation, technical problem solving, and systems where correctness is more valuable than minimum token cost. A focused evaluation can establish whether the Math Index signal transfers to your own workload.
The adoption decision should include two gates. First, validate task quality against representative prompts, including ambiguous inputs, adversarial cases, malformed data, and required output schemas. Second, validate product stability through the actual API route. The OpenAI Models documentation does not currently confirm the model’s context window, output ceiling, supported parameters, stable alias, or continued availability. Those are not minor details for a production integration. They affect prompt design, budgeting, observability, and migration planning.
| Choose GPT-5 (medium) when | Keep looking when |
|---|---|
| Mathematical accuracy is central to the workload | The workload is mostly routine text processing |
| A stronger answer can reduce review or retry cost | Output volume dominates the budget |
| You can verify endpoint availability and limits | Your system needs documented guarantees before launch |
| You are willing to run task-specific acceptance tests | You need evidence for coding, tools, or failure behavior that the brief does not provide |
The recommended posture is a bounded pilot with a fallback model and explicit stop conditions. The available evidence supports investigation, not blind standardization.
FAQ for developers
GPT-5 (medium) needs task-specific validation because its strongest public signal is mathematical performance, while important product and behavior details remain undocumented. The questions below address the selection risks that the benchmark table alone cannot answer.
Frequently asked questions
Is GPT-5 (medium) a good general-purpose model for production applications?
GPT-5 (medium) may support production use, but the available evidence does not establish it as a generally superior production model. Its broad Intelligence Index position is 109 of 578, and the official OpenAI model directory does not provide model-specific limits, availability details, or supported parameter documentation. Developers should validate their own coding, tool-use, structured-output, and reliability requirements before adopting it broadly.
What is GPT-5 (medium) strongest at according to the available data?
GPT-5 (medium) appears strongest at mathematical reasoning based on its rank of 18 of 265 on the Artificial Analysis Math Index at 91.7. That result supports testing equation-heavy analysis, quantitative problem solving, and constraint-based tasks. It does not establish equivalent performance for coding, retrieval, agent workflows, or general instruction following, because the supplied brief provides no direct evidence for those areas.
Is GPT-5 (medium) worth its price?
GPT-5 (medium) is worth its price when mathematical accuracy can reduce human review, retries, or downstream errors. Its blended price is $3.4375 per 1M tokens, with output priced at $10, so high-volume routine generation may be difficult to justify. Cheaper nearby references exist, but the brief does not provide enough task-level quality data to determine whether they are interchangeable.
Can developers rely on GPT-5 (medium) being available through the OpenAI API?
Developers should verify availability directly before building a dependency because the current OpenAI Models page does not list gpt-5-medium. The supplied research also does not identify a stable alias, confirmed endpoint, context window, output limit, or replacement relationship. The model may be accessible in a particular environment, but the available official documentation does not confirm that status.
Does GPT-5 (medium) have a documented output speed?
GPT-5 (medium) does not have a reported median output speed in the supplied data brief. The brief reports 0.3 seconds to first token, but that value does not describe sustained generation throughput. Developers building interactive or latency-sensitive features should measure complete responses under representative prompts rather than infer user experience from first-token latency alone.
Sources
- OpenAI ModelsChecking whether GPT-5 (medium) is officially listed and reviewing the general documented capability and API claims.
- OpenAI API PricingChecking whether GPT-5 (medium) has a current official listed price and reviewing the documented product line.
- Artificial AnalysisAttributing the benchmark, ranking, latency, throughput, and pricing snapshot supplied in the data brief.
Published: