Skip to content

GPT-5 (high) vs GPT-5.5 (medium): The Ultimate Performance & Pricing Comparison

Deep dive into reasoning, benchmarks, and latency insights.

The Final Verdict in the GPT-5 (high) vs GPT-5.5 (medium) Showdown

The current catalog does not contain complete performance evidence for both models, so this page does not declare an overall winner. Use the available fields as comparison signals and validate the models on your own workload.

Model Snapshot

Key decision metrics at a glance.

GPT-5 (high)GPT-5.5 (medium)
9.0
Reasoning
6.0
4.0
Coding
7.0
3.0
Multimodal
4.0
4.0
Long Context
6.0
$3.438
Blended Price / 1M tokens
$11.25
P95 Latency
Tokens per second

Machine-readable comparison data

ModelMetricValueUnitSource / snapshot
GPT-5 (high)Reasoning9.0benchmark or capability scoreArtificial Analysis · current catalog
GPT-5.5 (medium)Reasoning6.0benchmark or capability scoreArtificial Analysis · current catalog
GPT-5 (high)Coding4.0benchmark or capability scoreArtificial Analysis · current catalog
GPT-5.5 (medium)Coding7.0benchmark or capability scoreArtificial Analysis · current catalog
GPT-5 (high)Multimodal3.0benchmark or capability scoreArtificial Analysis · current catalog
GPT-5.5 (medium)Multimodal4.0benchmark or capability scoreArtificial Analysis · current catalog
GPT-5 (high)Long Context4.0benchmark or capability scoreArtificial Analysis · current catalog
GPT-5.5 (medium)Long Context6.0benchmark or capability scoreArtificial Analysis · current catalog
GPT-5 (high)Blended Price / 1M tokens$3.438USD per 1M tokensArtificial Analysis · current catalog
GPT-5.5 (medium)Blended Price / 1M tokens$11.25USD per 1M tokensArtificial Analysis · current catalog
GPT-5 (high)P95 LatencymillisecondsArtificial Analysis · current catalog
GPT-5.5 (medium)P95 LatencymillisecondsArtificial Analysis · current catalog
GPT-5 (high)Tokens per secondtokens per secondArtificial Analysis · current catalog
GPT-5.5 (medium)Tokens per secondtokens per secondArtificial Analysis · current catalog

Data provided by Artificial Analysis; live values use the current catalog.

Overall Capabilities

This radar chart visually maps the core capabilities (reasoning, coding, math proxy, multimodal, long context) of `GPT-5 (high)` vs `GPT-5.5 (medium)`.

IntelligenceCodingMathMultimodalLong Context
GPT-5 (high)GPT-5.5 (medium)

Benchmark Breakdown

This grouped bar chart provides a side-by-side comparison for each benchmark metric.

GPT-5 (high)GPT-5.5 (medium)

Speed & Latency

Lower time to first token is better; higher tokens per second is better.

Time to First Token · GPT-5 (high)
Time to First Token · GPT-5.5 (medium)
Tokens per Second · GPT-5 (high)
Tokens per Second · GPT-5.5 (medium)
Head to the playground to validate these results yourself

The Economics of GPT-5 (high) vs GPT-5.5 (medium)

Pricing Breakdown

Compare input and output pricing in USD per 1M tokens.

GPT-5 (high)GPT-5.5 (medium)

Real-World Cost Scenario

Per run: 1M input tokens + 250k output tokens

GPT-5 (high)$3.75

GPT-5.5 (medium)$12.5

GPT-5 (high) costs $8.75 less per run

Review the complete pricing and packaging strategy

GPT-5 (high) vs GPT-5.5 (medium): Which Model Should Developers Choose?

This article is a dated snapshot published on 2026-08-07. Live cards above use the current catalog; missing live fields are not inferred.

GPT-5 (high) vs GPT-5.5 (medium): Which Model Should Developers Choose?
  • Winner overall: GPT-5.5 (medium), with a 71.5 coding index versus 37.8 for GPT-5 (high)
  • Cheaper: GPT-5 (high) at $3.4375 vs $11.25 per 1M blended tokens
  • Faster: GPT-5 (high) and GPT-5.5 (medium) tie at 0.3 seconds latency
  • Pick GPT-5.5 (medium) when: coding quality and complex agent workflows matter more than token cost
  • Watch out: GPT-5.5 has no published benchmark results in the supplied official material, while GPT-5 has a 94.3 math index

GPT-5 (high) vs GPT-5.5 (medium)

GPT-5.5 (medium) is the stronger default for demanding software work, while GPT-5 (high) remains the better value for cost-sensitive systems.

The data brief gives GPT-5.5 (medium) a 71.5 coding index and GPT-5 (high) a 37.8 coding index. That is the clearest separation in the comparison. GPT-5.5 (medium) also leads the intelligence index, scoring 50.4 against 34.7. GPT-5 (high), however, has the only supplied math score, at 94.3, so the mathematical comparison is incomplete rather than settled.

The commercial gap points in the opposite direction. GPT-5 (high) costs $3.4375 per 1M blended tokens, compared with $11.25 for GPT-5.5 (medium). GPT-5 (high) also costs $1.25 per 1M input tokens and $10 per 1M output tokens, while GPT-5.5 (medium) costs $5 and $30. Both models show 0.3 seconds latency in the data brief, and neither has a supplied median output speed.

The product identities also differ. OpenAI describes GPT-5 as a reasoning model for coding, reasoning, and agentic tasks in GPT-5 for developers. OpenAI describes GPT-5.5 as a frontier model for complex professional work in the GPT-5.5 model page. Those descriptions support a practical conclusion: GPT-5.5 targets harder, broader workflows, but GPT-5 can still be rational when usage volume dominates the decision.

Executive summary for developers

GPT-5.5 (medium) is the capability pick, but GPT-5 (high) is the operationally safer choice for teams that cannot justify higher variable spend.

GPT-5.5 (medium) leads the supplied intelligence and coding measurements. Its coding index of 71.5 is far above GPT-5 (high) at 37.8. In production, that gap matters most when the model must understand several files, plan changes across a repository, preserve interfaces, and recover from tool errors. The score does not prove success on every codebase, but it is strong evidence for prioritizing GPT-5.5 in difficult coding evaluations.

GPT-5 (high) has two practical advantages beyond price. Its API documentation identifies a stable gpt-5 alias, while the fixed snapshot gpt-5-2025-08-07 is marked Deprecated in the GPT-5 model documentation. GPT-5.5 uses the direct ID gpt-5.5, and the supplied OpenAI model directory does not list it as deprecated. That makes GPT-5.5 the cleaner choice for a new integration, even though it is more expensive.

Neither model should be selected solely from community impressions. A Reddit report describes GPT-5 as useful for small bug fixes but less complete for full application and UI generation, based on an uncontrolled personal test in Tried GPT-5 Here Are My First Impressions. A separate Reddit report describes GPT-5.5 inside Claude Code as suitable for multi-repository and terminal-agent work, but it provides no reproducible benchmark in No one is talking about using GPT-5.5 inside Claude Code. These reports are useful signals, not performance guarantees.

Performance: what the chart does not show

GPT-5.5 (medium) is the better performance bet for repository-scale coding, although the supplied evidence does not establish its speed or reliability in production.

The coding-index gap is large enough to affect model routing. A higher coding result should matter when a request requires planning, code edits, test interpretation, and coordination across multiple files. It should matter less for short transformations, simple bug explanations, or deterministic extraction, where a cheaper model may already meet the acceptance threshold.

The intelligence-index gap points in the same direction. GPT-5.5 (medium) scores 50.4 versus 34.7 for GPT-5 (high). That supports using GPT-5.5 for tasks with ambiguous requirements, competing constraints, or several dependent decisions. It does not establish that GPT-5.5 will produce fewer regressions, because the supplied material contains no controlled error-rate study.

The speed result is less decisive. Both models have 0.3 seconds latency in the data brief, and neither has a supplied median output-tokens-per-second value. Developers therefore cannot infer that GPT-5.5 will feel slower or faster during long generations. The missing speed measurement is a material evidence gap.

GPT-5 has published external benchmark results, including 74.9% on SWE-bench Verified and 88% on Aider polyglot, with the stated evaluation conditions documented in GPT-5 for developers. The supplied GPT-5.5 model page publishes no comparable benchmark result or test method. GPT-5.5 may still be better, as the data brief suggests, but the cross-source comparison is not methodologically symmetrical.

GPT-5 (high)GPT-5.5 (medium)
37.8
ARTIFICIAL ANALYSIS CODING
71.5
34.7
ARTIFICIAL ANALYSIS INTELLIGENCE
50.4
94.3
ARTIFICIAL ANALYSIS MATH
Performance: what the chart does not show · Data provided by Artificial Analysis; live values use the current catalog.

Cost: when the cheaper model can become expensive

GPT-5 (high) is the clear cost winner, but GPT-5.5 (medium) can be cheaper overall if its stronger coding output prevents retries and human rework.

The blended comparison prices GPT-5 (high) at $3.4375 per 1M tokens and GPT-5.5 (medium) at $11.25. GPT-5.5 (medium) therefore requires a materially higher success rate to justify its price through avoided attempts. That tradeoff depends on the workload, because a low-cost model can lose its advantage when agents need repeated corrections, extra tool calls, or manual review.

Output-heavy workloads expose the largest direct gap. GPT-5 (high) costs $10 per 1M output tokens, while GPT-5.5 (medium) costs $30. Systems that generate long patches, explanations, test plans, or structured records should measure output volume separately from request count. Input-heavy workloads still favor GPT-5 (high), at $1.25 versus $5 per 1M input tokens.

GPT-5.5 adds a special long-context risk. The official GPT-5.5 model page and OpenAI pricing page state that inputs above 272K tokens trigger higher pricing for the full session. That rule can reverse an apparently reasonable architecture when an agent repeatedly sends large repository context. Prompt caching, selective file retrieval, and staged context should be tested before adopting GPT-5.5 for very large inputs.

The supplied materials do not provide token counts, retry rates, or human-review costs for either model. No honest break-even point can therefore be calculated. Teams should compare cost per accepted task, not cost per API request, using their own representative workload.

GPT-5 (high)GPT-5.5 (medium)
$1.25
Input Pricing
$5
$10
Output Pricing
$30
$3.438
Blended Price / 1M tokens
$11.25

GPT-5 (high) leads on 3 of 3 metrics

Cost: when the cheaper model can become expensive · Data provided by Artificial Analysis; live values use the current catalog.

Recommendation by workload

GPT-5.5 (medium) is the recommended starting point for new, high-value coding agents, while GPT-5 (high) fits predictable workloads with strict cost limits.

Choose GPT-5.5 (medium) when the model must work across repositories, use terminal-style tools, interpret incomplete requirements, or make coordinated changes. The supplied coding index of 71.5 supports that choice. Its official model page also lists functions, structured outputs, file search, web search, hosted shell, computer use, MCP, and other tools at GPT-5.5 model page. Those capabilities make it a strong candidate for agent orchestration, provided the application validates every edit.

Choose GPT-5 (high) when requests are shorter, traffic is high, or the task can tolerate a narrower capability ceiling. Its blended price of $3.4375 is substantially below GPT-5.5 (medium) at $11.25. GPT-5 also supports function calling, structured outputs, streaming, and constrained custom tools, according to GPT-5 for developers. These features are sufficient for many production pipelines that already have strong retrieval, tests, and guardrails.

Do not treat gpt-5-5-medium as an API model ID. The official documentation says medium is a reasoning-effort setting for gpt-5.5, not a separate model name. Route requests to gpt-5.5 and configure the reasoning effort explicitly through the documented API surface.

For migration planning, GPT-5.5 has the cleaner current lifecycle signal. GPT-5's fixed snapshot is marked Deprecated in GPT-5 model documentation, while GPT-5.5 does not appear in the supplied OpenAI deprecations list. This does not guarantee long-term availability, so production systems should still isolate model IDs, maintain evaluation fixtures, and keep a fallback route.

Before you choose

GPT-5.5 (medium) is the safer capability-first choice, but the evidence remains incomplete for speed, reliability, and mathematical breadth.

The most important unresolved issue is benchmark comparability. GPT-5 has official results with disclosed conditions, while GPT-5.5 has no supplied official benchmark score. The data brief favors GPT-5.5 on coding and intelligence, but developers should validate the exact tasks that determine business value.

The second unresolved issue is real-world failure behavior. Community posts suggest useful coding workflows for both models, yet neither report provides a controlled task set, sample count, or reproducible measurement. A team should treat those reports as hypothesis generators.

The third issue is context economics. GPT-5.5 offers a much larger documented context window, but large inputs can trigger higher full-session charges. A bigger context window helps only when retrieval quality, prompt structure, and budget controls make that context useful.

A practical evaluation should include accepted-task rate, retry count, review time, tool-call success, regression rate, latency, and total cost. The supplied sources do not provide those measurements, so the final choice should remain conditional on a representative private test.

Sources

  1. GPT-5 for developersGPT-5 positioning, reasoning controls, tool calling, custom tools, and official benchmark methodology
  2. GPT-5 model documentationGPT-5 API alias, snapshot status, pricing, modalities, endpoints, and model limitations
  3. Tried GPT-5 Here Are My First ImpressionsUncontrolled community observations about GPT-5 coding, debugging, UI generation, and failure risks
  4. GPT-5.5 model pageGPT-5.5 model ID, reasoning configuration, context, tools, pricing rules, modalities, and benchmark evidence gap
  5. OpenAI model directoryCurrent product-line placement and GPT-5.5 model identity
  6. OpenAI pricingGPT-5.5 Standard, Batch, Flex, Fast mode, and long-context pricing rules
  7. OpenAI deprecationsChecking whether GPT-5.5 is listed as deprecated
  8. No one is talking about using GPT-5.5 inside Claude CodeUncontrolled community observations about GPT-5.5 in multi-repository and terminal-agent workflows
  9. Artificial AnalysisData attribution for the supplied model comparison metrics and pricing snapshot

Your Questions about the GPT-5 (high) vs GPT-5.5 (medium) Comparison

Is GPT-5.5 (medium) better than GPT-5 (high) for coding?

GPT-5.5 (medium) is the stronger coding choice according to the supplied data, with a 71.5 coding index versus 37.8 for GPT-5 (high). The comparison still lacks a matched official benchmark method for GPT-5.5.

Which model is cheaper for production API traffic?

GPT-5 (high) is cheaper for direct API usage, priced at $3.4375 per 1M blended tokens versus $11.25 for GPT-5.5 (medium). GPT-5.5 may offset some spend only if it materially reduces retries, rework, or human review.

Do GPT-5 and GPT-5.5 have different latency?

The supplied data shows a tie, with both GPT-5 (high) and GPT-5.5 (medium) listed at 0.3 seconds latency. Neither model has a supplied median output-tokens-per-second measurement, so streaming speed remains uncertain.

Should developers call gpt-5-5-medium directly?

Developers should call gpt-5.5 and set reasoning effort to medium. The supplied official documentation does not list gpt-5-5-medium as an independent API model ID, so using it may cause routing failure.

Which model is safer for a new integration?

GPT-5.5 (medium) has the cleaner lifecycle signal for a new integration because GPT-5's fixed snapshot is marked Deprecated, while GPT-5.5 is not listed in the supplied deprecations material. Teams should still isolate model IDs and test fallback behavior.