Skip to content

GPT-5.5 (Non-reasoning)

Available

OpenAI · 2026-04-23 · 400,000 tokens

An AI model from OpenAI, suited to a broad range of AI workloads.

Supported modalities:textvideocode

Quick Overview

Text Generation4/10
Code Generation6/10
Reasoning6/10
Multimodal3/10

Benchmark Results

Scores from leading benchmark suites.

artificial analysis intelligence35.8
artificial analysis coding56.5

Performance Metrics

Latency and throughput performance.

P50 Latency
0tokens/sec

Dive Deeper

AI model analysis

GPT-5.5 (Non-reasoning) Review: Strong Coding Rank, Difficult Value Case

GPT-5.5 (Non-reasoning) Review: Strong Coding Rank, Difficult Value Case
Summary

- **Where it stands:** GPT-5.5 (Non-reasoning) ranks 54 of 202 on the Artificial Analysis Coding Index at 56.5 - **Price:** $11.25 per 1M blended tokens - **Speed:** 0.3s to first token, with output speed not reported - **Pick it when:** You need a high-ranking coding model and already depend on OpenAI’s API ecosystem - **Watch out:** Official documentation does not clearly identify the model’s context window, output limit, or dedicated capability profile

01

GPT-5.5 (Non-reasoning) is a credible coding choice, but its public specification is incomplete

GPT-5.5 (Non-reasoning) looks more compelling for software work than for broad intelligence tasks, based on its relative evaluation positions. The model ranks 54 of 202 on the Artificial Analysis Coding Index, while it ranks 94 of 578 on the Artificial Analysis Intelligence Index. Those positions make coding the clearer reason to evaluate it.

The main complication is product clarity. OpenAI’s model documentation describes GPT-5.6 Sol as a model for complex reasoning and coding, but does not provide a dedicated GPT-5.5 (Non-reasoning) profile. The same documentation does not state this model’s context window, maximum output size, supported input modalities, or official benchmark results.

OpenAI’s pricing documentation does list the API alias gpt-5.5. That confirms a publicly priced API entry, but the page does not explicitly confirm that the alias is identical to the product label “GPT-5.5 (Non-reasoning).” Developers should therefore treat the benchmark identity and API identity as related evidence, not as a fully documented equivalence.

Data provided by https://artificialanalysis.ai/.

02

The model’s strongest argument is coding rank, while the price and documentation weaken the case

GPT-5.5 (Non-reasoning) makes the strongest case for teams that value coding performance and OpenAI platform continuity more than minimum token cost. Its coding position is substantially better than its intelligence position, so a single overall label would hide the most important selection signal.

The nearby models show why the decision is not automatic:

Reference point What it suggests for GPT-5.5 (Non-reasoning)
GLM-5.1 (Non-reasoning) Matches GPT-5.5 (Non-reasoning) on the Artificial Analysis Intelligence Index, with a lower blended price of $2.135 per 1M tokens.
Grok 4.3 (low) Matches the same intelligence score, costs $1.5625 per 1M blended tokens, and reports 144.042 output tokens per second.
Kimi K2.5 (Reasoning) Matches the intelligence score, has a coding score of 46.8, and costs $1.2000000000000002 per 1M blended tokens.
MiMo-V2-Omni Has an intelligence score of 35, costs $15 per 1M blended tokens, and therefore sits closer to GPT-5.5 (Non-reasoning) on blended price.
Claude Sonnet 4.6 (Non-reasoning, High Effort) Scores 35.9 on the intelligence index and costs $6 per 1M blended tokens.

These references do not prove that another model will perform better on a specific repository. They do show that GPT-5.5 (Non-reasoning) carries a meaningful price premium over several nearby alternatives. The premium needs a concrete justification, such as better code transformation quality, stronger integration fit, or lower operational friction. The available materials do not provide controlled task-level comparisons for those factors.

03

GPT-5.5 (Non-reasoning) has a stronger coding signal than intelligence signal, but task behavior remains unverified

GPT-5.5 (Non-reasoning) should be tested first on coding workflows because its 54 of 202 coding rank is more favorable than its 94 of 578 intelligence rank. The ranking does not identify which software tasks drive that result. It cannot, by itself, establish performance on repository navigation, debugging, test generation, refactoring, migration work, or code review.

The practical interpretation is directional. Developers can reasonably place GPT-5.5 (Non-reasoning) on a coding evaluation shortlist. They should not assume that the ranking guarantees reliable autonomous implementation. The research brief contains no independently documented community benchmark, no disclosed testing method, and no verified Reddit, Hacker News, or X post focused specifically on this model. OpenAI’s model page also does not publish a dedicated GPT-5.5 capability statement or official benchmark result.

Latency adds a useful but limited signal. The data brief reports 0.3 seconds to first token, while median output speed is not reported. That makes the initial response feel measurable, but it leaves sustained generation throughput unknown. A developer building an interactive coding assistant should test both the first-token experience and the full completion time on representative prompts.

The non-reasoning label also matters for evaluation design. It suggests a different operating profile from a deliberate reasoning model, but the supplied materials do not define how that label changes tool use, planning depth, or error correction. Those behaviors require direct testing.

A sensible evaluation should use the team’s own repository tasks. Measure patch acceptance, test pass rate, review edits, tool-call recovery, and completion time. The supplied evidence supports prioritizing that test, not skipping it.

04

GPT-5.5 (Non-reasoning) is expensive unless its coding advantage reduces downstream work

GPT-5.5 (Non-reasoning) is difficult to justify on token economics alone because its blended price is $11.25 per 1M tokens. The data brief places GLM-5.1 (Non-reasoning) at $2.135, Grok 4.3 (low) at $1.5625, Kimi K2.5 (Reasoning) at $1.2000000000000002, and Claude Sonnet 4.6 (Non-reasoning, High Effort) at $6 per 1M blended tokens.

The gap matters most for high-volume workloads. Classification, routine extraction, simple transformations, and other repeated calls may not generate enough value from GPT-5.5 (Non-reasoning) to offset the premium. Coding agents can have a different cost structure because one accepted patch may replace several correction cycles. The data brief does not measure that cycle-level outcome, so the economic case remains conditional.

OpenAI’s pricing page lists standard short-context pricing of $5 per 1M input tokens, $0.50 per 1M cached input tokens, and $30 per 1M output tokens for the gpt-5.5 alias. It also lists Batch and Flex prices of $2.50 for input, $0.25 for cached input, and $15 for output in short context. Fast mode lists $12.50 for input, $1.25 for cached input, and $75 for output. These modes create different cost positions, but the documentation does not explain whether every mode applies identically to the evaluated model label.

Fast mode also lacks a listed long-context price. That omission does not prove a long-context restriction. It is simply an unresolved pricing-documentation gap. Teams should confirm the applicable mode and context pricing before committing to a budget.

05

GPT-5.5 (Non-reasoning) is worth a focused coding trial, not a default choice for every workload

GPT-5.5 (Non-reasoning) is worth piloting for code-heavy products that can convert higher model spend into fewer rejected patches and less developer review time. The recommendation follows from its relative coding rank, not from a documented superiority claim. Its coding position is 54 of 202, while its intelligence position is 94 of 578. That profile gives developers a clear hypothesis to test.

Choose GPT-5.5 (Non-reasoning) when the application needs consistent code generation, the team already operates on OpenAI APIs, and engineering time costs more than marginal token savings. The model may also suit workflows where a 0.3-second time to first token improves interactive feedback. Since output tokens per second are not reported, the full streaming experience still needs measurement.

Avoid making it the universal default for cost-sensitive, high-volume calls. Several adjacent models have the same intelligence score of 35.4 at lower blended prices. Grok 4.3 (low) also reports 144.042 output tokens per second, although that single metric does not establish better end-to-end performance. Claude Sonnet 4.6 (Non-reasoning, High Effort) has an intelligence score of 35.9 at $6 per 1M blended tokens, which makes it a relevant economic reference.

The most important unresolved issue is operational documentation. OpenAI’s model documentation does not state the model’s context window, maximum output tokens, adjustable parameters, input modalities, or deprecation status. Developers should confirm those details before architecture decisions.

The final decision should follow a repository-based bake-off. Compare accepted patches, test outcomes, correction turns, latency, and total cost on the same task set. Promote GPT-5.5 (Non-reasoning) only if its coding advantage appears in those measures.

06

What developers should verify before adopting GPT-5.5 (Non-reasoning)

GPT-5.5 (Non-reasoning) requires direct validation of model identity, limits, and coding behavior before production adoption. The public evidence confirms a priced gpt-5.5 alias, but it does not fully connect that alias to every attribute of the evaluated model label. OpenAI’s pricing documentation confirms the alias and its listed modes, while OpenAI’s model documentation provides no dedicated GPT-5.5 (Non-reasoning) specification.

Developers should verify the exact API model name, context window, maximum output size, tool behavior, supported input modalities, and applicable pricing mode. They should also measure sustained output speed because the data brief reports time to first token but does not report median output tokens per second.

The evidence is sufficient to prioritize GPT-5.5 (Non-reasoning) for a coding trial. It is not sufficient to claim a universal best model, a documented successor relationship, or a known failure profile. No reliable, method-disclosed community review was found for this specific model.

Frequently asked questions

Is GPT-5.5 (Non-reasoning) a good choice for coding?

GPT-5.5 (Non-reasoning) is a reasonable coding candidate because it ranks 54 of 202 on the Artificial Analysis Coding Index, but developers should validate repository-level accuracy before adopting it broadly.

Is GPT-5.5 (Non-reasoning) cost-effective?

GPT-5.5 (Non-reasoning) is cost-effective only when its coding performance reduces review and correction work enough to offset its $11.25 per 1M blended-token price.

How fast is GPT-5.5 (Non-reasoning)?

GPT-5.5 (Non-reasoning) has a reported 0.3-second time to first token, but the supplied data does not report median output tokens per second, so sustained generation speed remains unknown.

What information is missing from the official documentation?

GPT-5.5 (Non-reasoning) lacks a dedicated official specification covering its context window, maximum output tokens, adjustable parameters, input modalities, official benchmarks, and explicit deprecation status.

Should developers use GPT-5.5 (Non-reasoning) as a general default?

GPT-5.5 (Non-reasoning) should not be a universal default because several nearby models match its intelligence score at lower blended prices, while its task-level advantages remain unverified.

Sources

  1. OpenAI ModelsOfficial model positioning, GPT-5.6 product recommendations, and the absence of a dedicated GPT-5.5 (Non-reasoning) capability profile.
  2. OpenAI PricingThe gpt-5.5 API alias and standard, Batch, Flex, and Fast mode pricing.
  3. Artificial AnalysisAttribution for the benchmark, ranking, latency, and pricing data supplied in the data brief.

Published: