Skip to content

GPT-5.4 (xhigh)

Available

OpenAI · 2026-03-05 · 400,000 tokens

An AI model from OpenAI, suited to a broad range of AI workloads.

Supported modalities:textvideocode

Quick Overview

Text Generation5/10
Code Generation7/10
Reasoning6/10
Multimodal4/10

Benchmark Results

Scores from leading benchmark suites.

artificial analysis intelligence53.1
artificial analysis coding71.1

Performance Metrics

Latency and throughput performance.

P50 Latency
0tokens/sec

Dive Deeper

AI model analysis

GPT-5.4 (xhigh) Developer Review: Strong Quality, Selective Value

GPT-5.4 (xhigh) Developer Review: Strong Quality, Selective Value
Summary

- **Where it stands:** GPT-5.4 (xhigh) ranks 19 of 578 on the Artificial Analysis Intelligence Index at 51.4 - **Price:** $5.625 per 1M blended tokens - **Speed:** 0.3s to first token; output tokens per second are not reported in the data brief - **Pick it when:** You need a model ranked 22 of 202 on coding with 0.3s first-token latency for difficult developer work - **Watch out:** Sessions above 272K input tokens trigger 2x input and 1.5x output pricing multipliers

01

GPT-5.4 (xhigh) is a strong but selective developer model

GPT-5.4 (xhigh) is a strong but no longer obvious default for demanding developer workloads.

OpenAI positions GPT-5.4 for complex professional work, coding, tool use, and computer operation (GPT-5.4 model documentation, OpenAI release announcement). The supplied Artificial Analysis snapshot places GPT-5.4 (xhigh) at 19 of 578 on the Intelligence Index, with a score of 51.4, and at 22 of 202 on the Coding Index, with a score of 71.1 (Artificial Analysis snapshot).

Those positions make GPT-5.4 (xhigh) a serious candidate for difficult reasoning, code work, and tool-rich workflows. They do not prove that it will win every repository task, produce the fastest long answer, or justify its price in high-volume applications.

GPT-5.4 uses the stable API alias gpt-5.4. OpenAI describes a 1,050,000-token context window and a maximum output of 128,000 tokens, while the model documentation lists text and image input, text output, streaming, function calling, structured outputs, and several Responses API tools (GPT-5.4 model documentation). Data provided by https://artificialanalysis.ai/.

02

The practical tradeoff versus adjacent models

GPT-5.4 (xhigh) deserves serious consideration, but its ranking does not make it the best default for every developer workload.

The closest-model data shows a crowded field. GPT-5.4 sits near several newer or lower-cost choices, so selection should depend on workflow fit, output quality under your prompts, and total request volume. The current OpenAI model overview also places GPT-5.4 outside the newest model choices, even though the dedicated model page still documents direct API use (OpenAI models overview, GPT-5.4 model documentation).

Reference model What the supplied snapshot changes Selection implication
GPT-5.6 Luna (max) A newer, much lower-cost reference with nearly matching index readings GPT-5.4 needs a workflow or output-quality advantage to justify selection
GPT-5.6 Terra (xhigh) A newer OpenAI option with a lower blended price and closely related measured performance Teams should test both under the same prompts before standardizing
GLM-5.2 (max) A lower-cost non-OpenAI alternative with weaker supplied index readings GPT-5.4 is easier to defend for quality-sensitive work, but not for every budget
Muse Spark 1.1 (xhigh) A lower-cost option with coding results close to GPT-5.4 in the snapshot GPT-5.4 needs observed task quality to earn its premium
Claude Opus 5 (Adaptive Reasoning, Low Effort) A higher-priced reference with lower supplied index readings GPT-5.4 offers a stronger measured value case in this narrow comparison

The evidence supports a shortlist decision, not a universal verdict. The snapshot does not reveal how these models behave on your repository, tools, prompt style, or error budget.

03

What the rankings mean for real developer work

GPT-5.4 (xhigh) delivers front-rank measured performance, with the strongest case in broad intelligence rather than guaranteed coding dominance.

A position of 19 of 578 on the Artificial Analysis Intelligence Index places GPT-5.4 (xhigh) near the front of a broad comparison field. That is useful evidence for tasks that combine planning, synthesis, judgment, and several dependent steps. The coding position of 22 of 202 is also a leading result, but it should be read as evidence for serious evaluation rather than proof of universal repository reliability (Artificial Analysis snapshot).

In practical terms, GPT-5.4 (xhigh) belongs in early tests for debugging, architectural reasoning, code transformation, structured analysis, and tool-assisted work. OpenAI explicitly presents the model for professional knowledge work, coding, tool calling, and computer operation (OpenAI release announcement). The model documentation also lists function calling, structured outputs, and Responses API tools, which matter for production systems that need controlled interactions (GPT-5.4 model documentation).

The data brief reports a 0.3s latency value, but it does not report median output tokens per second. That gap matters. First-token latency can describe how quickly a response begins, but it cannot establish how quickly a long answer, code patch, or multi-step agent run finishes.

A Hacker News comment also describes GPT-5.4 latency as inconsistent in higher-reasoning situations, but it provides no request size, network conditions, or statistical method (latency discussion). Treat that feedback as a test prompt, not a service-level expectation.

04

The xhigh label changes how the evidence should be read

GPT-5.4 (xhigh) changes character across reasoning settings, so the xhigh ranking should not be read as a universal production baseline.

OpenAI describes none, low, medium, high, and xhigh as reasoning effort settings for the gpt-5.4 API model. The documentation identifies none as the default and treats xhigh as a parameter, not a separate model alias (GPT-5.4 model documentation).

That distinction creates an important evaluation boundary. A team using the default setting may see different cost, latency, and answer behavior from the supplied xhigh snapshot. A team using xhigh may obtain stronger answers on difficult tasks while accepting longer or more expensive runs. The Artificial Analysis result is therefore most useful when your production configuration matches the measured configuration (Artificial Analysis snapshot).

Prompt and context organization also deserve controlled testing. A Hacker News participant reported testing GPT-5.4 on a frontend feature, but the discussion does not disclose the feature requirements, code size, repetition count, or scoring method (frontend testing comment). That report can inform test design, but it cannot establish a general coding result.

The right conclusion is narrow: GPT-5.4 (xhigh) has strong evidence for shortlist inclusion, while the evidence remains insufficient for a blanket claim about every coding workflow or reasoning setting.

05

The price makes workload shape part of model quality

GPT-5.4 (xhigh) is expensive enough that output volume and context length can decide whether its quality pays back.

The supplied blended price is $5.625 per 1M tokens. Standard pricing lists $2.50 per 1M input tokens and $15.00 per 1M output tokens (OpenAI API pricing). That mix makes the blended figure easy to misread. Short, concise responses may fit the average, while verbose reasoning, code generation, and agent loops can push spend toward the output rate.

The cost risk becomes sharper for long-context systems. OpenAI states that once input exceeds 272K tokens, the entire session is charged at 2x the input price and 1.5x the output price (GPT-5.4 model documentation). A large repository, repeated tool history, or long document chain can therefore change the economics before the model reaches its maximum context window.

The adjacent-model snapshot reinforces the point. GPT-5.6 Luna (max), GPT-5.6 Terra (xhigh), GLM-5.2 (max), and Muse Spark 1.1 (xhigh) all provide lower-cost reference points in the supplied data (Artificial Analysis snapshot). GPT-5.4 should earn its premium through better completed-task quality, lower correction effort, or better integration with the surrounding system.

Community feedback supports caution without proving a universal cost problem. One long-term user described substantial spending during sustained GPT-5.4 use and also said results depended on context, configuration, and prompts (long-term usage comment). That is a credible reason to add budget monitoring, but it is not a reproducible cost benchmark.

GPT-5.4 is a poor fit for routine, high-volume generation when a cheaper adjacent model passes the same acceptance tests. It is easier to justify for difficult requests where a failed attempt, manual correction, or weak tool decision costs more than the token premium.

06

Who should choose GPT-5.4 (xhigh)

GPT-5.4 (xhigh) is a selective buy for difficult, tool-rich work where correctness matters more than minimum token cost.

Choose GPT-5.4 (xhigh) when the workload combines hard reasoning with code, structured outputs, function calls, or computer interaction. OpenAI explicitly supports these use cases and exposes them through its API documentation and release positioning (GPT-5.4 model documentation, OpenAI release announcement). The measured positions of 19 of 578 for intelligence and 22 of 202 for coding make it a reasonable first candidate for difficult evaluations (Artificial Analysis snapshot).

Workload Recommendation Reason
Difficult analysis with tools Start with GPT-5.4 (xhigh) The model combines strong measured results with broad official tool support
Repository debugging or code transformation Test GPT-5.4 against adjacent models Its coding rank is strong, but the snapshot does not prove universal repository success
High-volume routine generation Prefer a lower-cost model if quality remains acceptable GPT-5.4’s output price can make repeated verbose responses expensive
Audio or video input Do not choose GPT-5.4 for the input path The official model documentation lists text and image input, not audio or video input
Fine-tuning-dependent product design Do not choose GPT-5.4 as the primary model OpenAI lists fine-tuning as unsupported for this model
Very long-context applications Use only with explicit cost controls Sessions above 272K input tokens receive higher pricing multipliers

GPT-5.4 is not the newest choice in the current OpenAI model overview, so new systems should compare it with the newer options before locking in a default (OpenAI models overview). The final verdict is positive but conditional: GPT-5.4 is worth using when your acceptance tests reward difficult-task quality, while lower-cost models deserve priority for predictable volume.

Evidence remains insufficient for a claim that GPT-5.4 has a single stable coding failure mode. The supplied research also lacks a reproducible independent test that isolates GPT-5.4 across prompts, reasoning settings, repositories, and tool environments. Teams should measure those variables directly.

07

Before adopting GPT-5.4, validate the configuration

GPT-5.4 (xhigh) needs workload-specific validation before teams turn strong benchmark ranks into a blanket production choice.

The strongest evidence supports shortlist inclusion, not automatic adoption. The ranking data is useful for comparing broad intelligence and coding position. The official documentation clarifies the model’s tools, context limits, pricing trigger, and reasoning settings (GPT-5.4 model documentation).

The remaining gaps are practical. Output throughput is not reported in the data brief. Community comments are anecdotal and use undisclosed prompts or task methods (Hacker News discussion). The research also does not establish a universal, repeatable coding failure pattern. A serious evaluation should therefore keep prompts, repositories, tool permissions, reasoning effort, and acceptance criteria fixed while tracking completed-task quality and spend.

Frequently asked questions

Is GPT-5.4 (xhigh) a separate model?

GPT-5.4 (xhigh) is GPT-5.4 with the xhigh reasoning effort setting, not a separate API model alias. OpenAI identifies the stable API alias as gpt-5.4 and offers none, low, medium, high, and xhigh reasoning settings (GPT-5.4 model documentation).

Is GPT-5.4 (xhigh) good for coding?

GPT-5.4 (xhigh) is a strong coding candidate, ranking 22 of 202 on the Artificial Analysis Coding Index at 71.1. That position supports serious testing, but it does not prove success across every repository, language, prompt, or tool environment (Artificial Analysis snapshot).

Is GPT-5.4 (xhigh) worth the cost?

GPT-5.4 (xhigh) is worth the cost for difficult work where better task completion can offset a higher token bill. Its supplied blended price is $5.625 per 1M tokens, while output is priced at $15.00 per 1M tokens, so routine high-volume use needs a cheaper-model comparison (OpenAI API pricing).

How fast is GPT-5.4 (xhigh)?

GPT-5.4 (xhigh) has a reported latency value of 0.3s to first token in the supplied data. Median output tokens per second are not reported, so the available evidence cannot establish completion time for long answers or agent runs (Artificial Analysis snapshot).

Does GPT-5.4 support long-context applications?

GPT-5.4 supports a 1,050,000-token context window and a maximum output of 128,000 tokens according to OpenAI’s model documentation. Teams must still budget carefully because sessions above 272K input tokens receive 2x input and 1.5x output pricing multipliers (GPT-5.4 model documentation).

Sources

  1. Artificial AnalysisAll supplied ranking, score, blended price, and latency values used in the article.
  2. GPT-5.4 ModelReasoning settings, API alias, context window, output limit, supported inputs, tools, limitations, and long-context pricing trigger.
  3. Models | OpenAI APICurrent OpenAI model overview and GPT-5.4's position relative to newer model choices.
  4. Pricing | OpenAI APIStandard GPT-5.4 input and output pricing.
  5. Introducing GPT-5.4OpenAI's positioning of GPT-5.4 for professional work, coding, tool use, and computer operation.
  6. GPT 5.4 in practice – Stinks?Anecdotal community feedback about configuration sensitivity and mixed practical experience.
  7. Hacker News comment 47704353Anecdotal feedback about sustained-use cost and dependence on context and configuration.
  8. Hacker News comment 47686482An anecdotal frontend-task testing report and its methodological limitations.
  9. Hacker News comment 47704323Anecdotal feedback about latency during higher-reasoning use.

Published: