GPT-5.4 (low)
AvailableOpenAI · 2026-03-05 · 400,000 tokens
An AI model from OpenAI, suited to a broad range of AI workloads.
Quick Overview
Benchmark Results
Scores from leading benchmark suites.
Performance Metrics
Latency and throughput performance.
Dive Deeper
AI model analysis
GPT-5.4 (low) Review: Strong Benchmark Position, Unclear Product Fit

- **Where it stands:** GPT-5.4 (low) ranks 69 of 578 on the Artificial Analysis Intelligence Index at 39.1 - **Price:** $5.625 per 1M blended tokens - **Speed:** unavailable output tokens per second, 0.3s to first token - **Pick it when:** you need a high-ranked OpenAI option and can validate the low configuration through your own workload - **Watch out:** official sources do not clearly document the low configuration, its alias, or its independent benchmark method
GPT-5.4 (low) looks capable, but its identity is still under-documented
GPT-5.4 (low) is a high-ranking model configuration with enough uncertainty to require validation before production adoption. The supplied evaluation places GPT-5.4 (low) at position 69 of 578 on the Artificial Analysis Intelligence Index, with a score of 39.1. That result puts the model in a strong part of a large field, but it does not establish superiority for coding, tool use, long-context work, or agentic workflows.
The main product issue is naming clarity. OpenAI’s official model directory does not list a separate gpt-5-4-low alias. OpenAI’s official pricing page lists gpt-5.4, but does not separately explain the low configuration. Developers therefore cannot assume that the evaluated name maps cleanly to a stable, directly callable API model.
The practical verdict is conditional. GPT-5.4 (low) deserves consideration when OpenAI compatibility, a strong general intelligence ranking, and low measured first-token latency matter. It deserves less confidence when predictable documentation, transparent configuration semantics, or low operating cost are essential.
The model’s strongest case is general capability, not value for money
GPT-5.4 (low) offers a credible general-purpose capability signal, but nearby alternatives make its price difficult to justify without an OpenAI-specific requirement. Its Intelligence Index score of 39.1 is close to the scores of the supplied neighboring models, including GLM-5 (Reasoning) at 39.5, Gemini 3 Pro Preview (high) at 39.6, Qwen3.7 Plus at 39, and Agnes 2.5 Pro Alpha at 38.8.
That neighborhood matters because GPT-5.4 (low) does not stand alone as a clear quality outlier. Gemini 3 Pro Preview (high) scores slightly higher on the same intelligence index, while GLM-5 (Reasoning) also scores slightly higher. Qwen3.7 Plus and Agnes 2.5 Pro Alpha sit close enough that workflow fit, availability, and integration may matter more than the headline score.
| Decision factor | GPT-5.4 (low) | Nearby alternatives |
|---|---|---|
| General benchmark position | Strong, at 69 of 578 | Several nearby models have similar index scores |
| Provider fit | OpenAI ecosystem | Other providers offer materially different integration choices |
| Cost posture | Premium blended price | Several nearby models are listed at lower blended prices |
| Evidence quality | Independent index signal, limited official detail | The supplied brief also lacks full methodology for these comparisons |
Data provided by https://artificialanalysis.ai/ should be treated as the basis for the ranking discussion. The ranking supports serious consideration, but it does not answer whether the low configuration preserves the behavior developers expect from the broader GPT-5.4 product.
Performance is promising for first-response latency, but task-level evidence is missing
GPT-5.4 (low) appears responsive at the start of a request, but the supplied evidence cannot establish sustained throughput or task-specific reliability. The measured latency is 0.3 seconds to first token, while median output speed is unavailable. That combination supports a narrow conclusion: the model may begin responding quickly, but developers cannot infer how fast it will complete long generations or multi-step tasks.
The ranking position provides useful context. A score of 39.1 at position 69 of 578 indicates broad capability that is materially competitive across the evaluated index. It does not reveal the distribution of strengths behind that score. The supplied brief does not provide a coding index for GPT-5.4 (low), a mathematics score, a tool-use score, a long-context result, or a disclosed test method for this exact configuration.
This limitation is especially important for developer workloads. A model can rank well overall while behaving differently on repository changes, structured extraction, function calling, refusal boundaries, or iterative debugging. The supplied research found no reliable community tests for GPT-5.4 (low), and no official OpenAI page publishes separate benchmarks or known failure modes for it. The official model directory describes current OpenAI models at a product-family level, including text and image input, text output, multilingual capability, Responses API access, and official SDK support. Those statements are not specifically confirmed for the low configuration.
Choose GPT-5.4 (low) for latency-sensitive prototypes only when your own acceptance tests cover completion speed, tool correctness, and output quality. Evidence is insufficient to recommend it for high-volume generation solely because its first token arrives quickly.
GPT-5.4 (low) is expensive beside similarly ranked models
GPT-5.4 (low) becomes difficult to defend on cost when a workload can use a similarly ranked alternative. The supplied blended price is $5.625 per 1M tokens, with standard input priced at $2.5 per 1M tokens and output at $15 per 1M tokens. Those values make generated output the main cost driver for response-heavy applications.
The nearby-model data shows a wide pricing gap. Qwen3.7 Plus is listed at $0.7 per 1M blended tokens, while Agnes 2.5 Pro Alpha is listed at $0.5625. GLM-5 (Reasoning) is listed at $1.55, and Gemini 3 Pro Preview (high) is listed at $4.5. JT-4.1 Flash 236B A21B is listed at $15, so GPT-5.4 (low) is not the most expensive option in the comparison set. Its issue is that several alternatives sit near its intelligence score at a much lower blended price.
OpenAI’s pricing documentation adds another uncertainty. It describes prices for gpt-5.4, including standard, batch, flex, and fast modes, but does not separately identify gpt-5-4-low. Developers should therefore verify which billing schedule applies before building a cost forecast.
GPT-5.4 (low) can still be economical when its outputs reduce review time, improve completion quality, or avoid migration work. The supplied evidence does not quantify any of those benefits. Without workload-specific gains, the price looks like a premium for provider fit and presumed capability rather than demonstrated efficiency.
Use GPT-5.4 (low) selectively, with an explicit validation gate
GPT-5.4 (low) is best treated as a candidate for controlled evaluation, not as a default production choice. The model’s position at 69 of 578 supports testing for general assistant work, code generation, and reasoning-heavy tasks where OpenAI integration has clear operational value. The 0.3-second first-token latency also makes it reasonable to test in interactive experiences.
The recommendation changes when the main objective is cost efficiency. Qwen3.7 Plus and Agnes 2.5 Pro Alpha have nearby intelligence scores in the supplied data and lower blended prices. Their listed coding index values are 55.9 and 58.8 respectively, but GPT-5.4 (low) has no corresponding coding value in the brief. That prevents a fair coding conclusion. Developers should benchmark the exact tasks they care about instead of treating general intelligence ranking as a proxy.
| Choose GPT-5.4 (low) when | Prefer another candidate when |
|---|---|
| OpenAI compatibility materially reduces integration or migration effort | Blended token cost is the primary constraint |
| Fast initial response matters in an interactive product | Sustained output speed is a core requirement and must be proven |
| You can run private acceptance tests before launch | You need a clearly documented, stable model alias |
| General capability matters more than a specialized benchmark | Coding, mathematics, or tool-use evidence is required before selection |
The first validation gate should confirm that the evaluated configuration is callable under the intended API name. The next gate should test representative prompts, tool calls, structured outputs, retries, and long responses. The supplied research does not establish the model’s context window, maximum output length, exact low-mode parameters, or independent failure patterns. Those are decision blockers for workloads that depend on them.
What developers should verify before adoption
GPT-5.4 (low) requires documentation and workload checks before developers can make a confident production decision. The ranking data answers where the model sits in a broad intelligence comparison. It does not answer whether the name is stable, whether the configuration is directly available, or whether the measured behavior transfers to a particular application.
OpenAI’s Models documentation and Pricing documentation are the relevant official references for model availability, supported capabilities, and billing. The supplied research found no separate official entry for gpt-5-4-low, no independent official benchmark for that configuration, and no reliable community review that fills the gap.
A sensible evaluation should record success rate, correction rate, tool-call validity, structured-output compliance, response completion time, and total token spend. These measurements should come from the application’s real prompt distribution. The available evidence supports a shortlist decision, but it does not support a blanket claim that GPT-5.4 (low) is the best choice for developers.
Frequently asked questions
Is GPT-5.4 (low) an officially documented OpenAI model?
GPT-5.4 (low) is not separately documented in the supplied OpenAI model directory, which lists GPT-5.4 but does not explain the low configuration or confirm its stable API alias.
Is GPT-5.4 (low) good for coding?
GPT-5.4 (low) may be suitable for coding evaluation because its general intelligence ranking is strong, but the supplied data contains no coding score or reliable coding-specific community test.
Is GPT-5.4 (low) worth its price?
GPT-5.4 (low) is worth its premium price only when OpenAI integration, response quality, or reduced engineering effort creates measurable value that lower-priced nearby models cannot match.
Does GPT-5.4 (low) respond quickly?
GPT-5.4 (low) has a measured first-token latency of 0.3 seconds, but its median output speed is unavailable, so sustained completion speed remains unverified.
What is the biggest risk when adopting GPT-5.4 (low)?
GPT-5.4 (low) carries documentation risk because the supplied official sources do not clarify whether the low configuration is directly callable, stable, or priced separately from GPT-5.4.
Sources
- OpenAI ModelsModel directory, general capability descriptions, API access information, and the absence of a separately documented GPT-5.4 low alias.
- OpenAI API PricingGPT-5.4 pricing, billing modes, and the absence of separately documented pricing for GPT-5.4 low.
- Artificial AnalysisAttribution for the supplied intelligence index, ranking, pricing, latency, and neighboring-model data.
Published: