GPT-5.6 Luna (high)
AvailableOpenAI · 2026-07-09 · 400,000 tokens
An AI model from OpenAI, suited to a broad range of AI workloads.
Quick Overview
Benchmark Results
Scores from leading benchmark suites.
Performance Metrics
Latency and throughput performance.
Dive Deeper
AI model analysis
GPT-5.6 Luna (high) Review: Strong Coding Value, Unclear API Status

- **Where it stands:** GPT-5.6 Luna (high) ranks 34 of 578 on the Artificial Analysis Intelligence Index at 46.1 - **Price:** $0.45 per 1M blended tokens - **Speed:** 164.222 output tokens per second, 0.3s to first token - **Pick it when:** you need high-volume coding or general-purpose generation with a low blended token cost - **Watch out:** the evidence does not confirm that GPT-5.6 Luna (high) is an independent callable API model
GPT-5.6 Luna (high) at a glance
GPT-5.6 Luna (high) looks like a strong value candidate for developers who prioritize throughput and operating cost over maximum benchmark leadership. The model ranks 34 of 578 on the Artificial Analysis Intelligence Index and 37 of 202 on the Artificial Analysis Coding Index, placing it in a high-performing tier across both broad intelligence and coding evaluation. Data provided by https://artificialanalysis.ai/
The most important qualification is operational rather than statistical. OpenAI’s official model directory lists gpt-5.6-luna, but the research brief found no separate official entry for GPT-5.6 Luna (high) or gpt-5-6-luna-high (OpenAI Models). The pricing documentation also does not confirm the high variant as an independent API model (OpenAI Pricing).
For a developer evaluating this model, the practical conclusion is simple: the benchmark profile is attractive, but the API identity must be verified before implementation. This review therefore treats the measured model as a useful performance and cost reference, while separating that evidence from claims about production availability.
Executive summary for developers
GPT-5.6 Luna (high) is most compelling when a developer needs a capable model for repeated coding and generation tasks at a comparatively low blended price. Its intelligence ranking places it ahead of the nearby GPT-5.6 Terra (medium) reference on the Artificial Analysis Intelligence Index, while its coding position suggests meaningful strength without establishing category leadership.
The closest reference models reveal the tradeoff. Qwen3.7 Max and Gemini 3.1 Pro Preview offer higher coding scores in the supplied comparison set, but their listed blended prices are higher. GPT-5.6 Terra (medium) is a closer same-vendor reference with a lower intelligence score and a higher blended price. Gemini 3.5 Flash (medium) is another cost-conscious reference, though the supplied brief does not include its coding score.
| Decision factor | GPT-5.6 Luna (high) implication |
|---|---|
| General capability | Strong enough to rank near the front of a large evaluation field |
| Coding | Competitive, but some nearby models score higher |
| Economics | Attractive for high call volume and mixed input-output workloads |
| Responsiveness | Suitable for interactive workflows based on the measured latency and output rate |
| API certainty | Insufficient evidence confirms the high name as a direct public model identifier |
OpenAI’s official documentation positions the stable gpt-5.6-luna alias toward cost-sensitive, high-throughput workloads (OpenAI Models). That positioning supports the economic interpretation, but it does not prove that the measured high variant shares every documented capability or endpoint behavior.
What the benchmark position means in practice
GPT-5.6 Luna (high) offers a credible general-purpose performance profile for developers who need consistent quality across many ordinary tasks rather than a specialist model for the hardest coding problems. Ranking 34 of 578 on the Artificial Analysis Intelligence Index indicates that the model sits close to the leading group in a broad field. Ranking 37 of 202 on the Artificial Analysis Coding Index gives coding a similarly favorable position.
That combination matters for production systems. A model with strong results in both indexes can plausibly support code explanation, test generation, implementation drafts, structured extraction, technical writing, and routine agent steps within one deployment. The supplied rankings do not prove success on any particular repository, language, framework, or tool-use pattern. They do support choosing Luna as a serious baseline for an evaluation suite.
The speed profile strengthens that case. GPT-5.6 Luna (high) records 164.222 median output tokens per second and 0.3 seconds to first token in the supplied data. Those measurements favor interactive coding assistants, review tools, and applications where users see value quickly. They do not establish end-to-end application latency, because network time, queueing, prompt length, tool calls, streaming behavior, and orchestration overhead are not described.
Developers should interpret the coding ranking as evidence of broad coding capability, not as a guarantee of superior repository-level performance. The research brief found no reliable community reports describing coding experience, speed perception, model habits, or recurring failure modes. That absence makes local testing especially important for complex refactors, long-running agents, and tasks that require strict adherence to an existing codebase.
Where the performance conclusion can change
GPT-5.6 Luna (high) may be a poor coding choice when the workload rewards peak specialist performance more than balanced quality and cost. The supplied neighboring models include Qwen3.7 Max, Gemini 3.1 Pro Preview, and Kimi K3 (low), each with a higher coding score than Luna in the comparison data. That evidence suggests developers seeking the strongest coding result should benchmark those alternatives directly before standardizing on Luna.
The comparison is not a two-model verdict. The adjacent models also differ in price and output speed, and the brief does not describe evaluation prompts, task composition, variance, or statistical confidence. A higher index score therefore indicates a reason to test, not proof that another model will produce better pull requests for a specific team.
GPT-5.6 Luna (high) also has important unknowns. The research brief found no confirmed context window, maximum output length, API parameter set, official benchmark methodology, failure-mode catalog, or community testing evidence. OpenAI’s documentation describes current model families as supporting text and image input, text output, multilingual ability, and vision through the Responses API and OpenAI Client SDKs, but the brief explicitly does not prove that those statements apply to gpt-5-6-luna-high (OpenAI Models).
The safest performance plan is a focused bake-off using representative repository tasks, tool calls, long prompts, structured outputs, and regression tests. Evidence is insufficient to predict the result without that validation.
Why the price is attractive, and when it is not
GPT-5.6 Luna (high) is financially attractive for high-volume workloads because its supplied blended price is $0.45 per 1M tokens, with listed input pricing of $0.2 and output pricing of $1.2 per 1M tokens. Data provided by https://artificialanalysis.ai/
The cost advantage is most useful when traffic is steady, prompts are reasonably compact, and the application benefits from a capable model on many calls. Code review queues, documentation generation, classification with explanations, support automation, and routine agent substeps are natural candidates. The measured speed also reduces the risk that lower cost comes with visibly slow interaction.
The price becomes less compelling when a cheaper model can meet the task’s quality threshold, or when Luna’s output requires frequent retries, human correction, or escalation to a more capable model. The benchmark data alone cannot reveal that total cost. Developers should measure accepted outputs, retry rates, tool-call success, review time, and downstream defects rather than comparing token prices in isolation.
The official pricing page lists Standard, Batch, Flex, and Fast mode prices for gpt-5.6-luna, including different short-context and long-context rates (OpenAI Pricing). Those documented modes provide useful planning context, but the research brief does not confirm that the measured high variant has a separate pricing or routing identity. The blend between the data brief and the official price page should therefore be treated as a validation item before forecasting production spend.
For asynchronous workloads, the supplied official pricing information makes Batch and Flex relevant candidates. For latency-sensitive applications, Fast mode may alter the economics. Evidence is insufficient to recommend a mode without the application’s traffic pattern and API availability being confirmed.
Recommendation: who should choose GPT-5.6 Luna (high)
GPT-5.6 Luna (high) is worth testing first for developers building cost-sensitive, high-throughput systems that need strong general and coding performance from one model. Its ranking profile and token economics make it a sensible baseline for production experiments, especially where response speed and operating cost both matter.
Choose Luna when:
- the workload contains many routine coding, analysis, transformation, or technical-writing calls;
- the team values a low blended token cost;
- interactive response speed matters;
- the application can tolerate a short verification phase around model naming and routing;
- local evaluation can measure quality on real prompts and repositories.
Keep another model in the evaluation when:
- the product depends on the highest available coding score;
- tasks involve difficult multi-file changes or failure-sensitive automation;
- the workload needs a confirmed context limit or output limit;
- the team requires a clearly documented public API identifier;
- image input, multilingual behavior, or vision support is a hard requirement for this exact variant.
Qwen3.7 Max, Gemini 3.1 Pro Preview, and Kimi K3 (low) are useful comparison candidates because the supplied data shows stronger coding scores for each, while GPT-5.6 Terra (medium) provides a same-vendor reference. The decision should remain centered on Luna’s measured quality per unit cost, not on a single leaderboard position.
The final recommendation is conditional: adopt GPT-5.6 Luna (high) only after confirming that the intended identifier is callable and that representative tests reproduce acceptable quality. The research brief provides insufficient evidence to skip either check.
FAQ before you integrate
GPT-5.6 Luna (high) requires API identity validation before integration because the official documentation confirms gpt-5.6-luna, not the exact high identifier (OpenAI Models). The questions below address the main implementation risks.
Frequently asked questions
Is GPT-5.6 Luna (high) an officially documented OpenAI API model?
The available evidence does not confirm GPT-5.6 Luna (high) as an independent documented API model, because OpenAI’s model directory lists gpt-5.6-luna but not the exact high identifier (OpenAI Models).
Is GPT-5.6 Luna (high) good for coding?
GPT-5.6 Luna (high) appears strong for coding, ranking 37 of 202 on the Artificial Analysis Coding Index, although several supplied neighboring models record higher coding scores. Data provided by https://artificialanalysis.ai/
Is GPT-5.6 Luna (high) fast enough for interactive developer tools?
GPT-5.6 Luna (high) appears suitable for interactive tools because the supplied measurements show 164.222 median output tokens per second and 0.3 seconds to first token, subject to application overhead. Data provided by https://artificialanalysis.ai/
Is GPT-5.6 Luna (high) cost-effective for production?
GPT-5.6 Luna (high) is potentially cost-effective for high-volume workloads at $0.45 per 1M blended tokens, but retries, review effort, routing, and confirmed API pricing can change the total cost.
Should developers choose GPT-5.6 Luna (high) over a higher-scoring coding model?
Developers should choose GPT-5.6 Luna (high) when its lower listed blended cost and strong speed outweigh the need for the highest coding score; representative task tests should decide the tradeoff.
Sources
- OpenAI ModelsVerifying the documented model alias, official positioning, stated capabilities, API availability, and the absence of a separately confirmed GPT-5.6 Luna (high) entry.
- OpenAI PricingVerifying the listed pricing modes and documenting the distinction between official pricing for gpt-5.6-luna and the measured high variant.
- Artificial AnalysisAttributing the supplied benchmark rankings, speed measurements, latency, and model pricing snapshot.
Published: