GPT-5.6 Luna (max)
AvailableOpenAI · 2026-07-09 · 400,000 tokens
An AI model from OpenAI, suited to a broad range of AI workloads.
Quick Overview
Benchmark Results
Scores from leading benchmark suites.
Performance Metrics
Latency and throughput performance.
Dive Deeper
AI model analysis
GPT-5.6 Luna Review: Strong Reasoning at a Low API Cost

- **Where it stands:** GPT-5.6 Luna ranks 20 of 578 on the Artificial Analysis Intelligence Index at 51.2 - **Price:** $0.45 per 1M blended tokens - **Speed:** 175.726 output tokens per second, 0.3s to first token - **Pick it when:** You need coding-heavy reasoning with a model ranked 20 of 202 on the Coding Index - **Watch out:** Rank 20 of 202 is aggregate evidence, not proof of framework-level reliability
GPT-5.6 Luna at a glance
GPT-5.6 Luna is a strong default for developers who need high-end reasoning at a cost suited to frequent API calls. OpenAI positions GPT-5.6 Luna as a reasoning model for cost-sensitive, high-call-volume workloads in its model directory. The GPT-5.6 Luna model page lists text and image input, text output, the Responses API, Chat Completions API, Batch API, structured outputs, function calling, prompt caching, and several tools.
Those capabilities make Luna a practical application model rather than a narrow benchmark artifact. The supplied Artificial Analysis snapshot places Luna at rank 20 of 578 on intelligence and rank 20 of 202 on coding, while its blended price is $0.45 per 1M tokens. That combination supports a favorable default recommendation for teams serving many requests.
The evidence does not establish a universal lead for every language, framework, or agent pattern. OpenAI has not published a separate Luna benchmark table, and the supplied research found no reliable public community tests that fill the gap. Data provided by https://artificialanalysis.ai/.
The main trade-off
GPT-5.6 Luna presents the clearest quality-cost trade-off in the supplied comparison set, while the evidence stops short of proving a universal task-level lead.
The comparison data places Luna near the front of both broad intelligence and coding evaluations. Its practical advantage comes from combining that position with a lower listed blended cost than the nearby alternatives. The main question is therefore not whether Luna deserves testing. The main question is whether a more expensive model produces enough additional success on your specific workload to justify its operating cost.
| Reference model | What the comparison suggests | Practical choice |
|---|---|---|
| GLM-5.2 (max) | Similar general evaluation territory, with a higher listed blended cost | Consider only if its task behavior fits your stack better |
| GPT-5.4 (xhigh) | Comparable aggregate quality, but a much more expensive operating profile | Reserve for workloads where deeper reasoning clearly pays back |
| GPT-5.6 Terra (xhigh) | Same GPT-5.6 family, with a higher-cost and lower-throughput profile in the supplied data | Test when quality differences matter more than volume economics |
| Claude Opus 5 (Adaptive Reasoning, Low Effort) | Lower supplied rankings, higher cost, and lower measured speed | Choose only for a demonstrated workflow-specific advantage |
| Muse Spark 1.1 (xhigh) | Very close coding reference, but with a higher listed blended cost | Useful as a direct alternative in coding evaluations |
Luna also fits a broad application surface. The official model page documents API access, structured outputs, function calling, file search, web search, image input, and prompt caching. That makes Luna especially suitable for systems that need one model to handle reasoning, extraction, tool use, and code-oriented tasks.
What the rankings mean in real workloads
GPT-5.6 Luna earns rank 20 of 578 on the Artificial Analysis Intelligence Index and rank 20 of 202 on its Coding Index, supporting a serious production shortlist.
The matching positions matter because they show a consistent aggregate signal across broad intelligence and coding evaluation. Luna is not merely a general-purpose model with weak coding evidence, nor is it a coding specialist with a narrow overall profile. Developers can reasonably begin with Luna when a product mixes planning, code generation, structured reasoning, and tool calls.
The ranking still cannot answer several questions that determine production quality. The supplied data does not identify performance by programming language, framework, repository size, test style, or repair difficulty. A strong aggregate coding position therefore supports a shortlist decision, not a claim that Luna will reliably fix every repository issue. Teams should test representative prompts, inspect generated diffs, run automated tests, and measure recovery after failed tool calls.
The reported 0.3s latency and 175.726 median output tokens per second also support interactive use. Those figures suggest that Luna can return early responses quickly while sustaining high output throughput. They do not reveal tail latency, reasoning variance, queue behavior, or tool-call overhead. Applications with strict response budgets should measure complete user-visible latency rather than relying on model-only figures.
OpenAI’s GPT-5.6 Luna model page documents the model’s interfaces and capabilities, but it does not publish Luna-specific benchmark results. The supplied research also found no reliable community test set for coding behavior. That absence is the central performance limitation: Luna looks strong in aggregate, but task-level certainty remains unavailable without local evaluation.
When the low price helps, and when it does not
GPT-5.6 Luna is most compelling when reasoning quality must survive high call volume and token spend is a primary operating constraint.
The supplied blended price is $0.45 per 1M tokens. That figure gives Luna a meaningful economic advantage over nearby references such as GLM-5.2 (max), which is listed at $2.15 per 1M blended tokens. The gap matters most for assistants, code review systems, document pipelines, and agent loops that make many model calls rather than a few occasional requests.
The blended figure should not become the only budget input. Standard pricing lists short-context input at $0.20 per 1M tokens and output at $1.20 per 1M tokens, so output-heavy workflows can behave differently from input-heavy workflows. Cached prompts, repeated system instructions, and large tool results also change the cost profile. Developers should model their actual input and output mix before committing to a monthly budget. The official OpenAI pricing page provides the relevant pricing modes.
Batch and Flex options can improve economics for workloads that tolerate asynchronous processing. Fast mode carries a different price profile, so latency-sensitive systems should compare its value against normal service behavior rather than assuming faster execution is automatically cheaper.
Very long inputs also require special attention. The GPT-5.6 Luna model page describes additional billing multipliers for requests beyond a long-input threshold. That means a large-context workflow can erase part of Luna’s apparent price advantage. Retrieval, summarization, prompt caching, and selective context assembly may produce a better result than sending every available document on every request.
Luna is therefore inexpensive in the right workload shape, not universally inexpensive. High-volume, short-context, output-disciplined applications are the strongest fit. Large-context workflows and response-heavy agents need a workload-specific cost test.
Who should choose GPT-5.6 Luna
GPT-5.6 Luna is the best first deployment candidate for high-volume assistants, code agents, and document workflows that need strong reasoning without premium-model economics.
| Situation | Recommendation | Reason |
|---|---|---|
| High-volume assistant with structured responses | Start with Luna | The model combines strong aggregate rankings, low blended cost, and structured output support |
| Coding agent for repositories | Test Luna first | Its coding rank supports serious evaluation, but framework-level reliability remains unverified |
| Large document or knowledge workflow | Use Luna with retrieval and context controls | Long inputs can change cost behavior, and the documented maximum input still matters |
| Premium workflow with low call volume | Compare against GPT-5.4 or Terra | A more expensive model is justified only when local tests show a meaningful quality gain |
| Audio or video generation | Do not use Luna as the output model | The documented output modality is text, not audio or video |
| Fresh or rapidly changing information | Pair Luna with retrieval or web search | The model’s knowledge horizon cannot replace current application data |
Luna’s official tool list includes web search, file search, code execution, computer use, MCP, Tool Search, Apply Patch, and Skills. The GPT-5.6 Luna model page does not specify quotas, regional availability, latency guarantees, or error-recovery behavior for each tool. Production teams should validate those details in the exact deployment environment.
Teams should also verify the model identifier and lifecycle status before rollout. The OpenAI changelog documents GPT-5.6 series changes and alias information. The OpenAI deprecations page does not list a deprecation plan for Luna in the supplied research context.
The final recommendation is straightforward: make Luna the first model in the evaluation matrix when cost and throughput matter, then keep it only if it passes representative task tests.
What the available evidence still cannot answer
GPT-5.6 Luna still needs task-specific validation because public evidence shows strong aggregate rankings but leaves important behavior-level questions unanswered.
The available material does not establish which programming languages or frameworks Luna handles best. It does not show repository repair rates, long-context recall, tool-call recovery, refusal patterns, or consistency across repeated prompts. The absence of a Luna-specific official benchmark also makes it difficult to separate model behavior from evaluation methodology.
Reasoning configuration deserves a separate test. The broader OpenAI reasoning guide describes max reasoning effort and persistent reasoning for the GPT-5.6 series, but the Luna model page does not document every model-specific limitation. Teams should therefore verify effort settings, token behavior, and latency under their own API configuration.
Tool support is another open area. The official documentation lists many tools, but the supplied research does not confirm their quotas, regional constraints, failure modes, or production latency. A strong model ranking cannot substitute for an end-to-end agent test.
These gaps do not weaken Luna’s shortlist position. They define the boundary of what the current evidence can responsibly support.
Frequently asked questions
Is GPT-5.6 Luna good for coding agents?
GPT-5.6 Luna is a strong coding-agent candidate, but developers should confirm repository-level accuracy with their own tests because rank 20 of 202 does not identify languages, frameworks, or repair rates.
Is GPT-5.6 Luna cost-effective at scale?
GPT-5.6 Luna is cost-effective for many high-volume workloads because its blended price is $0.45 per 1M tokens, but output-heavy and long-context requests require separate budget modeling.
Does GPT-5.6 Luna support tools and structured outputs?
GPT-5.6 Luna supports structured outputs, function calling, web search, file search, image input, and other documented tools through supported OpenAI APIs, according to the model page.
Should developers use GPT-5.6 Luna for current information?
GPT-5.6 Luna should use retrieval, web search, or application-provided data for current information because the model’s documented knowledge horizon cannot guarantee coverage of later events.
Is GPT-5.6 Luna ready for production evaluation?
GPT-5.6 Luna is ready for production evaluation through the API, but teams should verify tool behavior, cost under real prompts, identifier stability, and task accuracy before making it the default model.
Sources
- Artificial Analysis model comparison dataIntelligence and coding rankings, price, latency, throughput, adjacent-model comparisons, and data snapshot attribution.
- Models | OpenAI APIOfficial model positioning and catalog context.
- GPT-5.6 Luna Model | OpenAI APIModel capabilities, APIs, tools, modalities, pricing limitations, and documented operational constraints.
- Pricing | OpenAI APIInput, output, caching, Batch, Flex, and Fast mode pricing context.
- Changelog | OpenAI APIGPT-5.6 series changes and model alias context.
- Deprecations | OpenAI APILifecycle and deprecation status checks.
- Reasoning guide | OpenAI APIReasoning effort and persistent reasoning context.
Published: