Skip to content

GPT-5.6 Luna (low)

Available

OpenAI · 2026-07-09 · 400,000 tokens

An AI model from OpenAI, suited to a broad range of AI workloads.

Supported modalities:textvideocode

Quick Overview

Text Generation3/10
Code Generation4/10
Reasoning6/10
Multimodal3/10

Benchmark Results

Scores from leading benchmark suites.

artificial analysis intelligence33.9
artificial analysis coding44.2

Performance Metrics

Latency and throughput performance.

P50 Latency
143.382tokens/sec

Dive Deeper

AI model analysis

GPT-5.6 Luna (low) Review: Fast, Affordable, and Better for Coding Than General Intelligence

GPT-5.6 Luna (low) Review: Fast, Affordable, and Better for Coding Than General Intelligence
Summary

- **Where it stands:** GPT-5.6 Luna (low) ranks 118 of 578 on the Artificial Analysis Intelligence Index at 33.3 - **Price:** $0.45 per 1M blended tokens - **Speed:** 166.399 output tokens per second, 0.3s to first token - **Pick it when:** You need a fast, cost-sensitive coding and production automation model - **Watch out:** Public sources do not confirm the model’s context limit, failure patterns, or the meaning of the low tier

01

GPT-5.6 Luna (low) is a production-oriented model with strong speed and moderate overall capability

GPT-5.6 Luna (low) is best understood as a throughput-focused model for developers who value response speed and operating cost. OpenAI positions the official gpt-5.6-luna model for cost-sensitive, high-throughput workloads in its model directory. The official documentation also lists text and image input, text output, multilingual capability, and vision support. Developers can call the model through the Responses API and OpenAI Client SDK.

The available evaluation data supports that positioning, but it also narrows the likely role of this model. GPT-5.6 Luna (low) ranks 118 of 578 on the Artificial Analysis Intelligence Index, with a score of 33.3. Its coding position is stronger, ranking 85 of 202 with a score of 44.2. Data provided by https://artificialanalysis.ai/. These results suggest a model that may fit coding assistance, structured automation, and high-volume application features better than demanding general reasoning work.

The central qualification is evidence quality. OpenAI does not publish official benchmark scores for this model in the reviewed documentation. Public sources also do not establish the context window, maximum output length, supported parameter limits, or model-specific failure cases. The official API name is gpt-5.6-luna, while gpt-5-6-luna-low appears in the evaluation dataset rather than as a confirmed public API model ID. Developers should validate the deployment configuration before treating the low variant as a separately selectable endpoint.

02

The model’s main advantage is operational value, not frontier-level intelligence

GPT-5.6 Luna (low) offers a more attractive production trade-off than its intelligence rank alone suggests. The evaluation data places its coding score above its general intelligence score, while its measured throughput is high and its blended price is low. That combination matters for applications that process many short or medium requests, especially where a small quality gap is acceptable.

Decision area GPT-5.6 Luna (low) Nearby reference point Practical reading
General intelligence Lower-ranked than GPT-5.5 Instant and LongCat 2.0 in the supplied comparison set GPT-5.5 Instant and LongCat 2.0 each score 33.5 Do not select Luna for maximum broad reasoning quality
Coding Coding score of 44.2 LongCat 2.0 scores 45.3 Luna is close to the nearby coding reference while costing less
Cost $0.45 blended per 1M tokens LongCat 2.0 costs $1.30 blended Luna has a substantial cost advantage in the supplied set
Math signal Not provided for Luna Grok 4 scores 92.7 and Gemini 3 Pro Preview scores 86.7 Math-heavy selection remains unresolved

The comparison is directional, not a substitute for task-specific testing. Grok 4, MiMo-V2-Flash, Gemini 3 Pro Preview, GPT-5.5 Instant, and LongCat 2.0 are useful reference points because they appear in the supplied nearest-model set. They do not prove that Luna will behave better or worse on a particular codebase, language, tool chain, or prompt format.

The strongest conclusion is narrower: GPT-5.6 Luna (low) looks compelling when request volume, latency, and cost matter more than leading general intelligence. The evidence is insufficient to determine whether its lower tier changes reasoning depth, instruction following, or output reliability relative to the official gpt-5.6-luna listing. OpenAI’s model documentation does not separately explain that tier.

03

GPT-5.6 Luna (low) should perform well in responsive coding workflows, but ranking does not guarantee reliability

GPT-5.6 Luna (low) is a plausible choice for interactive coding workflows because its coding rank is materially better than its overall intelligence rank and its measured response speed is high. The model ranks 85 of 202 on the Artificial Analysis Coding Index with a score of 44.2. It also records 166.399 median output tokens per second and 0.3 seconds to first token in the supplied data from Artificial Analysis.

For developers, that profile points toward repository navigation, code explanation, test drafting, routine refactoring, API glue code, and structured transformations. These tasks benefit from quick feedback and reasonable code competence. The result does not establish that Luna can safely handle architectural changes, subtle concurrency issues, security-sensitive code, or long multi-step debugging sessions without review.

The coding comparison is close rather than dominant. LongCat 2.0 has a supplied coding score of 45.3, only slightly above Luna’s 44.2, while Luna has the stronger cost position. That makes the selection depend on error cost. If a small quality difference creates expensive human review, LongCat or a stronger model may be preferable. If the workload contains many routine requests, Luna’s speed and price can outweigh a narrow benchmark gap.

The largest performance unknowns concern model limits and behavior. The reviewed OpenAI model page does not disclose a verified context window, maximum output length, complete parameter support, or known failure modes. The research brief also found no reliable community tests with disclosed methods. Developers should therefore treat benchmark rank as a screening signal, not as evidence of stable production quality across every task.

04

GPT-5.6 Luna (low) is inexpensive enough for high-volume use, but long prompts and generated output can change the economics

GPT-5.6 Luna (low) is financially attractive when applications keep prompts compact, reuse context, and avoid unnecessary generation. The supplied blended price is $0.45 per 1M tokens, based on a 3 to 1 input-output mix. The corresponding official short-context Standard prices are $0.20 for input and $1.20 for output per 1M tokens, according to OpenAI Pricing.

That price profile favors workloads where the model handles frequent classification, code suggestions, extraction, summarization, test generation, and agent substeps. It also makes Luna a reasonable candidate for routing systems. A product can send routine work to Luna and reserve a more capable model for failures, ambiguity, or high-impact decisions. The supplied comparison places Luna below LongCat 2.0’s $1.30 blended price and far below Grok 4’s $6 blended price, which strengthens the case for volume-sensitive deployments.

The economics become less obvious when requests are output-heavy. Luna’s official output price is higher than its input price, so verbose answers, large patches, and repeated agent traces consume budget faster than short responses. The official pricing page also lists separate short-context and long-context prices, but the reviewed documentation does not define the token boundary between them. Developers cannot make a reliable long-context cost forecast from the public page alone.

Batch and Flex pricing are lower than Standard, while Fast mode is higher. That creates a simple operational split: use lower-cost modes for asynchronous workloads, Standard for ordinary interactive traffic, and Fast mode only when its latency benefit justifies the added price. The public sources do not explain whether the evaluation measurements correspond to any particular service mode. Cost testing should therefore use the exact deployment mode planned for production.

05

Choose GPT-5.6 Luna (low) for fast, repetitive developer work when quality gates can catch mistakes

GPT-5.6 Luna (low) is recommended for cost-sensitive coding assistance and high-throughput automation, not as a default model for every reasoning task. Its coding position, speed, and blended price form a coherent production profile. The supplied data shows a coding rank of 85 of 202, 166.399 median output tokens per second, and a blended price of $0.45 per 1M tokens. Data provided by https://artificialanalysis.ai/.

A good fit includes code completion, boilerplate generation, issue triage, documentation drafts, test scaffolding, structured extraction, and internal tools with human review. Luna is also a candidate for first-pass agent steps, provided the system can detect low-confidence results and escalate them. Its 0.3-second time to first token supports interfaces where users expect immediate feedback.

Luna is a weaker default for math-intensive work, high-stakes decisions, security review, novel architecture, or tasks where an incorrect answer costs more than additional inference. The comparison set includes stronger math signals for Grok 4 and Gemini 3 Pro Preview, but Luna has no supplied math score. The data therefore does not justify a math recommendation.

Before adoption, run a task sample from the target codebase. Measure compilation success, test success, tool-call correctness, review edits, escalation rate, and total cost. Include long prompts and verbose outputs because the official pricing documentation separates short and long context pricing without defining the boundary. Also confirm that the production endpoint is the supported gpt-5.6-luna alias described in OpenAI’s model directory, since the evaluated low-tier identifier is not confirmed there.

06

Questions developers should answer before deployment

GPT-5.6 Luna (low) is suitable for a controlled pilot when the workload values speed and price more than maximum reasoning depth. The evidence supports a practical trial, but public documentation leaves several deployment questions unanswered. Developers should confirm endpoint availability, context behavior, output limits, and failure handling before committing the model to critical paths.

The safest rollout uses measurable quality gates. Start with reversible workloads, compare Luna with one stronger fallback, and record both successful outputs and human corrections. This approach tests the actual trade-off that the supplied rankings cannot settle: whether the model’s lower cost offsets additional review or retry work.

Frequently asked questions

Is GPT-5.6 Luna (low) a good model for coding?

GPT-5.6 Luna (low) is a credible coding model for routine developer tasks because it ranks 85 of 202 on the supplied coding index and responds quickly. Developers should still test repository-specific reliability before using it for security-sensitive or architectural changes.

Who should choose GPT-5.6 Luna (low) instead of a stronger model?

GPT-5.6 Luna (low) suits teams processing many repetitive requests where low cost and fast feedback matter more than the highest general intelligence score. Teams should choose a stronger model when review is expensive or failures carry material business risk.

Is GPT-5.6 Luna (low) officially available as a separate API model?

GPT-5.6 Luna (low) is not confirmed as a separate public API model ID in the reviewed OpenAI documentation. OpenAI lists the stable alias gpt-5.6-luna, so developers should verify the exact endpoint before implementation.

Does GPT-5.6 Luna (low) support long context?

GPT-5.6 Luna (low) has no verified public context limit in the reviewed sources. OpenAI lists separate short-context and long-context pricing, but the pricing page does not state the token boundary or confirm the evaluated low-tier configuration.

What is the biggest risk when deploying GPT-5.6 Luna (low)?

GPT-5.6 Luna (low) has an evidence gap around failure patterns, maximum output limits, supported parameters, and context behavior. The absence of public community tests means teams must establish those limits through their own evaluation before production use.

Sources

  1. OpenAI ModelsVerifying the official gpt-5.6-luna alias, workload positioning, supported input and output modalities, multilingual and vision capabilities, API availability, and the absence of separately documented low-tier details.
  2. OpenAI PricingVerifying Standard, Batch, Flex, and Fast mode pricing, plus the separate short-context and long-context pricing presentation.
  3. Artificial AnalysisAttributing the supplied ranking, benchmark, latency, throughput, and blended pricing data.

Published: