GPT-5.6 Terra (max)
AvailableOpenAI · 2026-07-09 · 400,000 tokens
An AI model from OpenAI, strongest at code generation, suited to a broad range of AI workloads.
Quick Overview
Benchmark Results
Scores from leading benchmark suites.
Performance Metrics
Latency and throughput performance.
Dive Deeper
AI model analysis
GPT-5.6 Terra Review: A High-Value Reasoning Model for Developers

- **Where it stands:** GPT-5.6 Terra ranks 11 of 578 on the Artificial Analysis Intelligence Index at 55 - **Price:** $4.50 per 1M blended tokens - **Speed:** 144.252 output tokens per second, 0.3s to first token - **Pick it when:** you need coding-heavy workflows where rank 6 of 202 on the Artificial Analysis Coding Index matters - **Watch out:** the broad intelligence result is rank 11 of 578, not a task-specific guarantee
GPT-5.6 Terra at a glance
GPT-5.6 Terra is a strong general-purpose reasoning choice for developers who need high coding quality without paying frontier-model prices. OpenAI describes GPT-5.6 Terra as a reasoning model designed to balance intelligence and cost, with a position roughly comparable to the earlier GPT-5 mini tier. (GPT-5.6 Terra Model)
OpenAI places GPT-5.6 Terra in its Frontier models product line, which confirms that the model targets serious production workloads rather than lightweight experimentation. (Models) The model joined the GPT-5.6 family on 2026-07-09. (OpenAI API Changelog)
The practical appeal is broad developer coverage. GPT-5.6 Terra accepts text and image inputs, returns text, and supports the Responses API, Chat Completions API, and Batch API. OpenAI also lists structured outputs, function calling, file search, web search, Code Interpreter, hosted shell, Apply Patch, MCP, and other tools. (GPT-5.6 Terra Model)
The model therefore fits coding agents, document analysis, tool-driven workflows, and applications that need a fast reasoning loop. The main qualification is evidence quality. OpenAI does not publish a Terra-specific benchmark table, latency guarantee, or success-rate guarantee on the official model page. The assessment below uses the independent ranking snapshot instead. Data provided by https://artificialanalysis.ai/.
The short verdict
GPT-5.6 Terra offers the clearest value case in its comparison set: strong coding, near-front general intelligence, and a lower listed blended price. The Artificial Analysis snapshot places Terra at rank 11 of 578 for general intelligence and rank 6 of 202 for coding, which makes its coding result the strongest part of the buying argument. (Artificial Analysis)
| Reference model | Where GPT-5.6 Terra is stronger | What the reference may still offer |
|---|---|---|
| GPT-5.5 (xhigh) | Lower listed cost, stronger coding position, and reported output speed | A familiar higher-spend baseline for teams already using GPT-5.5 |
| Claude Opus 4.8 | Lower listed cost and stronger coding position | A close general-intelligence alternative that deserves task-level testing |
| GPT-5.6 Sol (high) | Lower listed cost and faster reported output | Slightly higher index results for teams willing to pay more |
| Grok 4.5 (high) | Stronger intelligence and coding positions in the snapshot, with faster reported output | Lower listed cost for applications where budget dominates |
| Claude Opus 5 (medium) | Stronger coding position | Slightly stronger general-intelligence result in the snapshot |
The correct interpretation is not that Terra wins every task. Terra earns a strong default position because its coding ranking, response profile, and price align unusually well for developer workloads. Teams choosing between close models should still compare repository changes, test repair, structured output accuracy, tool calls, and refusal behavior on their own data. The supplied evidence supports a confident pilot recommendation, not a universal quality claim.
Performance: what the rankings mean in practice
GPT-5.6 Terra looks strongest as a fast coding-oriented reasoning model, with broad intelligence results that remain near the front of the evaluated field. Terra ranks 6 of 202 on the Artificial Analysis Coding Index, while it ranks 11 of 578 on the Artificial Analysis Intelligence Index. (Artificial Analysis)
That coding position matters for work where the model must transform instructions into repository changes. It supports a favorable hypothesis for code generation, test creation, bug repair, code review, and tool-mediated edits. It does not prove that Terra will understand every codebase, preserve every invariant, or produce correct patches without verification. Index rankings screen model capability; they do not replace task-level acceptance tests.
The response profile strengthens the interactive case. The snapshot reports 144.252 median output tokens per second and 0.3 seconds of latency. Those figures support editor assistants, agent loops, and applications where users wait for a visible response before taking the next action. They do not establish sustained throughput under every concurrency pattern, because the supplied materials do not include a Terra-specific latency guarantee or load curve.
GPT-5.6 Terra also has a broad reasoning and tool surface. OpenAI documents standard and pro reasoning modes, several reasoning-effort settings, and reasoning context that can remain available across turns. (Reasoning models) That design should help multi-step workflows, but it introduces a budgeting concern: reasoning tokens consume the same output budget controlled by max_output_tokens. The reasoning guide warns that an overly low limit can produce an incomplete response before visible text appears. (Reasoning models)
The evidence gap is important for developers. No reliable Terra-specific public failure catalog was supplied, and the official documentation does not quantify success rates by task. Treat Terra’s ranking as a strong screening signal, then validate it with real repositories, real tools, and real completion criteria.
Cost: attractive by default, less predictable at large context
GPT-5.6 Terra is economically attractive for mixed, interactive workloads, but large-context billing can weaken that advantage. The snapshot reports a $4.50 blended price per 1M tokens, which is lower than the listed GPT-5.5, GPT-5.6 Sol, and Claude references. (Artificial Analysis)
That price creates a useful middle position. Terra costs more than Grok 4.5 in the comparison set, but its stronger coding and intelligence rankings make the premium defensible for applications where failed generations, manual review, or slow agent loops carry operational cost. A low token price is not automatically a low system cost if the model needs more retries or produces weaker patches.
Large documents require extra care. OpenAI’s model page applies higher input and output billing rules to very large requests, so Terra’s headline blended price should not be treated as a universal rate for every context size. (GPT-5.6 Terra Model) This matters for repository-wide analysis, long legal files, persistent agent histories, and repeated document evaluations. Cache-aware prompt design can improve economics when the same instructions or reference material recur.
Asynchronous workloads have another option. OpenAI’s pricing page lists lower Batch and Flex rates, which can make offline evaluation, bulk extraction, and scheduled code analysis more economical than interactive execution. (Pricing) Fast mode carries a separate premium, so teams should reserve it for latency-sensitive paths rather than apply it to every request.
The safest cost policy is simple: measure cost per accepted task, not cost per generated token alone. Set max_output_tokens high enough for the selected reasoning effort, monitor incomplete responses, and test large-context behavior before committing to a fixed budget. Public materials do not provide a Terra-specific cost curve by task, so workload measurement remains necessary.
Recommendation: who should choose GPT-5.6 Terra
GPT-5.6 Terra should be the default candidate for coding-heavy applications that need fast responses and controlled model spend. The recommendation is strongest for teams building developer tools, coding agents, repository assistants, automated review systems, and structured workflows that combine reasoning with function calls.
Choose GPT-5.6 Terra when:
- Coding quality matters more than selecting the lowest listed token price.
- Users need a short response wait and high output throughput.
- The application needs text reasoning over image inputs, documents, or visual references.
- Tool calls, structured outputs, file search, web search, or patch-oriented workflows are central to the product.
- You want a model positioned between cheaper general options and higher-cost frontier configurations.
Consider GPT-5.6 Sol (high) when a slightly higher coding or intelligence result justifies higher spend and slower reported output. Consider Grok 4.5 (high) when unit cost dominates and your own evaluation accepts its lower comparison-set rankings. Claude Opus 4.8 and Claude Opus 5 remain credible alternatives when organization standards, existing integrations, or general reasoning behavior outweigh Terra’s coding advantage.
Do not choose Terra as the sole endpoint for native image, audio, or video generation. OpenAI documents image input but text output for this model. (GPT-5.6 Terra Model) Also avoid treating the ranking as proof of reliability on high-risk tasks without a domain-specific evaluation.
Adoption risk appears manageable for a staged rollout. The supplied official deprecation page does not list gpt-5.6-terra as deprecated, while the model documentation identifies gpt-5.6-terra as the stable model ID and current snapshot. (Deprecations) Pin that ID, log incomplete responses, and keep a fallback until your acceptance data confirms the choice.
Questions to answer before adoption
GPT-5.6 Terra deserves a staged production trial because its rankings are strong while model-specific public evidence remains limited. The right pilot should use representative repositories, real tool calls, realistic prompt histories, structured output validation, and image inputs if the product depends on them.
Before approval, test how reasoning effort changes completion quality, latency, and token use. Confirm that max_output_tokens leaves enough room for hidden reasoning and visible output, because the official reasoning guide documents incomplete responses when the budget is exhausted. (Reasoning models)
Test large-context economics separately from ordinary requests. The official model page documents higher billing behavior for very large inputs, so a successful quality test does not automatically prove a sustainable cost profile. (GPT-5.6 Terra Model)
Finally, verify the integration surface your product needs. GPT-5.6 Terra supports text and image inputs, text output, structured outputs, function calling, and several hosted tools, but teams needing generated media require another endpoint. Pin the stable model ID, record the snapshot used for evaluation, and revisit the decision when Terra-specific public benchmark or failure evidence becomes available.
Frequently asked questions
Is GPT-5.6 Terra good for coding?
GPT-5.6 Terra is a strong coding candidate, because its Artificial Analysis coding index ranks it 6 of 202 while its general intelligence ranking remains 11 of 578. That combination favors repository changes, test repair, code review, and tool-driven development, but teams should still run their own task set.
Is GPT-5.6 Terra worth its price?
GPT-5.6 Terra is worth its price when faster responses and coding quality matter more than the lowest token bill. Its $4.50 blended price sits below the listed OpenAI and Anthropic nearby references, but Grok 4.5 is cheaper, so cost-only buyers should compare acceptance rates.
Does the large context make GPT-5.6 Terra economical for large documents?
GPT-5.6 Terra does not make very large-context work automatically economical, because the official model page applies higher billing rules to very large inputs. Cache-aware prompts and Batch or Flex processing can improve the business case, but actual savings depend on workload reuse.
Does GPT-5.6 Terra support multimodal output?
GPT-5.6 Terra accepts text and image inputs but produces text only, so it fits visual analysis and document workflows but cannot serve as a native image, audio, or video generation endpoint.
Is the available evidence sufficient for production use?
GPT-5.6 Terra has enough evidence for a focused production pilot, but not enough public evidence for a universal quality guarantee. The broad rankings are strong, while official Terra-specific benchmark and community failure data remain limited.
Sources
- GPT-5.6 Terra ModelModel positioning, stable model ID, modalities, APIs, tools, reasoning capabilities, and large-context billing limitations
- ModelsOpenAI Frontier models product-line positioning
- PricingStandard, Batch, Flex, cached-input, and Fast mode pricing guidance
- OpenAI API ChangelogGPT-5.6 Terra release timing and GPT-5.6 family information
- Reasoning modelsReasoning modes, effort settings, reasoning context, output-token budgeting, and incomplete-response behavior
- DeprecationsCurrent deprecation status for gpt-5.6-terra
- Artificial AnalysisBenchmark rankings, scores, pricing snapshot, latency, output speed, and adjacent-model comparisons
Published: