GPT-5.6 Terra (xhigh)
AvailableOpenAI · 2026-07-09 · 400,000 tokens
An AI model from OpenAI, suited to a broad range of AI workloads.
Quick Overview
Benchmark Results
Scores from leading benchmark suites.
Performance Metrics
Latency and throughput performance.
Dive Deeper
AI model analysis
GPT-5.6 Terra (xhigh) Review: Is the Balanced Frontier Model Worth Its Price?

- **Where it stands:** GPT-5.6 Terra (xhigh) ranks 18 of 578 on the Artificial Analysis Intelligence Index at 51.6 - **Price:** $4.500000000000001 per 1M blended tokens - **Speed:** 121.137 output tokens per second, 0.3s to first token - **Pick it when:** Mixed reasoning and coding workflows need a fast default, with a coding score of 70.6 - **Watch out:** GPT-5.6 Luna (max) scores 71.4 on coding, so Terra is not the automatic coding pick
GPT-5.6 Terra (xhigh): a balanced frontier model for production reasoning
GPT-5.6 Terra (xhigh) is a strong production candidate for developers who need frontier-level reasoning without choosing the most expensive nearby option (Artificial Analysis, OpenAI model directory).
OpenAI places GPT-5.6 Terra in its Frontier models category and describes the GPT-5.6 family as balancing intelligence and cost (OpenAI model directory). The supplied benchmark snapshot supports that positioning. Terra ranks near the front of a large intelligence field, while its coding result remains strong enough for mixed engineering workloads.
The reviewed label also needs careful interpretation. The documented API model is gpt-5.6-terra; xhigh is a reasoning.effort setting rather than a separate stable model ID (GPT-5.6 Terra model page, model parameters migration guide). Developers should therefore reproduce this evaluation with the stable model ID and the matching reasoning setting.
The official model page lists text and image input, structured outputs, function calling, web search, file search, prompt caching, and several tool integrations (GPT-5.6 Terra model page). That makes Terra relevant for agents, document analysis, coding assistants, and multimodal workflows.
The evidence has limits. OpenAI does not publish concrete Terra benchmark scores on the checked model pages, and the supplied research found no reproducible community tests. Data provided by https://artificialanalysis.ai/ therefore supplies the main quantitative basis for this review.
The short verdict: strong default, conditional value
GPT-5.6 Terra (xhigh) offers the clearest default case for mixed reasoning and coding workloads, but its premium makes the choice workload-dependent.
Terra is most compelling when one model must handle difficult analysis, code, tool use, and multimodal input with responsive output. The benchmark position supports a high-quality generalist thesis. It does not prove that Terra wins every coding task, every long-context workflow, or every cost-per-success comparison.
The adjacent models clarify the decision without changing Terra’s central role:
| Reference model | What the snapshot suggests | Implication for Terra |
|---|---|---|
| GPT-5.4 (xhigh) | A coding edge at a higher blended price | Terra is the better mixed-workload default unless coding quality dominates the workload |
| GPT-5.6 Luna (max) | Lower blended price, higher coding result, and faster output | Terra needs a measurable reasoning or workflow advantage to justify its premium |
| GLM-5.2 (max) | Lower blended price and faster output, with lower observed index results | Terra buys a stronger observed quality position, subject to governance and task fit |
| Claude Opus 5, Adaptive Reasoning, Low Effort | Higher blended price and lower observed index results in this snapshot | A higher price alone does not establish better value |
| Muse Spark 1.1 (xhigh) | Lower blended price, faster output, and a higher coding result | Terra’s case rests on broader reasoning position and deployment fit |
Terra is therefore worth a serious pilot, not an automatic fleet-wide standard. The supplied research contains no reliable community experience reports, so claims about coding feel, reliability, or developer ergonomics remain unverified.
Performance: what the rankings mean for real developer work
GPT-5.6 Terra (xhigh) is strongest as a broad reasoning model with credible coding ability, rather than as an unqualified coding specialist.
Terra ranks 18 of 578 on the Artificial Analysis Intelligence Index at 51.6, and 23 of 202 on the Artificial Analysis Coding Index at 70.6 (Artificial Analysis). Those placements indicate a model that performs near the front of broad evaluation fields. The practical implication is important: Terra should be considered for tasks that mix planning, explanation, implementation, and judgment, not only for isolated code completion.
The coding result is strong, but Terra does not lead the adjacent set. GPT-5.6 Luna (max), GPT-5.4 (xhigh), and Muse Spark 1.1 (xhigh) all score higher on coding in the supplied comparison (Artificial Analysis). Developers building repository agents should therefore test bug fixing, test generation, refactoring, and multi-file changes separately. A strong aggregate coding index does not reveal which of those tasks drives production cost or user satisfaction.
The latency snapshot gives Terra a useful interactive profile. The model records 0.3 seconds to first token and a median output rate of 121.137 tokens per second (Artificial Analysis). That supports chat, coding assistance, and agent loops where users need early feedback. It does not establish full task latency, because tool calls, network access, prompt size, and reasoning depth can dominate the user-visible wait.
The xhigh setting is another performance variable. OpenAI recommends higher reasoning effort only when it creates a measurable quality gain, and does not guarantee that higher effort helps every workload (model parameters migration guide). Run evaluations at the exact effort level used in production.
OpenAI’s checked model page and model directory do not publish independent Terra benchmark scores (GPT-5.6 Terra model page, OpenAI model directory). Artificial Analysis is useful evidence, but task-specific pass rates and failure modes remain insufficiently documented.
Cost: premium balanced pricing, with a narrow value case
GPT-5.6 Terra (xhigh) is priced as a premium balanced model, so its value depends on avoiding quality-related rework.
The data brief reports a blended price of $4.500000000000001 per 1M tokens (Artificial Analysis). That places Terra below GPT-5.4 (xhigh) at $5.625 and Claude Opus 5 at $10, but above GPT-5.6 Luna (max) at $0.45, GLM-5.2 (max) at $2.15, and Muse Spark 1.1 (xhigh) at $2 (Artificial Analysis). Terra is therefore not a low-cost volume choice. It is a quality-oriented middle option among nearby frontier models.
The price is defensible when better answers reduce retries, human review, tool misuse, or failed code changes. The supplied evidence does not measure any of those outcomes. It also does not provide average prompt length, cache hit rate, output mix, or cost per successful task. No exact production return estimate can be supported from this brief.
OpenAI documents separate Standard, Batch, Flex, and fast service paths, each with its own pricing rules (OpenAI API pricing). That gives developers room to match service mode to urgency, but the correct choice depends on workload scheduling and availability requirements.
Long-context workflows can also change the economics. OpenAI states that requests beyond its long-context threshold receive higher input and output billing multipliers (GPT-5.6 Terra model page). A workflow that looks affordable under a blended average can become expensive when prompts contain large repositories, repeated documents, or long agent histories.
Choose Terra for quality-sensitive interactive work. Prefer a cheaper adjacent model for high-volume tasks unless a controlled evaluation shows that Terra’s successful-task rate offsets the price difference.
Recommendation: pilot Terra as the mixed-workload default
GPT-5.6 Terra (xhigh) deserves a default pilot for mixed reasoning, coding, tool use, and multimodal workflows, not a blind fleet-wide rollout.
The adjacent-model comparison below uses the supplied Artificial Analysis snapshot:
| Candidate | Best reason to consider it | Recommendation for Terra buyers |
|---|---|---|
| GPT-5.4 (xhigh) | Stronger coding position at a higher price | Keep Terra if the workload includes substantial non-coding reasoning |
| GPT-5.6 Luna (max) | Lower cost, faster output, and stronger coding result | Move to Luna if task quality remains equal in your evaluation |
| GLM-5.2 (max) | Lower cost and faster output | Compare total operating constraints, especially quality and deployment fit |
| Claude Opus 5, Adaptive Reasoning, Low Effort | A different provider and reasoning approach | Pay more only when task-specific results justify it |
| Muse Spark 1.1 (xhigh) | Lower cost, faster output, and stronger coding result | Treat Terra’s broader intelligence position as a hypothesis to validate |
Use the documented API identity in production. OpenAI identifies gpt-5.6-terra as the model and xhigh as a reasoning setting (GPT-5.6 Terra model page, model parameters migration guide). The benchmark site’s gpt-5-6-terra-xhigh slug should be treated as an evaluation label, not as the API model ID.
Terra is a good fit when one endpoint must support Responses API, Chat Completions API, or Batch API workflows, plus structured outputs, function calling, search, files, and image input (GPT-5.6 Terra model page). If the workload runs through Amazon Bedrock, re-check the tool surface. OpenAI documents deployment differences involving Hosted File Search, remote MCP servers, Computer Use, and image generation (Amazon Bedrock support).
The final decision should use successful task completion as the gate. Test real prompts, code repositories, tool calls, long documents, and human review effort. Keep Terra if it reduces failure and rework enough to justify the premium. Choose a neighbor if it matches quality at lower cost or higher speed. The supplied research does not provide enough evidence to make that choice universally.
Questions to answer before adopting GPT-5.6 Terra
GPT-5.6 Terra (xhigh) should pass a workload-specific review before developers commit to its production role.
The official documentation supports a broad API and tool surface, but documentation alone cannot predict task success. The model page describes text and image input, structured outputs, function calling, search, file access, prompt caching, and other tools (GPT-5.6 Terra model page). Each capability should be tested in the exact endpoint and service tier that the application will use.
Context handling deserves its own check. The documented context window is larger than the maximum input allowance, so developers should not treat the entire advertised window as usable prompt space (GPT-5.6 Terra model page). Long prompts can also trigger higher billing multipliers, which makes repository agents and document workflows especially sensitive to prompt construction.
Deployment path matters as well. OpenAI’s Bedrock guidance lists missing or different tool support for several capabilities, including Hosted File Search, remote MCP, Computer Use, and image generation (Amazon Bedrock support). A direct OpenAI API result should not be assumed to transfer unchanged to Bedrock.
Finally, the available evidence does not establish a universal quality lead, a stable community consensus, or a cost-per-success advantage. Use Terra as a serious candidate, then compare it with at least one cheaper and one coding-focused neighbor on production tasks.
Frequently asked questions
Is GPT-5.6 Terra (xhigh) worth the price for developers?
GPT-5.6 Terra (xhigh) is worth piloting for mixed reasoning workloads, but the supplied evidence does not prove a universal return on its premium price. Artificial Analysis places it 18 of 578 on intelligence at 51.6, while nearby models offer lower cost or stronger coding results (Artificial Analysis). Choose it after measuring successful task completion, retries, and human review on your own prompts.
Is GPT-5.6 Terra (xhigh) good for coding?
GPT-5.6 Terra (xhigh) is a strong coding candidate, but it is not the automatic coding winner in the supplied comparison. Its coding result ranks 23 of 202 at 70.6, while GPT-5.6 Luna (max), GPT-5.4 (xhigh), and Muse Spark 1.1 (xhigh) score higher in the adjacent set (Artificial Analysis). Run repository-level tests before standardizing on Terra.
Does xhigh represent a separate GPT-5.6 Terra model ID?
GPT-5.6 Terra (xhigh) should be called through gpt-5.6-terra with reasoning.effort set to xhigh, not through a separate xhigh model suffix. OpenAI’s migration guide describes xhigh as a reasoning setting and says the general gpt-5.6 alias routes to gpt-5.6-sol (model parameters migration guide). This distinction matters for reproducibility and deployment configuration.
Is GPT-5.6 Terra (xhigh) fast enough for interactive applications?
GPT-5.6 Terra (xhigh) is fast enough for interactive API workflows in the supplied latency snapshot, with 0.3 seconds to first token and median output at 121.137 tokens per second (Artificial Analysis). Those figures do not establish end-to-end completion time when prompts trigger tools, long reasoning, or network calls, so production tests should measure full user-visible latency.
Does GPT-5.6 Terra have the same capabilities on Amazon Bedrock?
GPT-5.6 Terra (xhigh) does not expose the same tool surface across every deployment path. OpenAI’s Bedrock guidance lists restrictions affecting Hosted File Search, remote MCP servers, Computer Use, and image generation, so developers using Bedrock must validate the exact integration before migrating (Amazon Bedrock support).
Sources
- Artificial AnalysisBenchmark rankings, scores, pricing, throughput, latency, and adjacent-model comparisons.
- GPT-5.6 Terra model pageDocumented model ID, capabilities, context handling, endpoint support, and long-context billing rules.
- Model parameters migration guideThe relationship between gpt-5.6-terra and reasoning.effort xhigh, plus the gpt-5.6 alias routing rule.
- OpenAI model directoryFrontier model positioning and the absence of published Terra benchmark scores in the checked directory.
- OpenAI API pricingStandard, Batch, Flex, and fast service path pricing rules.
- Amazon Bedrock supportDeployment-specific tool and capability differences on Amazon Bedrock.
Published: