Skip to content

AI model analysis

Artificial Analysis Intelligence Index Model Ranking for Developers

A developer-focused reading of the Artificial Analysis Intelligence Index leaderboard, covering aggregate capability, coding performance, cost tradeoffs, production limits, and model selection.

Artificial Analysis Intelligence Index Model Ranking for Developers
Summary

- **Top of the board:** Claude Opus 5 (Adaptive Reasoning, Max Effort) at 60.7 - **Best value:** DeepSeek V4 Flash 0731 (Reasoning, Max Effort), 49.9 at $0.17500000000000002 per 1M blended tokens - **Biggest gap:** Intelligence Index, Claude Opus 5 at 60.7 is higher than DeepSeek V4 Pro at 44.3 - **Pick GPT-5.6 Terra (max) when:** throughput matters, with 55 at $4.500000000000001 and 144.252 median output tokens per second - **Watch out:** The 60.7 score comes from a text, English index and does not directly represent multimodal or multilingual ability

01

Claude Opus 5 leads the benchmark

Claude Opus 5 (Adaptive Reasoning, Max Effort) leads the supplied Artificial Analysis Intelligence Index snapshot at 60.7.

The result makes Anthropic’s highest-ranked configuration the clearest choice for developers seeking the strongest aggregate score in this table. The lead is a benchmark result, not a promise that every repository, tool chain, or production workflow will behave the same way.

The index combines broad capability areas, while the data snapshot also exposes coding scores, blended token prices, median output speed, and latency for comparison. That makes the leaderboard useful for shortlisting models before task-specific validation.

Data provided by https://artificialanalysis.ai/.

02

What the Intelligence Index measures

The Artificial Analysis Intelligence Index is best read as a broad capability signal with an agent-oriented scope.

Artificial Analysis describes a composite evaluation covering agents, coding, scientific reasoning, and general tasks. The methodology is available at https://artificialanalysis.ai/methodology/intelligence-benchmarking. The evaluation page at https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index presents score, per-task cost, output token count, and decoding time, so developers can inspect quality and operating tradeoffs together.

That scope answers which models perform well under the benchmark, but it leaves several production questions open. The research brief does not establish a reliable relationship between index score and repository maintenance, live incident response, or long-running agent workflows. It also does not quantify how provider, region, rate limits, caching, or reasoning budgets change actual spend and latency.

The evaluation is text-based and English, so the index should not stand alone for Chinese, multimodal, voice, vision, or structured tool-calling requirements. Use it to narrow the field, then test representative tasks with the same tools and acceptance criteria as production.

03

Full model ranking

The Artificial Analysis Intelligence Index ranking preserves every supplied score from rank 1 through rank 50.

Rank Model Intelligence Index Coding Index Math Index
1 Claude Opus 5 (Adaptive Reasoning, Max Effort) 60.7 78 Not reported
2 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) 60.1 77 Not reported
3 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) 59.9 76.5 Not reported
4 Claude Opus 5 (Adaptive Reasoning, High Effort) 58.9 76.5 Not reported
5 GPT-5.6 Sol (max) 58.9 77.4 Not reported
6 GPT-5.6 Sol (xhigh) 57.7 78.3 Not reported
7 Kimi K3 (max) 57.1 76.2 Not reported
8 Claude Opus 5 (Adaptive Reasoning, Medium Effort) 56.3 74.3 Not reported
9 GPT-5.6 Sol (high) 55.9 77.2 Not reported
10 Claude Opus 4.8 (Adaptive Reasoning, Max Effort) 55.7 74.3 Not reported
11 GPT-5.6 Terra (max) 55 76.7 Not reported
12 GPT-5.5 (xhigh) 54.8 74.9 Not reported
13 Grok 4.5 (high) 53.8 72.4 Not reported
14 GPT-5.6 Sol (medium) 53.6 76.3 Not reported
15 Claude Opus 4.7 (Adaptive Reasoning, Max Effort) 53.5 73.6 Not reported
16 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) 53.4 71.5 Not reported
17 GPT-5.5 (high) 53.1 71.6 Not reported
18 GPT-5.6 Terra (xhigh) 51.6 70.6 Not reported
19 GPT-5.4 (xhigh) 51.4 71.1 Not reported
20 GPT-5.6 Luna (max) 51.2 71.4 Not reported
21 GLM-5.2 (max) 51.1 68.8 Not reported
22 Claude Opus 5 (Adaptive Reasoning, Low Effort) 50.6 66.9 Not reported
23 Muse Spark 1.1 (xhigh) 50.6 71.3 Not reported
24 GPT-5.5 (medium) 50.4 71.5 Not reported
25 Gemini 3.5 Flash (high) 50.2 70.1 Not reported
26 Gemini 3.6 Flash (high) 50.1 69.2 Not reported
27 DeepSeek V4 Flash 0731 (Reasoning, Max Effort) 49.9 69.1 Not reported
28 GPT-5.6 Sol (low) 49.4 69.7 Not reported
29 GPT-5.6 Luna (xhigh) 49.1 68.6 Not reported
30 GPT-5.6 Terra (high) 49 67.1 Not reported
31 Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) 47.2 63 Not reported
32 Kimi K3 (low) 46.6 72 Not reported
33 Gemini 3.1 Pro Preview 46.5 68.8 Not reported
34 GPT-5.6 Luna (high) 46.1 63.3 Not reported
35 Qwen3.7 Max 46 66 Not reported
36 GPT-5.6 Terra (medium) 45.6 64.7 Not reported
37 Gemini 3.5 Flash (medium) 45.4 Not reported Not reported
38 MiniMax-M3 44.4 58.6 Not reported
39 DeepSeek V4 Pro (Reasoning, Max Effort) 44.3 59.4 Not reported
40 GPT-5.3 Codex (xhigh) 44.3 Not reported Not reported
41 Kimi K2.6 44.2 61.8 Not reported
42 Motif 3 (Beta) 44.1 62 Not reported
43 Claude Opus 4.6 (Adaptive Reasoning, Max Effort) 43.7 Not reported Not reported
44 GPT-5.5 (low) 43.5 60.9 Not reported
45 DeepSeek V4 Pro (Reasoning, High Effort) 43.1 58.7 Not reported
46 Muse Spark 43.1 58.6 Not reported
47 Claude Opus 4.7 (Non-reasoning, High Effort) 42.7 Not reported Not reported
48 GPT-5.2 (xhigh) 42.2 Not reported 99
49 MiMo-V2.5-Pro 42.2 60.2 Not reported
50 Kimi K2.7 Code 41.9 60.8 Not reported

The Intelligence Index column determines the displayed order. Coding and math values appear only where the supplied data includes them. The table keeps benchmark scores separate from interpretation so developers can inspect the underlying ranking directly.

04

Performance: aggregate intelligence is not coding quality

Claude Opus 5 (Adaptive Reasoning, Max Effort) leads broad capability, while GPT-5.6 Sol (xhigh) shows why coding and overall rankings can diverge.

Claude Opus 5 records 60.7 on Intelligence Index and 78 on Coding Index. GPT-5.6 Sol (xhigh) records 57.7 on Intelligence Index and 78.3 on Coding Index. The second model therefore presents a stronger coding result in the supplied fields without taking the overall lead. The reason is structural: the broader index covers agent, scientific, general, and coding work, so a specialist advantage can be diluted by other task families.

That pattern matters for developers. A high overall rank supports a broad shortlist, but it does not tell an engineering team whether the model will configure an environment, install dependencies, complete a debugging loop, improve test coverage, or understand a long-lived codebase. The research brief identifies those failure modes as information gaps rather than settled findings.

Claude Opus 5 max and xhigh also score 60.7 and 60.1, showing that reasoning effort settings can move a model’s position. The supplied data does not isolate how reasoning tokens, thinking budgets, or maximum output limits cause that movement.

Sebastian Raschka’s analysis at https://sebastianraschka.com/llm-architecture-gallery/aa-intelligence-index/ warns that model scores reflect training data, post-training, tool use, reasoning settings, and test-time compute. Teams should interpret leaderboard movement as a combined model-and-evaluation-setting signal.

05

Cost: price changes the practical shortlist

DeepSeek V4 Flash 0731 is the strongest cost-first candidate in the supplied snapshot, while GPT-5.6 Terra (max) offers a higher overall score at a higher blended price.

DeepSeek V4 Flash 0731 records 49.9 at $0.17500000000000002 per 1M blended tokens. GPT-5.6 Luna (max) records 51.2 at $0.45 and 175.726 median output tokens per second. GPT-5.6 Terra (max) records 55 at $4.500000000000001 and 144.252 median output tokens per second. Claude Opus 5 records 60.7 at $10 and 60.088 median output tokens per second. These examples show a real selection tradeoff between broad score, spend, and output rate.

Cost is not a final bill. The benchmark’s blended price is a test input, not a production guarantee. Actual totals can change with provider, region, rate limits, retries, cached input, tool calls, prompt shape, and reasoning budgets. The research brief does not quantify those effects.

Latency is listed as 0.3 in the supplied rows, but that does not answer time to useful answer for a multi-step tool workflow. Developers should measure end-to-end success and cost with their own prompts.

LinkedIn commentary at https://www.linkedin.com/pulse/announcing-artificial-analysis-intelligence-index-vzl8c uses GPT-5.5, Claude Opus 4.8, and DeepSeek V4 Pro to frame quality and cost as a joint choice, but that commentary does not prove universal savings. A LocalLLaMA discussion at https://www.reddit.com/r/LocalLLaMA/comments/1rljbix/artificial_analysis_intelligence_index_vs_weighted_model_size/ likewise shows developers connecting the index to parameter size, effective parameters, and local deployment feasibility. That community signal is useful for hardware planning, not proof of stable production coding performance.

06

Recommendations by developer workload

GPT-5.6 Terra (max) fits throughput-sensitive developer workflows when the supplied score, price, and output rate match the application’s priorities.

Choose Claude Opus 5 (Adaptive Reasoning, Max Effort) when broad benchmark leadership is the primary requirement. Choose GPT-5.6 Sol (xhigh) when coding is the dominant task and the coding score deserves more weight than the aggregate position. Choose DeepSeek V4 Flash 0731 for budget-constrained batch inference, especially when a lower token price matters more than leading the overall table. Choose Gemini 3.5 Flash (high) when output rate is the main constraint.

These recommendations are starting hypotheses, not final procurement decisions. Each model should be tested against representative repositories, tool permissions, failure recovery, and review standards. The available research does not provide enough reproducible evidence about sustained development speed, coding failure rates, debugging behavior, or long-cycle agent reliability.

For Chinese, multimodal, voice, vision, or structured tool-calling applications, pair this benchmark with specialized evaluations. The English text scope is too narrow to answer those product questions by itself.

07

Version control and production blind spots

The Artificial Analysis Intelligence Index requires version pinning before teams compare model decisions over time.

The research brief notes that task sets and weights can change the ranking, and the community discussion at https://www.reddit.com/r/singularity/comments/1q8b3pp/big_change_in_artificialanalysisai_benchmarks/ flags visible leaderboard shifts after benchmark changes. Model variants and effort settings also appear as separate rows, so teams must record the exact model label, effort setting, provider route, prompt, tool access, and evaluation date.

Raschka’s article at https://sebastianraschka.com/llm-architecture-gallery/aa-intelligence-index/ separates the official total score from category profiles, so a profile chart cannot be used to reconstruct the official total.

Open-weight comparisons add another blind spot. The LocalLLaMA discussion connects model size and MoE effective parameters with the index, but the discussion does not establish a stable relationship between parameter count and codebase outcomes.

For production agents, add task-level tests for environment setup, dependency installation, debugging loops, test coverage, long-context retrieval, and failure recovery. The supplied research does not yet provide enough reproducible evidence for those outcomes.

08

Artificial Analysis Intelligence Index FAQ

The Artificial Analysis Intelligence Index supports model selection, but the benchmark cannot answer every production question.

Developers should treat the FAQ below as a decision boundary: the index helps compare candidates, while workflow tests establish whether a candidate is reliable for a specific product.

Artificial Analysis’s methodology at https://artificialanalysis.ai/methodology/intelligence-benchmarking defines the benchmark scope, and the evaluation page at https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index supplies the comparative fields used in this article.

Frequently asked questions

What does the Artificial Analysis Intelligence Index measure?

The Artificial Analysis Intelligence Index measures broad text-based English capability across agent, coding, scientific reasoning, and general tasks, producing a composite signal for comparing model configurations under the published evaluation setup.

Does the highest-ranked model guarantee the best developer experience?

No, Claude Opus 5 (Adaptive Reasoning, Max Effort) leads this snapshot, but the sources do not establish that the highest benchmark score guarantees better repository maintenance, incident response, or long-running agent performance.

Which model is the best cost-first choice?

DeepSeek V4 Flash 0731 is the strongest cost-first candidate in this snapshot because it combines a 49.9 Intelligence Index score with $0.17500000000000002 per 1M blended tokens, subject to production validation.

Should developers use this index for Chinese or multimodal products?

No, the index should not stand alone for Chinese, multimodal, voice, vision, or structured tool-calling systems because the methodology describes a text-based English evaluation scope.

Can teams compare scores across benchmark versions?

Teams should version-pin the benchmark and exact model configuration before comparing results over time because the research brief reports that task sets, weights, and model settings can shift leaderboard positions.

Sources

  1. Artificial AnalysisData attribution for the benchmark snapshot.
  2. Artificial Analysis Intelligence IndexBenchmark scope, model scores, evaluation fields, cost, output speed, and latency.
  3. Artificial Analysis Intelligence Benchmarking MethodologyComposite benchmark design, task families, evaluation scope, and methodological limitations.
  4. Announcing Artificial Analysis Intelligence IndexBenchmark direction and quality-cost examples discussed in the research brief.
  5. Artificial Analysis Intelligence Index | Sebastian RaschkaDistinction between official total scores and profiles, plus factors affecting model scores.
  6. Artificial Analysis Intelligence Index vs weighted model size of open-source modelsCommunity discussion connecting benchmark scores with model size, effective parameters, and local deployment.
  7. Big Change in artificialanalysis.ai benchmarksCommunity discussion of ranking changes after benchmark updates.

Published: