Skip to content

Qwen3.5 397B A17B (Reasoning)

Available

Other · 2026-02-16 · 32,000 tokens

An AI model from Other, suited to a broad range of AI workloads.

Supported modalities:textcode

Quick Overview

Text Generation3/10
Code Generation5/10
Reasoning6/10
Multimodal3/10

Benchmark Results

Scores from leading benchmark suites.

artificial analysis intelligence34.3
artificial analysis coding48.2

Performance Metrics

Latency and throughput performance.

P50 Latency
82.671tokens/sec

Dive Deeper

AI model analysis

Qwen3.5 397B A17B (Reasoning) Review: A Low-Cost Coding Bet with Limited Evidence

Qwen3.5 397B A17B (Reasoning) Review: A Low-Cost Coding Bet with Limited Evidence
Summary

- **Where it stands:** Qwen3.5 397B A17B (Reasoning) ranks 109 of 578 on the Artificial Analysis Intelligence Index at 33.7 - **Price:** $1.35 per 1M blended tokens - **Speed:** 64.156 output tokens per second, 0.3s to first token - **Pick it when:** You need inexpensive reasoning-oriented coding capacity and can validate outputs in your own workflow - **Watch out:** Public documentation, stable API details, failure reports, and community testing were not identified in the supplied research Data provided by https://artificialanalysis.ai/

01

Qwen3.5 397B A17B (Reasoning) is an inexpensive model with a stronger coding position than its general intelligence rank suggests.

Qwen3.5 397B A17B (Reasoning) is most interesting as a budget-oriented developer experiment, not as a fully documented production default. The available data places the model at rank 109 of 578 on the Artificial Analysis Intelligence Index, with a score of 33.7, and rank 75 of 202 on the Artificial Analysis Coding Index, with a score of 48.2 (Artificial Analysis).

That ranking pattern supports a cautious interpretation. Coding is the clearer reason to test this model. The evidence does not establish whether the model is reliable for repository-scale changes, tool use, long-running agents, or production-critical reasoning. The supplied research found no verifiable official announcement, developer documentation, pricing page, stable API alias, current callable status, replacement guidance, or reliable community testing for this model.

The practical verdict is therefore conditional. Qwen3.5 397B A17B (Reasoning) may deserve a place in a controlled coding evaluation because its measured coding rank is better than its general intelligence rank. Developers should treat availability, operational behavior, and failure modes as unresolved until their own provider and test suite answer those questions.

02

Qwen3.5 397B A17B (Reasoning) offers a low-cost trial profile, but the evidence base is too thin for an unconditional recommendation.

Qwen3.5 397B A17B (Reasoning) makes the strongest case where cost matters and human or automated validation is already available. Its blended price is $1.35 per 1M tokens, while its coding position is rank 75 of 202 on the Artificial Analysis Coding Index (Artificial Analysis). That combination is attractive for experiments, batch coding assistance, and candidate generation.

The main weakness is not a documented technical failure. It is the absence of enough public information to judge production risk. The supplied research found no official product page, stable API identity, callable-status confirmation, community discussion, or tested limitation specific to this model. That gap makes deployment planning harder than the price suggests.

Decision area Qwen3.5 397B A17B (Reasoning) Nearby reference point
General evaluation Rank 109 of 578 The listed nearby models share the score of 33.7 on this index
Coding evaluation Rank 75 of 202 KAT Coder Pro V2 is listed at 59.5
Blended price $1.35 per 1M tokens GLM-4.7 is listed at $1, and MiniMax-M2.5 is listed at $0.525
Selection posture Test first, then decide Cheaper or better-ranked options may change the decision

The comparison is directional, not a substitute for task-level testing. The model’s value depends on whether its coding score transfers to the developer’s languages, repository patterns, and acceptance tests.

03

Qwen3.5 397B A17B (Reasoning) is more compelling for coding evaluation than for broad model selection.

Qwen3.5 397B A17B (Reasoning) earns its clearest performance case from a coding rank of 75 of 202, rather than from its broader intelligence rank of 109 of 578 (Artificial Analysis). The difference in ranking position suggests that developers should begin testing with programming tasks, not assume uniform strength across every reasoning workload.

A coding index can support model discovery, but it cannot answer several questions that matter in a real codebase. The available materials do not show how the model handles ambiguous requirements, unfamiliar frameworks, multi-file edits, test repair, dependency changes, security-sensitive code, or explanations that must remain consistent across a long interaction. No official documentation or reliable community record was found for those scenarios.

The model’s measured speed is also relevant to workflow design. Qwen3.5 397B A17B (Reasoning) records 64.156 median output tokens per second and 0.3s latency to first token (Artificial Analysis). Those values make interactive trials plausible, but they do not establish end-to-end task speed. Total completion time still depends on prompt size, output length, retries, tool calls, validation, and provider behavior.

Developers should evaluate it with a small acceptance suite. Include code generation, bug fixing, test writing, refactoring, and explanation tasks. Record pass rates, edit size, rejected suggestions, retry frequency, and review time. Those measurements would resolve the evidence gap more directly than the index alone.

The conclusion can change if the workload is primarily general reasoning, mathematical work, or autonomous repository operation. The supplied data provides a nearby math score for some reference models, but no math score for Qwen3.5 397B A17B (Reasoning). That comparison is therefore insufficient for a math-led selection.

04

Qwen3.5 397B A17B (Reasoning) is cheap enough to test seriously, but not automatically cheap enough to operate blindly.

Qwen3.5 397B A17B (Reasoning) has a blended price of $1.35 per 1M tokens, with input priced at $0.6 and output priced at $3.6 (Artificial Analysis). The input-output spread matters because reasoning workflows can produce substantial output. A low blended figure may look favorable while output-heavy tasks create a different cost profile.

The price is competitive against the listed nearby reference models, but price alone does not determine value. GLM-4.7 is listed at $1 per 1M blended tokens, while MiniMax-M2.5 and KAT Coder Pro V2 are each listed at $0.525 (Artificial Analysis). KAT Coder Pro V2 also has a listed coding score of 59.5, compared with Qwen3.5 397B A17B (Reasoning) at 48.2. That makes a direct task-level cost comparison necessary before choosing Qwen3.5 for coding volume.

A developer should prefer Qwen3.5 when its output quality reduces review work, retries, or tool calls enough to offset its output price. The supplied data does not measure any of those operational factors. It also does not verify a stable API, a current callable endpoint, or a vendor-supported service path. Therefore, the apparent economic advantage remains conditional.

For a fair trial, hold prompts, test cases, and acceptance criteria constant. Compare cost per accepted change, not only cost per generated token. If the model requires frequent correction, the cheaper token price may become a false economy. If it produces accepted code with fewer iterations, its price can be justified even without being the lowest listed option.

05

Qwen3.5 397B A17B (Reasoning) is worth a controlled coding pilot, but it should not be the sole production model on current evidence.

Qwen3.5 397B A17B (Reasoning) should be selected for a controlled coding pilot when low token cost, fast first response, and developer-side validation are the main priorities. Its coding rank of 75 of 202 provides a reasonable basis for testing, while its price of $1.35 per 1M blended tokens keeps experimentation accessible (Artificial Analysis).

The model should be avoided as an unverified production dependency when uptime, contractual support, stable naming, documented limits, or known failure behavior are mandatory. The supplied research found no reliable evidence for those areas. This does not prove that the model lacks them. It means the decision cannot safely assume they exist.

Choose Qwen3.5 for Keep another option available for
Coding experiments with strong tests Production workflows needing documented support
Cost-sensitive candidate generation Tasks where failure modes are not yet observable
Interactive developer evaluation Autonomous changes without human review
A secondary model in a routing policy A single-provider, single-model dependency

The best next step is a bounded pilot with representative repositories and explicit acceptance tests. Promote the model only if it meets the team’s quality, review-time, reliability, and availability requirements. Those thresholds are not supplied here, so they must come from the deployment context.

06

Qwen3.5 397B A17B (Reasoning) requires local validation before production adoption.

Qwen3.5 397B A17B (Reasoning) is suitable for a validation-first workflow because public evidence does not yet answer the operational questions developers need (Artificial Analysis). The available benchmark data supports a coding-focused pilot, but it does not document the model’s API stability, context behavior, failure patterns, or production support.

Developers should treat the model as a candidate to measure, not a conclusion to accept. A useful evaluation should test real repositories, include automated checks, measure accepted changes, and track retries. It should also confirm that the selected provider can actually serve the exact model identity used in the benchmark data.

Frequently asked questions

Is Qwen3.5 397B A17B (Reasoning) good for coding?

Qwen3.5 397B A17B (Reasoning) is promising enough for a controlled coding pilot because it ranks 75 of 202 on the Artificial Analysis Coding Index with a score of 48.2, but the supplied research contains no task-specific failure records or production coding evidence.

Is Qwen3.5 397B A17B (Reasoning) worth its price?

Qwen3.5 397B A17B (Reasoning) can be worth its $1.35 per 1M blended-token price when validation is cheap and accepted output reduces retries, but the supplied data does not measure review effort, correction rate, or total task cost.

How fast is Qwen3.5 397B A17B (Reasoning)?

Qwen3.5 397B A17B (Reasoning) records 64.156 median output tokens per second and 0.3s latency to first token, although these measurements do not establish total completion time for tool-using or repository-scale workflows.

Should developers use Qwen3.5 397B A17B (Reasoning) in production?

Qwen3.5 397B A17B (Reasoning) should enter production only after a provider and task-level pilot confirm availability, reliability, output quality, and support, because the supplied research found no verifiable official documentation or stable API information.

What is the biggest risk when evaluating Qwen3.5 397B A17B (Reasoning)?

The biggest evaluation risk is evidence scarcity rather than a documented model defect, because no reliable public record was supplied for the model’s official positioning, limitations, failure modes, or current callable status.

Sources

  1. Artificial AnalysisBenchmark rankings, evaluation scores, pricing figures, output speed, latency, and comparisons with nearby models.

Published: