Skip to content

Qwen3.5 397B A17B (Non-reasoning)

Available

Other · 2026-02-16 · 32,000 tokens

An AI model from Other, suited to a broad range of AI workloads.

Supported modalities:textcode

Quick Overview

Text Generation3/10
Code Generation6/10
Reasoning6/10
Multimodal3/10

Benchmark Results

Scores from leading benchmark suites.

artificial analysis intelligence32.7

Performance Metrics

Latency and throughput performance.

P50 Latency
83.02tokens/sec

Dive Deeper

AI model analysis

Qwen3.5 397B A17B Non-reasoning Review: Strong Mid-Frontier Results, With Important Evidence Gaps

Qwen3.5 397B A17B Non-reasoning Review: Strong Mid-Frontier Results, With Important Evidence Gaps
Summary

- **Where it stands:** Qwen3.5 397B A17B (Non-reasoning) ranks 126 of 578 on the Artificial Analysis Intelligence Index at 32 - **Price:** $1.35 per 1M blended tokens - **Speed:** 66.396 output tokens per second, 0.3s to first token - **Pick it when:** You need a reasonably fast general-purpose model with a mid-frontier benchmark position and predictable blended-token economics - **Watch out:** Public evidence does not verify the model’s provider, API availability, context window, coding behavior, or failure patterns

01

Qwen3.5 397B A17B Non-reasoning is a credible benchmark performer with limited product evidence

Qwen3.5 397B A17B (Non-reasoning) looks like a practical candidate for general-purpose evaluation, but its public product story is incomplete. The available data places the model at position 126 of 578 on the Artificial Analysis Intelligence Index, with a score of 32. That ranking is high enough to justify a serious trial, yet it does not establish that the model will be dependable for a specific production workload. The model has no verified vendor identity in the supplied record, and the research brief found no reliable official announcement, developer documentation, stable API alias, current availability statement, or pricing page. Developers should therefore treat the benchmark record as the strongest available evidence, not as a complete product specification.

The model’s non-reasoning label also matters. It suggests a response mode aimed at direct generation rather than visible or extended reasoning, but the supplied research does not verify how that mode behaves in coding, extraction, structured output, or long multi-step tasks. No reliable public community reports were found to confirm coding quality, perceived speed, or characteristic failure cases. That absence does not prove poor behavior. It means the selection decision needs a workload-specific validation step.

Data provided by Artificial Analysis.

02

The model’s main appeal is balanced utility, not category leadership

Qwen3.5 397B A17B (Non-reasoning) offers a balanced position for teams that value general capability, moderate cost, and usable streaming speed over a clear specialty. Its Intelligence Index score of 32 matches DeepSeek V3.2 (Reasoning), while Qwen3.5 122B A10B (Reasoning) scores 32.3 and GLM-5 (Non-reasoning) scores 32.4. These neighboring results suggest that the model sits in a competitive cluster rather than separating itself through a decisive benchmark advantage. The ranking of 126 of 578 supports that interpretation: Qwen3.5 397B A17B (Non-reasoning) is materially above the broad field, but it is not near the top of the listed comparison set.

For a developer, that position creates a clear trade-off. DeepSeek V3.2 (Reasoning) has the same Intelligence Index score and reported coding and math results, so it may be a stronger fit when explicit reasoning or technical depth is central. Qwen3.5 122B A10B (Reasoning) is slightly higher on the Intelligence Index and reports a coding score of 45.7, while also showing an output speed of 138.285 tokens per second. Qwen3.5 397B A17B (Non-reasoning) remains attractive when a direct-response mode and its own latency-throughput balance fit the application better. The supplied evidence does not show whether those differences persist on your prompts.

Model Selection signal
Qwen3.5 397B A17B (Non-reasoning) Balanced general benchmark position and direct-response profile
DeepSeek V3.2 (Reasoning) Same Intelligence Index score, with reported coding and math metrics
Qwen3.5 122B A10B (Reasoning) Slightly higher Intelligence Index score and higher reported output speed

Data provided by Artificial Analysis.

03

The ranking supports broad usefulness, but it cannot predict task-level reliability

Qwen3.5 397B A17B (Non-reasoning) is strong enough on the available ranking to merit broad task testing, but the ranking alone cannot identify where it will fail. A position of 126 of 578 indicates a model with meaningful general capability relative to the measured population. It does not tell a developer whether the model follows schemas, preserves instructions across long prompts, writes correct code, handles adversarial inputs, or produces stable answers across repeated requests. Those are separate production properties, and the research brief contains no verified public evidence for them.

The most important evidence gap is coding. Developers often use a general model for repository questions, patch generation, test writing, and debugging. Yet the supplied record does not provide a coding index for Qwen3.5 397B A17B (Non-reasoning), and the research brief found no reliable community reports that could fill that gap. The presence of coding scores for neighboring models does not allow a direct conclusion about this model. Qwen3.5 122B A10B (Reasoning) reports a coding score of 45.7, and Qwen3.6 35B A3B (Reasoning) reports 41.9, but those values describe different models and modes.

The same caution applies to reasoning-heavy workflows. The model is explicitly labeled non-reasoning, while the closest alternatives include reasoning variants. That naming difference may matter for planning, mathematical derivation, and complex debugging, but the available material does not document the behavioral boundary. Test the model with representative prompts, strict output validation, tool-call scenarios, and retry policies before assigning it a critical role.

04

Qwen3.5 397B A17B (Non-reasoning) has usable responsiveness for interactive generation

Qwen3.5 397B A17B (Non-reasoning) should feel responsive in interactive streaming workloads because its measured latency is 0.3 seconds and its median output speed is 66.396 tokens per second. Those figures support applications such as chat interfaces, drafting tools, classification with explanations, and developer assistants where users see output progressively. They do not guarantee equal performance under concurrency, large prompts, rate limits, cold starts, or a particular hosting provider. The supplied research brief found no verified availability or infrastructure documentation, so operational behavior remains unconfirmed.

The speed result is best interpreted as a workload filter. It makes the model plausible for user-facing flows that need an answer to begin quickly. It does not automatically make the model suitable for high-volume batch generation. Batch economics depend on completion length, retry rates, queueing, and output quality. A slower model with fewer corrections can be cheaper in practice, while a faster model that needs human review can be expensive. The supplied data cannot resolve that trade-off.

Compared with the listed neighboring models, Qwen3.5 122B A10B (Reasoning) reports 138.285 output tokens per second, but that comparison does not establish equivalent quality, serving conditions, or response behavior. The practical question is whether Qwen3.5 397B A17B (Non-reasoning) reaches acceptable task quality before its speed advantage or disadvantage affects user experience. Measure time to accepted result, not only time to first token.

05

The price is reasonable for a general model, but output-heavy workloads can change the verdict

Qwen3.5 397B A17B (Non-reasoning) is most defensible at its $1.35 per 1M blended-token price when the model reduces routing complexity and delivers acceptable first-pass quality. The listed input price is $0.6 per 1M tokens, while the output price is $3.6 per 1M tokens. That spread means applications that generate long answers, code patches, or repeated explanations need to watch output volume closely. The blended figure is useful for rough comparison, but it can hide the cost profile of a workload with unusually long completions.

The adjacent models show why price alone is not enough. DeepSeek V3.2 (Reasoning) is listed at $0.315 per 1M blended tokens, and Qwen3.5 122B A10B (Reasoning) is listed at $1.1. Qwen3 Max Thinking is listed at $15, despite an Intelligence Index score of 31.7. These comparisons suggest that Qwen3.5 397B A17B (Non-reasoning) is neither the cheapest option in the set nor priced like a premium specialist. Its value depends on whether its direct-response behavior, measured responsiveness, and general benchmark position reduce downstream review or fallback costs.

Do not select it solely because the blended price appears moderate. Track accepted-output rate, average completion length, retry frequency, and escalation frequency during a pilot. The research brief found no verified current pricing page or stable API information, so the listed price should be rechecked before procurement or capacity planning. The supplied evidence also does not show whether the price is available consistently across providers.

06

Choose Qwen3.5 397B A17B (Non-reasoning) for balanced workloads, not unverified critical decisions

Qwen3.5 397B A17B (Non-reasoning) is worth piloting for general developer workflows that can tolerate evidence gaps and validate outputs before release. Its position of 126 of 578, score of 32, 0.3-second latency, and 66.396 output tokens per second form a coherent case for testing interactive assistants, content transformation, support drafting, and internal productivity tools. These signals are strong enough to justify engineering time, but they are not enough to approve the model for unsupervised code changes, financial decisions, security-sensitive analysis, or other high-consequence tasks.

Use the model when the application benefits from direct answers and moderate blended-token economics. Keep a reasoning-oriented fallback when tasks require explicit multi-step planning, difficult mathematics, or repository changes that must compile and pass tests. Consider a lower-cost adjacent model when the workload is high-volume and quality differences are small. Consider a faster adjacent model when user-perceived completion time matters more than model size or mode. These recommendations are conditional because the supplied research brief contains no verified behavioral studies, provider documentation, or failure reports.

Decision Recommendation
General assistant pilot Recommended, with representative prompt tests
Interactive drafting Plausible, given the reported latency and output speed
Complex reasoning Validate against a reasoning model before adoption
Production coding agent Require compilation, tests, and human review
Procurement commitment Reconfirm provider, availability, and price first

The final choice should be based on accepted task outcomes, not the benchmark rank alone. Data provided by Artificial Analysis.

07

Questions developers should answer before adoption

Qwen3.5 397B A17B (Non-reasoning) requires a focused validation plan because the available research confirms benchmark and serving metrics but leaves product behavior largely undocumented.

The supplied research brief found no reliable official product page, developer documentation, stable API alias, current availability statement, pricing page, community post, limitation notice, or public failure case. That absence should shape the pilot design. It should not be turned into an unsupported claim that the model is unreliable. Developers should explicitly test the missing dimensions and record the provider, endpoint, model identifier, prompt format, output format, and observed failure rate.

A useful evaluation should include short and long prompts, structured JSON responses, code explanation, patch generation, multilingual input if relevant, tool use if supported, repeated sampling, and refusal or uncertainty handling. Compare accepted outputs with at least one adjacent model. The goal is to learn whether Qwen3.5 397B A17B (Non-reasoning) offers a better total workflow, not merely a better headline metric.

Frequently asked questions

Is Qwen3.5 397B A17B (Non-reasoning) worth evaluating?

Yes, Qwen3.5 397B A17B (Non-reasoning) is worth evaluating because its Intelligence Index score of 32 ranks 126 of 578, while its reported latency and output speed support interactive testing. However, no reliable public product or behavior documentation was found, so evaluation should precede production approval.

Is Qwen3.5 397B A17B (Non-reasoning) a good coding model?

The available evidence cannot establish that Qwen3.5 397B A17B (Non-reasoning) is a good coding model because its supplied record has no coding index and the research brief found no reliable coding reports. Test repository tasks, patch correctness, test generation, and instruction adherence directly.

Is Qwen3.5 397B A17B (Non-reasoning) fast enough for an interactive application?

Qwen3.5 397B A17B (Non-reasoning) appears suitable for interactive testing because the reported latency is 0.3 seconds and median output speed is 66.396 tokens per second. Actual user experience still depends on concurrency, prompt size, provider infrastructure, rate limits, and completion length.

Is Qwen3.5 397B A17B (Non-reasoning) cost-effective?

Qwen3.5 397B A17B (Non-reasoning) can be cost-effective for balanced workloads at $1.35 per 1M blended tokens, but output-heavy applications pay closer attention to the $3.6 output-token price. Compare accepted results, retries, review time, and fallback usage before choosing it.

When should developers avoid Qwen3.5 397B A17B (Non-reasoning)?

Developers should avoid assigning Qwen3.5 397B A17B (Non-reasoning) an unsupervised critical role until its provider, availability, coding behavior, reasoning limits, and failure patterns are verified. The research brief supplies no reliable evidence for those operational and behavioral requirements.

Sources

  1. Artificial AnalysisBenchmark position, Intelligence Index score, pricing, latency, output speed, and comparison data supplied in the data brief

Published: