Skip to content

Hy3-preview (Reasoning)

Available

Other · 2026-04-23 · 32,000 tokens

An AI model from Other, suited to a broad range of AI workloads.

Supported modalities:textcode

Quick Overview

Text Generation3/10
Code Generation6/10
Reasoning6/10
Multimodal3/10

Benchmark Results

Scores from leading benchmark suites.

artificial analysis intelligence34.4

Performance Metrics

Latency and throughput performance.

P50 Latency
0tokens/sec

Dive Deeper

AI model analysis

Hy3-preview (Reasoning) Review: Extremely Cheap, but Evidence Is Thin

Hy3-preview (Reasoning) Review: Extremely Cheap, but Evidence Is Thin
Summary

- **Where it stands:** Hy3-preview (Reasoning) ranks 115 of 578 on the Artificial Analysis Intelligence Index at 33.6 - **Price:** $0.1075 per 1M blended tokens - **Speed:** output tokens per second not reported, 0.3s to first token - **Pick it when:** you need a low-cost reasoning model for experiments, prototypes, or workloads with human review - **Watch out:** public evidence does not confirm availability, API behavior, context limits, or failure patterns

01

Hy3-preview (Reasoning) is a low-cost model with a respectable general benchmark position and unusually limited public documentation.

Hy3-preview (Reasoning) is difficult to evaluate as a production dependency because the available research brief contains no verified vendor announcement, developer documentation, pricing page, or community discussion. That absence limits what can be said about its interface, reliability, context handling, output limits, multimodal support, and practical coding behavior.

The available data still gives developers one useful signal. Hy3-preview (Reasoning) scores 33.6 on the Artificial Analysis Intelligence Index and ranks 115 of 578 models. That places the model in a meaningful middle-to-upper portion of the measured field, although the ranking alone does not establish that it is reliable for a particular application.

The strongest concrete argument for testing Hy3-preview (Reasoning) is cost. Its blended price is $0.1075 per 1M tokens, while its reported time to first token is 0.3 seconds. Output throughput is not reported. Developers should therefore treat this model as an inexpensive candidate for controlled evaluation, not as a proven replacement for a documented production model.

Data provided by https://artificialanalysis.ai/ supplies the benchmark, ranking, latency, and pricing snapshot used in this review.

02

Hy3-preview (Reasoning) looks most attractive when low inference cost matters more than operational certainty.

Hy3-preview (Reasoning) offers a rare combination of a 33.6 general intelligence score and a $0.1075 blended token price, but the evidence does not show whether that advantage survives real application workloads.

The closest measured models provide a useful reference point. Claude 4.1 Opus (Reasoning), GLM-4.7 (Reasoning), and GPT-5 (medium) each score 33.7 on the Artificial Analysis Intelligence Index. GPT-5.5 Instant (May 2026) scores 33.5. KAT Coder Pro V2 also scores 33.7, while its coding index is 59.5. These nearby scores suggest that Hy3-preview (Reasoning) is broadly competitive on this single aggregate measure, but they do not identify a clear capability advantage.

The trade-off becomes clearer when cost is included:

Model Main signal Developer implication
Hy3-preview (Reasoning) 33.6 intelligence score, $0.1075 blended price Very inexpensive experimentation, with substantial unknowns
KAT Coder Pro V2 33.7 intelligence score, 59.5 coding score, $0.525 blended price More visible coding evidence, at a higher price
GLM-4.7 (Reasoning) 33.7 intelligence score, 45.3 coding score, $1 blended price Better documented benchmark breadth in the supplied data
GPT-5 (medium) 33.7 intelligence score, $3.4375 blended price Higher cost, with stronger math evidence in the supplied data

Hy3-preview (Reasoning) is therefore a cost-first option. The brief does not provide enough evidence to claim that it is better for coding, mathematics, long-context work, or production stability.

03

Hy3-preview (Reasoning) is benchmark-competitive in aggregate, but the available data cannot predict task-level quality.

Hy3-preview (Reasoning) reaches 33.6 on the Artificial Analysis Intelligence Index and ranks 115 of 578, which supports testing for general reasoning tasks but does not prove consistent performance in software development.

The aggregate position is close to several supplied reference models. Claude 4.1 Opus (Reasoning), GLM-4.7 (Reasoning), GPT-5 (medium), and KAT Coder Pro V2 each have an intelligence score of 33.7. GPT-5.5 Instant (May 2026) scores 33.5. Hy3-preview (Reasoning) is therefore near its comparison group on the reported general index, rather than being an obvious outlier.

The missing task-level evidence matters. The data brief does not report a coding score, math score, tool-use result, instruction-following result, context-window measurement, or output limit for Hy3-preview (Reasoning). It also does not report median output tokens per second. Developers cannot infer coding quality from the intelligence rank alone, especially for repository changes, debugging, structured output, or agentic loops.

The reported 0.3-second latency is useful for interactive testing, but latency does not reveal completion speed or total task duration. A model can start quickly and still be inefficient on long reasoning tasks. The research brief also found no verified community reports describing coding experience, speed perception, model preferences, failure modes, or usage traps.

The practical conclusion is narrow: Hy3-preview (Reasoning) deserves a task-based trial because its general benchmark position is credible enough to justify one. The evidence is insufficient to predict whether it will produce correct patches, preserve existing behavior, follow strict schemas, or recover from tool errors.

04

Hy3-preview (Reasoning) is cheap enough to justify broad evaluation, but its low price does not automatically make it economical.

Hy3-preview (Reasoning) costs $0.1075 per 1M blended tokens, making it the clear cost leader among the supplied nearby models.

That price can change the economics of experimentation. Teams can afford more candidate prompts, regression cases, and human-reviewed trials before committing to a deployment decision. It may also suit batch classification, draft generation, lightweight code explanation, and internal prototypes where occasional failure is cheap to detect.

The comparison is substantial. KAT Coder Pro V2 costs $0.525 per 1M blended tokens. GLM-4.7 (Reasoning) costs $1. Claude 4.1 Opus (Reasoning) costs $30, GPT-5.5 Instant (May 2026) costs $11.25, and GPT-5 (medium) costs $3.4375. These figures show why Hy3-preview (Reasoning) can be attractive for volume-sensitive workloads.

The price advantage becomes less meaningful if the model requires repeated retries, extensive validation, manual correction, or routing to a stronger model. The supplied data does not include error rates, answer quality by task, retry behavior, or total cost per successful result. Developers should measure cost per accepted output, not token price alone.

The input and output prices also differ. Hy3-preview (Reasoning) is listed at $0.065 per 1M input tokens and $0.235 per 1M output tokens. Output-heavy reasoning workflows may therefore experience a different cost profile from short-answer workloads. The supplied blended figure is the most useful headline estimate, but real spending will depend on the token mix and the model’s tendency to produce long answers.

No verified current pricing page was found in the research brief. The listed price should be treated as a data snapshot, not confirmation of a currently callable public service.

05

Hy3-preview (Reasoning) is worth piloting for low-risk, cost-sensitive workloads, but not choosing as an unverified sole production model.

Hy3-preview (Reasoning) is a sensible pilot candidate when token cost is the main constraint and a human or deterministic validation layer can catch mistakes.

Good fits include prompt experiments, offline evaluation, low-risk content transformation, internal assistants, and workloads where the system can retry or route difficult cases elsewhere. Its 33.6 intelligence score gives it enough benchmark standing to justify testing. Its $0.1075 blended price makes that testing inexpensive relative to the supplied comparison models.

Developers should avoid making it the only model behind high-consequence decisions, unattended code changes, or workflows that require a known context window and stable API contract. The research brief found no verified official documentation, stable alias, replacement path, current availability, parameter reference, or known limitation. Those gaps are operational risks, independent of benchmark performance.

A practical pilot should focus on accepted-task rate, correction rate, retry count, latency under realistic prompts, output length, structured-output validity, and failure recovery. Those measurements are not available in the supplied materials, so they must come from the developer’s own workload. The comparison should include at least one nearby model with task-specific evidence, such as KAT Coder Pro V2 for coding-oriented tests or GPT-5 (medium) for math-oriented tests.

The recommendation is conditional: test Hy3-preview (Reasoning) as a cheap specialist or fallback candidate, then keep it only if its cost per accepted result beats the operational burden. Current evidence does not support a stronger production endorsement.

06

Questions developers should answer before adopting Hy3-preview (Reasoning)

Hy3-preview (Reasoning) should be adopted only after a small workload-specific evaluation confirms access, output quality, and operational fit.

The research brief contains no verified official or community source for product behavior, so the questions below separate measured facts from unresolved adoption risks.

Frequently asked questions

Is Hy3-preview (Reasoning) good enough for production use?

Hy3-preview (Reasoning) is not supported by enough public evidence for an unconditional production recommendation, despite ranking 115 of 578 on the Artificial Analysis Intelligence Index. Developers should first verify availability, API stability, task accuracy, structured-output reliability, and recovery behavior on their own workload.

Why consider Hy3-preview (Reasoning) instead of a better-known model?

Hy3-preview (Reasoning) is worth considering because its $0.1075 blended price is far below the supplied nearby models, while its intelligence score of 33.6 remains close to several models scoring 33.7. The trade-off is missing documentation and limited task-specific evidence.

Is Hy3-preview (Reasoning) suitable for coding?

Hy3-preview (Reasoning) may be suitable for coding experiments, but the supplied data does not include a coding score or verified developer reports. Teams should compare patch correctness, test-pass rate, repository navigation, and repair behavior against a coding-focused reference before deployment.

How fast is Hy3-preview (Reasoning)?

Hy3-preview (Reasoning) has a reported latency of 0.3 seconds to first token, but output tokens per second are not reported. That means the model may feel responsive at the start while its total completion time remains unknown for long reasoning or code-generation tasks.

What are the main risks of choosing Hy3-preview (Reasoning)?

Hy3-preview (Reasoning) has an evidence risk rather than a confirmed technical failure pattern. The research brief could not verify its context window, output limit, API parameters, multimodal support, current availability, pricing page, community behavior, or known failure scenarios.

Sources

  1. Artificial AnalysisBenchmark score, ranking, latency, pricing snapshot, and comparison-model data

Published: