Skip to content

Hy3-preview (Reasoning) vs o3: The Ultimate Performance & Pricing Comparison

Deep dive into reasoning, benchmarks, and latency insights.

The Final Verdict in the Hy3-preview (Reasoning) vs o3 Showdown

The current catalog does not contain complete performance evidence for both models, so this page does not declare an overall winner. Use the available fields as comparison signals and validate the models on your own workload.

Model Snapshot

Key decision metrics at a glance.

Hy3-preview (Reasoning)o3
6.0
Reasoning
9.0
6.0
Coding
6.0
3.0
Multimodal
3.0
4.0
Long Context
4.0
$0.108
Blended Price / 1M tokens
$3.5
P95 Latency
Tokens per second
128.056

Machine-readable comparison data

ModelMetricValueUnitSource / snapshot
Hy3-preview (Reasoning)Reasoning6.0benchmark or capability scoreArtificial Analysis · current catalog
o3Reasoning9.0benchmark or capability scoreArtificial Analysis · current catalog
Hy3-preview (Reasoning)Coding6.0benchmark or capability scoreArtificial Analysis · current catalog
o3Coding6.0benchmark or capability scoreArtificial Analysis · current catalog
Hy3-preview (Reasoning)Multimodal3.0benchmark or capability scoreArtificial Analysis · current catalog
o3Multimodal3.0benchmark or capability scoreArtificial Analysis · current catalog
Hy3-preview (Reasoning)Long Context4.0benchmark or capability scoreArtificial Analysis · current catalog
o3Long Context4.0benchmark or capability scoreArtificial Analysis · current catalog
Hy3-preview (Reasoning)Blended Price / 1M tokens$0.108USD per 1M tokensArtificial Analysis · current catalog
o3Blended Price / 1M tokens$3.5USD per 1M tokensArtificial Analysis · current catalog
Hy3-preview (Reasoning)P95 LatencymillisecondsArtificial Analysis · current catalog
o3P95 LatencymillisecondsArtificial Analysis · current catalog
Hy3-preview (Reasoning)Tokens per secondtokens per secondArtificial Analysis · current catalog
o3Tokens per second128.056tokens per secondArtificial Analysis · current catalog

Data provided by Artificial Analysis; live values use the current catalog.

Overall Capabilities

This radar chart visually maps the core capabilities (reasoning, coding, math proxy, multimodal, long context) of `Hy3-preview (Reasoning)` vs `o3`.

IntelligenceCodingMathMultimodalLong Context
Hy3-preview (Reasoning)o3

Benchmark Breakdown

This grouped bar chart provides a side-by-side comparison for each benchmark metric.

Hy3-preview (Reasoning)o3

Speed & Latency

Lower time to first token is better; higher tokens per second is better.

Time to First Token · Hy3-preview (Reasoning)
Time to First Token · o3
Tokens per Second · Hy3-preview (Reasoning)
Tokens per Second · o3
128.056
Head to the playground to validate these results yourself

The Economics of Hy3-preview (Reasoning) vs o3

Pricing Breakdown

Compare input and output pricing in USD per 1M tokens.

Hy3-preview (Reasoning)o3

Real-World Cost Scenario

Per run: 1M input tokens + 250k output tokens

Hy3-preview (Reasoning)$0.124

o3$4

Hy3-preview (Reasoning) costs $3.876 less per run

Review the complete pricing and packaging strategy

Hy3-preview (Reasoning) vs o3: Which Model Should Developers Choose?

This article is a dated snapshot published on 2026-08-07. Live cards above use the current catalog; missing live fields are not inferred.

Hy3-preview (Reasoning) vs o3: Which Model Should Developers Choose?
  • Winner overall: Hy3-preview (Reasoning), with an Artificial Analysis Intelligence Index of 33.6 vs o3 at 30.4, although operational evidence is incomplete
  • Cheaper: Hy3-preview (Reasoning) at $0.10750000000000001 vs $3.5 per 1M blended tokens
  • Faster: o3 at 128.056 median output tokens per second
  • Pick o3 when: You need a documented OpenAI model with a measured Artificial Analysis Math Index of 88.3
  • Watch out: Neither model has a confirmed context window in the supplied evidence, and current o3 availability and pricing remain unclear

Hy3-preview (Reasoning) vs o3 at a glance

Hy3-preview (Reasoning) leads the measured general intelligence score, while o3 offers the only supplied output-speed and mathematics measurements. The Artificial Analysis Intelligence Index places Hy3-preview (Reasoning) at 33.6 and o3 at 30.4. The supplied data also records identical latency of 0.3 seconds for both models. o3 reaches 128.056 median output tokens per second, while Hy3-preview (Reasoning) has no supplied output-speed value. Hy3-preview (Reasoning) costs $0.10750000000000001 per 1M blended tokens, compared with $3.5 for o3. Those measurements favor Hy3-preview (Reasoning) for low-cost experimentation, but they do not establish production readiness. The evidence contains no verified vendor documentation for Hy3-preview (Reasoning), and the supplied OpenAI model directory does not list o3 among its current models.

Executive summary for developers

Hy3-preview (Reasoning) is the stronger measured value choice, while o3 is the safer choice only where its documented ecosystem and mathematics result matter more than price. The score gap is visible but limited: Hy3-preview (Reasoning) records 33.6 on the Artificial Analysis Intelligence Index, and o3 records 30.4. That difference does not prove that Hy3-preview (Reasoning) will write better code, follow tool instructions more reliably, or solve a specific application’s tasks more accurately. No supplied source describes Hy3-preview (Reasoning)’s API, context window, output limit, multimodal support, stable alias, or failure modes. The supplied OpenAI model directory also does not provide those details for o3, because o3 is not listed there.

The practical comparison therefore has two layers. The benchmark layer favors Hy3-preview (Reasoning) on the available general index and gives o3 a mathematics score of 88.3. The operational layer is unresolved for both models, with especially high uncertainty around Hy3-preview (Reasoning). Developers should treat the Artificial Analysis measurements as useful screening signals, not as a complete deployment decision. The supplied data identifies Hy3-preview (Reasoning) as released on 2026-04-23 and o3 as released on 2025-04-16, but release dates do not establish current support or compatibility.

The strongest conclusion is conditional: Hy3-preview (Reasoning) deserves a controlled evaluation for cost-sensitive workloads, while o3 deserves consideration for teams that specifically value its measured mathematics result and known OpenAI association. Neither model can be selected confidently for context-heavy or API-sensitive work from the supplied evidence alone.

Performance: benchmark separation does not define developer experience

o3 is the better-supported performance candidate in the supplied data because it has an output-speed measurement and a mathematics score, while Hy3-preview (Reasoning) has fewer measured dimensions. o3 records 128.056 median output tokens per second, which gives developers a concrete signal for long generated responses, interactive coding sessions, and agent loops. Hy3-preview (Reasoning) has no supplied output-speed value, so its practical streaming behavior cannot be compared fairly. The equal latency value of 0.3 seconds suggests that initial response timing does not separate the models in this dataset. It does not show that total task completion time will be equal, because generation speed remains unknown for Hy3-preview (Reasoning).

Hy3-preview (Reasoning) leads the available Artificial Analysis Intelligence Index at 33.6 versus o3 at 30.4. That result may make it attractive for broad reasoning evaluations, but the supplied evidence does not identify the index’s task mix in enough detail to map the difference to code review, debugging, planning, or tool use. o3’s Artificial Analysis Math Index is 88.3, while no corresponding Hy3-preview (Reasoning) value is supplied. The missing value is evidence of an incomplete comparison, not proof that Hy3-preview (Reasoning) is weaker at mathematics.

The performance decision should therefore depend on the workload. Use o3 as the easier candidate to test for speed-sensitive generation and mathematical reasoning. Test Hy3-preview (Reasoning) directly before assigning it latency-sensitive production work. Developers should also measure complete task time, retries, tool-call correctness, and answer acceptance, because the supplied benchmark and latency fields do not cover those behaviors.

Hy3-preview (Reasoning)o3
33.6
ARTIFICIAL ANALYSIS INTELLIGENCE
30.4
ARTIFICIAL ANALYSIS MATH
88.3
Performance: benchmark separation does not define developer experience · Data provided by Artificial Analysis; live values use the current catalog.

Cost: Hy3-preview changes the economics, but workload shape still matters

Hy3-preview (Reasoning) is the clear price leader in the supplied data, but its low token cost does not remove the need to validate reliability and availability. Hy3-preview (Reasoning) is listed at $0.10750000000000001 per 1M blended tokens using the 3-to-1 input-to-output mix, while o3 is listed at $3.5. The input prices are $0.065 for Hy3-preview (Reasoning) and $2 for o3 per 1M input tokens. The output prices are $0.235 and $8 respectively. These figures make Hy3-preview (Reasoning) attractive for high-volume classification, drafting, evaluation, and exploratory agent workloads, assuming the model can be called consistently.

The blended figure is a modeling assumption, not a universal application cost. A product that produces unusually long answers will be more exposed to output pricing. A product that sends large prompts will be more exposed to input pricing. The supplied data does not provide request volumes, token distributions, retry rates, or task-success rates, so it cannot identify the true cost per successful task. A cheaper model can become more expensive operationally if it needs more retries, produces unusable code, or requires additional verification. Those failure rates are not available for either model.

Current o3 billing status is also uncertain. The supplied OpenAI API pricing page does not list o3 under Standard, Batch, Flex, or Fast mode pricing. The listed $3.5 blended value is therefore a data-brief measurement, not a confirmed current OpenAI price. Developers should confirm endpoint access and billing before committing architecture. Hy3-preview (Reasoning) has an even larger evidence gap because no verified pricing page or vendor documentation is supplied.

Hy3-preview (Reasoning)o3
$0.065
Input Pricing
$2
$0.235
Output Pricing
$8
$0.108
Blended Price / 1M tokens
$3.5

Hy3-preview (Reasoning) leads on 3 of 3 metrics

Cost: Hy3-preview changes the economics, but workload shape still matters · Data provided by Artificial Analysis; live values use the current catalog.

Recommendation: choose by evidence tolerance, not benchmark rank alone

Hy3-preview (Reasoning) is the best first candidate for cost-sensitive evaluation, while o3 is the better fallback when a mathematics signal and recognizable vendor context are priorities. Start with Hy3-preview (Reasoning) when the application can tolerate an unknown API contract, an unconfirmed context window, and missing production evidence. Its 33.6 Artificial Analysis Intelligence Index and $0.10750000000000001 blended price justify a controlled proof of concept. The proof of concept should test the application’s real prompts, structured outputs, coding tasks, tool calls, and recovery behavior.

Choose o3 when the application benefits from its 88.3 Artificial Analysis Math Index or when the team prefers evaluating an OpenAI-associated model. o3 also has the only supplied output-speed measurement, 128.056 median output tokens per second. Those advantages matter for mathematical workloads and interactive generation, but the supplied OpenAI model directory does not currently list o3. The supplied OpenAI API pricing page also does not list an o3 price. Availability, endpoint naming, and current billing must therefore be confirmed before implementation.

Do not make either model the default for context-heavy systems based on this brief. Both models have a null context-window value in the supplied dataset. Hy3-preview (Reasoning) lacks verified official or community sources altogether. o3 has official directory and pricing references, but those references establish absence from the current pages rather than a complete migration path. The selection should remain provisional until direct API tests answer the unanswered operational questions.

Evidence gaps that can reverse the decision

o3 has more identifiable evidence, but neither model has enough supplied documentation to support a fully confident production recommendation. The source material provides no verified Hy3-preview (Reasoning) announcement, developer documentation, pricing page, community discussion, or failure analysis. That absence prevents reliable conclusions about its API stability, output controls, context handling, coding behavior, and support lifecycle. The data brief still records a release date of 2026-04-23 and measurable pricing, but those fields do not replace a callable public interface or a documented contract.

The o3 evidence is narrower than a normal production assessment. The supplied OpenAI model directory does not list o3 and does not confirm a stable alias, endpoint, successor, context window, output limit, API parameters, or multimodal capability. The supplied OpenAI API pricing page does not list current o3 Standard, Batch, Flex, or Fast mode prices. No verified community test is available for either model, so coding feel, speed perception, instruction-following preferences, and failure scenarios remain unknown.

These gaps create a clear reversal risk. Hy3-preview (Reasoning) may win the measured value comparison but fail an access or reliability test. o3 may be the stronger mathematical or interactive option but become unsuitable if current access or pricing cannot be confirmed. Developers should report those unknowns explicitly rather than converting missing evidence into positive or negative capability claims.

Sources

  1. OpenAI ModelsVerifying the current model directory, o3 visibility, model availability signals, and documented capability gaps.
  2. OpenAI API PricingVerifying the current pricing page and the absence of listed o3 Standard, Batch, Flex, and Fast mode prices.

Your Questions about the Hy3-preview (Reasoning) vs o3 Comparison

Which model should developers test first?

Developers should test Hy3-preview (Reasoning) first when evaluation budget matters, because its supplied blended price is $0.10750000000000001 and its Intelligence Index is 33.6. The test must verify access, reliability, coding quality, and output behavior because no official Hy3-preview (Reasoning) documentation is supplied.

Is Hy3-preview (Reasoning) faster than o3?

The supplied evidence cannot establish that Hy3-preview (Reasoning) is faster than o3. Both models show 0.3 seconds of latency, but only o3 has a median output-speed measurement of 128.056 tokens per second, while Hy3-preview (Reasoning) has no supplied value.

Why choose o3 despite its higher listed cost?

Developers may choose o3 when its measured Artificial Analysis Math Index of 88.3 fits the workload, or when OpenAI-associated model context is important. The supplied evidence does not confirm current o3 availability or pricing, so teams must verify those details before committing.

Can this comparison confirm production readiness?

This comparison cannot confirm production readiness for either model. Both models have null context-window values in the supplied data, and the research brief lacks verified failure rates, tool-use behavior, support guarantees, retry data, and complete API contracts.

Does Hy3-preview (Reasoning) win the benchmark comparison?

Hy3-preview (Reasoning) leads the supplied Artificial Analysis Intelligence Index at 33.6 versus o3 at 30.4. The result does not establish a universal win because o3 has a supplied mathematics score of 88.3 and Hy3-preview (Reasoning) lacks several comparable measurements.