Skip to content

Granite 4.0 H 350M

Available

Other · 2025-10-28 · 32,000 tokens

An AI model from Other, suited to a broad range of AI workloads.

Supported modalities:textcode

Quick Overview

Text Generation1/10
Code Generation6/10
Reasoning1/10
Multimodal1/10

Benchmark Results

Scores from leading benchmark suites.

artificial analysis intelligence1.0
artificial analysis math1.3

Performance Metrics

Latency and throughput performance.

P50 Latency
0tokens/sec

Dive Deeper

AI model analysis

Granite 4.0 H 350M Review: A Free Model with Limited Evidence

Granite 4.0 H 350M Review: A Free Model with Limited Evidence
Summary

- **Where it stands:** Granite 4.0 H 350M ranks 582 of 595 on the Artificial Analysis Intelligence Index at 1 - **Price:** $0 per 1M blended tokens - **Speed:** 0 output tokens per second, 0s to first token - **Pick it when:** You need a no-cost, low-stakes experiment and can validate every output yourself - **Watch out:** No reliable official documentation or community testing was found, so availability, limits, and failure modes remain uncertain

01

Granite 4.0 H 350M is a free recorded model with serious evidence gaps

Granite 4.0 H 350M looks useful mainly as a low-cost experiment, not as a dependable production model.

The available data records a blended price of $0 per 1M tokens, but that figure does not prove that the model has a stable public API or an active hosted endpoint. The research brief found no verifiable IBM announcement, developer documentation, pricing page, or reliable community discussion for Granite 4.0 H 350M. That means important deployment questions remain unanswered, including the context window, maximum output length, supported API parameters, multimodal capability, current availability, and model naming.

The benchmark evidence is much clearer than the product evidence. Granite 4.0 H 350M ranks 582 of 595 on the Artificial Analysis Intelligence Index. It ranks 346 of 348 on MMLU-Pro and 559 of 567 on SciCode. Those positions place the model close to the bottom of broad knowledge and scientific coding evaluations.

The model is therefore hard to justify for tasks where incorrect answers create support, financial, operational, or engineering risk. A developer can still test it for simple classification, routing, local experimentation, or disposable prototypes. However, the zero price should be treated as a recorded data point rather than a complete commercial recommendation. Data provided by Artificial Analysis.

02

Granite 4.0 H 350M offers cost relief, but nearby small models show a stronger quality case

Granite 4.0 H 350M is attractive only when avoiding model fees matters more than answer quality and operational certainty.

The closest models in the supplied comparison set provide a useful reference point. Gemma 3 1B Instruct and Gemma 3 4B Instruct are also recorded at $0 per 1M blended tokens. Yet Gemma 3 4B Instruct has a MMLU-Pro score of 0.417, while Granite 4.0 H 350M has a score of 0.127. Gemma 3 4B Instruct also records 0.112 on LiveCodeBench, compared with Granite 4.0 H 350M at 0.019. These are not proof that Gemma will win every task, but they make Granite harder to select when both options appear free in the snapshot.

Gemma 3 1B Instruct is a closer size-oriented reference. It records 0.135 on MMLU-Pro, slightly above Granite 4.0 H 350M’s 0.127, and 0.017 on SciCode, equal to Granite 4.0 H 350M’s 0.017. This suggests that Granite’s small footprint does not create a clear benchmark advantage in the supplied data.

Apertus 8B Instruct is not free in the snapshot, with a blended price of $0.125 per 1M tokens. Its GPQA score is 0.256, almost the same as Granite 4.0 H 350M’s 0.257, but Granite’s lower model size and zero recorded price may still matter for tightly constrained experiments. The comparison does not establish hosting quality, speed, context capacity, or licensing terms for any model.

The practical summary is simple: Granite may reduce recorded inference cost, but the supplied evidence does not show that it delivers better value than other free small models. Data provided by Artificial Analysis.

03

Granite 4.0 H 350M is weak on broad capability and software-oriented evaluations

Granite 4.0 H 350M should be expected to struggle with complex reasoning, coding, and precise instruction-following workloads.

Its overall position is the strongest warning sign. Granite 4.0 H 350M ranks 582 of 595 on the Artificial Analysis Intelligence Index, placing it near the bottom of the measured field. The ranking does not explain why individual answers fail, and the research brief contains no verified failure analysis. Still, it is strong evidence against using the model as a general-purpose assistant without a strict validation layer.

The knowledge and reasoning results point in the same direction. Granite 4.0 H 350M ranks 346 of 348 on MMLU-Pro, with a score of 0.127. It ranks 558 of 575 on GPQA, with a score of 0.257. It ranks 312 of 567 on HLE, with a score of 0.064. The HLE position is less extreme than the MMLU-Pro and GPQA positions, so the model is not uniformly last across every evaluation. Even so, the combined profile does not support high-confidence answers to difficult factual or analytical questions.

Coding results are especially weak. Granite 4.0 H 350M ranks 339 of 343 on LiveCodeBench at 0.019 and 559 of 567 on SciCode at 0.017. It also ranks 385 of 432 on TerminalBench Hard with a score of 0. The supplied research brief found no reliable developer reports that explain whether these results reflect poor code generation, weak tool use, inadequate reasoning, or evaluation-specific behavior. Developers should therefore avoid assuming that a simple prompt or a different wrapper will remove the problem.

Instruction following is also a concern. Granite 4.0 H 350M ranks 446 of 450 on IFBench with a score of 0.175510204081633. It ranks 437 of 500 on LCR at 0 and 386 of 440 on tau2 at 0.146198830409357. These results suggest limited reliability when a task requires constraints, long chains of decisions, or consistent interaction behavior. They do not establish the exact failure mode.

The performance conclusion can change for narrow tasks. A model with weak general benchmarks may still be adequate for deterministic labels, short text cleanup, simple extraction, or a fallback path. The available materials do not include task-specific tests, latency observations from a live endpoint, or prompt-level examples. A developer should run a private acceptance set before trusting Granite with any user-facing workflow. Data provided by Artificial Analysis.

04

Granite 4.0 H 350M is cost-effective only if the free access is real and the quality threshold is low

Granite 4.0 H 350M has a recorded blended price of $0 per 1M tokens, but free inference does not automatically make it economical.

The zero price matters most for experiments that would otherwise incur meaningful request volume. It can support quick prototypes, offline tests, internal demos, or a fallback route where incorrect answers are cheap to detect. It may also be useful when the main goal is to study deployment behavior rather than maximize answer quality.

The value case weakens when developers must add review, retries, routing, or human correction. Granite 4.0 H 350M ranks 582 of 595 on the Artificial Analysis Intelligence Index and 339 of 343 on LiveCodeBench. If those results translate into a high correction rate for the intended task, engineering and review costs can exceed the saved token fees. The data brief does not provide error rates, request limits, uptime, hosting fees, or self-hosting requirements, so the total cost cannot be established.

The recorded speed values are 0 output tokens per second and 0 seconds to first token. These values should not be interpreted as proof of exceptional speed. They may indicate missing or unavailable serving measurements. The research brief also found no verified source for the model’s current endpoint or runtime behavior.

For cost-sensitive selection, Granite deserves a short controlled trial rather than an automatic production approval. Measure answer acceptance, correction effort, failure severity, and actual access conditions. Data provided by Artificial Analysis.

05

Granite 4.0 H 350M fits low-stakes trials, but developers should choose a stronger free model for important work

Granite 4.0 H 350M is reasonable for disposable experiments and difficult to recommend for production features that depend on reliable reasoning.

Choose Granite when all four conditions are true: the task is low stakes, the output can be checked automatically or manually, the recorded $0 price is valuable, and the model can be accessed through a dependable route. Examples include rough text categorization, simple routing tests, local model experiments, and prototypes where failure does not affect customers or business records.

Avoid Granite for autonomous coding, technical research, regulated advice, complex support answers, financial workflows, and tool-using agents. Its rankings on MMLU-Pro, LiveCodeBench, SciCode, and IFBench do not provide enough confidence for these uses. The research brief provides no official documentation or community evidence that could offset the benchmark concerns.

A stronger free alternative deserves an early comparison. Gemma 3 4B Instruct is also recorded at $0 per 1M blended tokens and has higher supplied scores on MMLU-Pro, LiveCodeBench, SciCode, and Math 500. Those figures do not settle deployment questions, and the article is not a full comparison between the two models. They do show that Granite does not have an obvious free-tier quality advantage in the available data.

The final decision should depend on a task-level test. Create a small set of real prompts, define what counts as an acceptable answer, and compare Granite against one stronger free model. If Granite does not win on correction-adjusted cost, do not select it merely because its token price is recorded as zero. Data provided by Artificial Analysis.

06

Granite 4.0 H 350M questions developers should answer before adoption

Granite 4.0 H 350M should be adopted only after access, licensing, and task-level quality have been verified by the developer.

The supplied research brief contains no official IBM documentation, reliable community discussion, or verified public pricing page. The benchmark snapshot is therefore useful for risk screening, but it cannot answer every deployment question. The following FAQ separates what the data supports from what remains unknown.

Frequently asked questions

Is Granite 4.0 H 350M good enough for production use?

Granite 4.0 H 350M is not a strong default for production use because it ranks 582 of 595 on the Artificial Analysis Intelligence Index and lacks verified deployment documentation. It may still fit a narrow, low-stakes task after task-specific testing.

Is Granite 4.0 H 350M really free?

Granite 4.0 H 350M has a recorded blended price of $0 per 1M tokens in the supplied data snapshot, but the research brief found no verified current pricing page or stable access route. Developers must confirm whether free access exists in their chosen environment.

Can Granite 4.0 H 350M write reliable code?

Granite 4.0 H 350M should not be assumed reliable for coding because it ranks 339 of 343 on LiveCodeBench at 0.019 and 559 of 567 on SciCode at 0.017. Use automated tests and human review before accepting generated code.

What is the biggest unknown about Granite 4.0 H 350M?

The biggest unknown is whether Granite 4.0 H 350M is currently available through a stable, documented service, because no verifiable official release, developer documentation, pricing page, or community testing was found in the research brief.

Should developers choose Granite 4.0 H 350M over Gemma 3 4B Instruct?

Developers should choose Granite 4.0 H 350M over Gemma 3 4B Instruct only when its access conditions or resource requirements are materially better for the specific task. The supplied data gives Gemma 3 4B Instruct stronger results across several important evaluations while recording the same $0 blended price.

Sources

  1. Artificial AnalysisBenchmark rankings, evaluation scores, recorded pricing, speed fields, model comparison data, and the required data attribution.

Published: