Skip to content

AI model analysis

Gemini 1.5 Pro (Sep '24) vs GPT-5.4 Pro (xhigh): Which Should Developers Choose?

A careful developer comparison of Gemini 1.5 Pro (Sep '24) and GPT-5.4 Pro (xhigh), covering measured capability, cost visibility, availability risk, and production selection criteria.

Gemini 1.5 Pro (Sep '24) vs GPT-5.4 Pro (xhigh): Which Should Developers Choose?
Summary

- **Winner overall:** Gemini 1.5 Pro (Sep '24), the only model with recorded evaluation results, including a 23.6 coding index, although current API availability is unverified - **Cheaper:** Gemini 1.5 Pro (Sep '24) at $0 vs $67.5 per 1M blended tokens - **Faster:** Gemini 1.5 Pro (Sep '24) and GPT-5.4 Pro (xhigh) tied at 0 median output tokens per second - **Pick Gemini 1.5 Pro (Sep '24) when:** you already have a verified working endpoint and need a model with recorded multimodal and long-context positioning - **Watch out:** GPT-5.4 Pro (xhigh) has no recorded evaluation values, while Gemini 1.5 Pro (Sep '24) has no current official price or active endpoint

01

Gemini 1.5 Pro vs GPT-5.4 Pro: The Practical Verdict

Gemini 1.5 Pro (Sep '24) is the more measurable option, while GPT-5.4 Pro (xhigh) is the less documented option, so neither model is a safe default for a new production integration without endpoint verification.

The comparison is unusual because the two models do not have symmetrical evidence. Gemini 1.5 Pro (Sep '24) has recorded evaluation results and a listed blended price of $0 in the supplied data. GPT-5.4 Pro (xhigh) has no recorded evaluation values and a blended price of $67.5 per 1M tokens. At the same time, current official documentation does not list either exact model entry.

That difference changes the buying question. Developers cannot simply ask which model scores higher or costs less. They must first ask whether the named model can still be called, whether its price is still valid, and whether the benchmark evidence describes the same version that the application will receive.

Google’s current Gemini API model documentation no longer presents an active model card for Gemini 1.5 Pro (Sep '24). OpenAI’s current model documentation does not list GPT-5.4 Pro (xhigh). The strongest conclusion is therefore operational: verify access before optimizing around either model.

02

What the Evidence Actually Shows

Gemini 1.5 Pro (Sep '24) has the stronger evidence record, but GPT-5.4 Pro (xhigh) cannot be judged as weaker from missing measurements alone.

The supplied evaluation data records Gemini 1.5 Pro (Sep '24) at 9.9 on the Artificial Analysis Intelligence Index and 23.6 on the Artificial Analysis Coding Index. It also records scores of 0.75 on MMLU Pro, 0.589 on GPQA, 0.046 on HLE, 0.316 on LiveCodeBench, 0.295 on SciCode, 0.876 on Math 500, and 0.23 on AIME. GPT-5.4 Pro (xhigh) has null values across these listed evaluations. Because one model has values and the other does not, the dataset does not establish a benchmark winner.

The product evidence points in a different direction. Google describes the Gemini 1.5 Pro family as a multimodal model for complex tasks, with text, image, video, and audio input and text output. Historical official material positioned the family around a context window of up to about 2 million tokens, but the supplied research could not verify a separate parameter record for the Sep '24 version. The current Gemini API model documentation does not preserve a live parameter table for this exact entry.

OpenAI’s current directory gives general information about newer OpenAI models, including text and image input, text output, multilingual ability, vision, Responses API access, and client SDK access. The same OpenAI model documentation does not confirm that those capabilities apply to GPT-5.4 Pro (xhigh).

The data snapshot is therefore useful for historical comparison, not sufficient for a production purchase decision. Data provided by.

03

Performance: Recorded Capability vs Unknown Capability

Gemini 1.5 Pro (Sep '24) is the only model with measured capability in this comparison, but the missing GPT-5.4 Pro data prevents a reliable quality ranking.

For developers, the recorded Gemini results suggest that the model has evidence across general reasoning, coding, scientific coding, mathematics, and instruction-oriented tasks. The 23.6 coding index and 0.316 LiveCodeBench result provide more decision support than a product description alone. They still do not answer whether Gemini will produce the right code for a specific repository, framework, toolchain, or agent loop.

GPT-5.4 Pro (xhigh) has no supplied value for the coding index, LiveCodeBench, MMLU Pro, GPQA, or any other listed evaluation. That absence is not a zero score. It means the comparison cannot establish whether GPT-5.4 Pro is better, equal, or worse on those tasks. Developers should not convert missing data into a negative quality judgment.

The speed comparison is equally limited. The supplied data records median output speed as 0 median output tokens per second and latency as 0 seconds for both models. Those values do not provide a usable real-world speed distinction. They may represent unavailable measurements rather than instantaneous responses. No documented community testing in the research brief supplies a replacement.

The practical implication is that Gemini is easier to evaluate experimentally because the dataset contains historical measurements. GPT-5.4 Pro requires a live test before any performance claim is credible. A useful test should use the developer’s own repository, prompts, tool calls, and acceptance checks. The supplied research does not provide reproducible community tests for either exact version.

04

Cost: The Cheapest Number May Not Be a Callable Price

Gemini 1.5 Pro (Sep '24) appears cheaper in the supplied data, but its $0 price should not be treated as a confirmed free production offer.

The data lists Gemini 1.5 Pro (Sep '24) at $0 per 1M blended tokens, $0 per 1M input tokens, and $0 per 1M output tokens. GPT-5.4 Pro (xhigh) is listed at $67.5 per 1M blended tokens, $30 per 1M input tokens, and $180 per 1M output tokens. On the visible price fields, Gemini is the clear lower-cost option.

The cost conclusion changes once availability enters the calculation. Google’s current Gemini API pricing page does not list Gemini 1.5 Pro (Sep '24) in a current free, paid, or batch pricing tier. The page also does not provide a current stable endpoint for the exact model. A $0 historical or dataset value is economically irrelevant if requests cannot be sent through a supported route.

GPT-5.4 Pro (xhigh) has the opposite problem. The supplied data gives it a price, but the current OpenAI pricing page does not list the exact model. The price therefore needs live account and endpoint confirmation before budgeting.

For a developer, the expensive model can become cheaper overall if it is callable, stable, and reduces retries, manual review, or migration work. The cheap model can become more expensive if its endpoint fails, its behavior changes, or the team must rebuild the integration around another model. The evidence does not quantify those operational costs, so no total-cost winner can be established.

05

Recommendation by Project Situation

Gemini 1.5 Pro (Sep '24) is the better experimental candidate when a team already has a verified endpoint, while GPT-5.4 Pro (xhigh) needs live validation before selection.

Choose Gemini 1.5 Pro (Sep '24) for a controlled prototype when your account can still call the exact model identifier and your workload benefits from its documented family positioning around multimodal inputs and long context. Its recorded 23.6 coding index and 9.9 intelligence index give the team a starting point for regression tests. They do not prove that it will outperform newer or differently configured systems on your codebase.

Choose GPT-5.4 Pro (xhigh) only after confirming that the exact model exists in the intended OpenAI account, API, and region. The supplied material contains no benchmark result, context window, output limit, stable endpoint, or exact capability record for this version. A team selecting it should treat every quality and reliability assumption as unverified until a live evaluation supports it.

For a new production system, the safest recommendation is to pause the model-level decision and verify two facts first: callable endpoint and current price. The official Gemini API model documentation, Gemini API pricing, OpenAI model documentation, and OpenAI pricing do not currently confirm the exact entries consistently.

The evidence gap is itself a selection result. Gemini wins on available historical measurements and visible dataset cost. Neither model wins on verified current production readiness.

06

Developer Questions Before Choosing

Gemini 1.5 Pro (Sep '24) answers more of the historical comparison questions, but neither model has enough current documentation for an automatic production approval.

The most important unanswered questions concern availability, pricing validity, context limits, output limits, and version identity. The research brief could not verify these details for the exact named models. That matters because developers often choose a model from a benchmark page, then discover that the production API serves a different alias, rejects the identifier, or applies a different price.

A short validation exercise should confirm the exact model name, send representative prompts, record response quality and errors, and check the account’s actual billing treatment. The supplied materials do not provide reproducible community tests for coding experience, speed, stability, or failure behavior. Any claim in those areas should remain provisional until the team produces its own evidence.

Frequently asked questions

Is Gemini 1.5 Pro (Sep '24) better than GPT-5.4 Pro (xhigh) for coding?

Gemini 1.5 Pro (Sep '24) is the only model with a recorded coding result, including a 23.6 Artificial Analysis Coding Index and a 0.316 LiveCodeBench score. GPT-5.4 Pro (xhigh) has no supplied coding measurements, so the evidence cannot prove which model is better for coding.

Which model is cheaper for API usage?

Gemini 1.5 Pro (Sep '24) is cheaper in the supplied dataset, with $0 per 1M blended tokens compared with $67.5 for GPT-5.4 Pro (xhigh). Google’s current pricing page does not list Gemini 1.5 Pro, so the $0 value requires live verification before budgeting.

Can developers still use Gemini 1.5 Pro (Sep '24) in a new project?

Developers should use Gemini 1.5 Pro (Sep '24) in a new project only after verifying a working endpoint and current account access. Google’s current model documentation does not show an active model card for this exact version, and the research found no confirmed stable endpoint.

Does GPT-5.4 Pro (xhigh) support a larger context window or better reasoning?

The supplied research does not establish a context advantage or reasoning advantage for GPT-5.4 Pro (xhigh). Current OpenAI documentation does not list the exact model, and the supplied dataset contains no evaluation values for it, so those claims require a live test or official model record.

Should a production team choose either model today?

A production team should choose neither model automatically until it verifies the exact endpoint, price, version behavior, and application-specific quality. Gemini has stronger historical evidence, while GPT-5.4 Pro has a listed dataset price but no confirmed current model entry or benchmark record.

Sources

  1. Gemini API model documentationVerified Gemini 1.5 Pro model status, historical family positioning, current model directory, and endpoint uncertainty.
  2. Gemini API pricingVerified that the current pricing page does not list Gemini 1.5 Pro (Sep '24) and does not provide a current exact-model price.
  3. OpenAI ModelsVerified that GPT-5.4 Pro (xhigh) is not listed and that general newer-model capabilities cannot be assigned to this exact version.
  4. OpenAI PricingVerified that the current pricing page does not list GPT-5.4 Pro (xhigh) or provide an exact alias mapping.
  5. Artificial AnalysisData attribution for the supplied benchmark, pricing, speed, latency, and model comparison snapshot.

Published: