Gemini 3.1 Pro Api Vs Gpt-5 Api Token Cost Comparison
Compare Gemini 3.1 Pro API vs GPT-5 API token costs. Review input, cached-input, and output prices to estimate AI usage expenses with official rates.
Please complete this field to continue.
Please complete this field to continue.
Please complete this field to continue.
Results
Gemini 3.1 Pro API vs GPT-5 API Per Token Cost
Quick answer: Gemini 3.1 Pro API vs GPT-5 API Per Token Cost compares the standard API prices of Google's Gemini 3.1 Pro Preview and OpenAI's GPT-5. It helps developers estimate input, output, and cached-input expenses using published prices per one million tokens.
Gemini 3.1 Pro and GPT-5 use different token-pricing structures, so comparing their input and output rates separately is important. Gemini 3.1 Pro Preview charges different rates depending on whether the prompt contains 200,000 tokens or fewer, while GPT-5 has separate standard input, cached-input, and output rates.
This comparison uses the standard paid API pricing published by Google and OpenAI. Prices are in US dollars and reflect the published rates available on October 11, 2026. Model prices and availability can change, so check the official pricing pages before making production budget decisions.
TL;DR / Key Takeaways
- Gemini 3.1 Pro Preview: $2.00 per million input tokens and $12.00 per million output tokens for prompts up to 200,000 tokens.
- GPT-5: $1.25 per million input tokens, $0.125 per million cached input tokens, and $10.00 per million output tokens.
- Long prompts: Gemini 3.1 Pro Preview's standard input rate rises to $4.00 and its output rate to $18.00 per million tokens when prompts exceed 200,000 tokens.
- Cost estimation: Actual expenses depend on input volume, generated output, cached tokens, and the selected model's billing rules.
API Pricing Comparison
The following table compares standard paid API rates per one million tokens. Gemini 3.1 Pro figures apply to prompts of 200,000 tokens or fewer unless otherwise specified.
| Pricing Metric | Gemini 3.1 Pro Preview | GPT-5 |
|---|---|---|
| Standard input | $2.00 / 1M tokens | $1.25 / 1M tokens |
| Cached input | $0.20 / 1M tokens | $0.125 / 1M tokens |
| Output | $12.00 / 1M tokens | $10.00 / 1M tokens |
| Input for prompts over 200K tokens | $4.00 / 1M tokens | Standard rate: $1.25 / 1M tokens |
| Output for prompts over 200K tokens | $18.00 / 1M tokens | Standard rate: $10.00 / 1M tokens |
Official references: Google Gemini API pricing and OpenAI API pricing.
Cached-input prices apply only when requests qualify for the provider's supported caching mechanism. Do not assume that every repeated prompt automatically receives the cached rate.
How to Use Gemini 3.1 Pro API vs GPT-5 API Per Token Cost?
- Identify the model: Select Gemini 3.1 Pro Preview or GPT-5 and confirm the API pricing tier you intend to use.
- Estimate input tokens: Count the tokens sent to the model, including relevant system instructions, conversation history, and other supported input content.
- Estimate output tokens: Determine how many tokens the model is expected to generate. For reasoning models, billed output may include reasoning or thinking tokens under the provider's rules.
- Account for caching: Separate eligible cached input from regular input when the API's billing model supports that distinction.
- Calculate the cost: Multiply each token category by its price per million tokens and add the resulting charges.
Input and Output Cost Example
Consider a representative workload consisting of 100,000 input tokens and 20,000 output tokens. Assume standard paid API rates, a Gemini prompt below the 200,000-token pricing threshold, no cached input, and no additional tool charges.
| Cost Component | Gemini 3.1 Pro Preview | GPT-5 |
|---|---|---|
| Input cost | 100,000 / 1,000,000 × $2.00 = $0.20 | 100,000 / 1,000,000 × $1.25 = $0.125 |
| Output cost | 20,000 / 1,000,000 × $12.00 = $0.24 | 20,000 / 1,000,000 × $10.00 = $0.20 |
| Total | $0.44 | $0.325 |
For this illustrative workload, GPT-5 costs $0.115 less per request, approximately 26.1% less than Gemini 3.1 Pro Preview. This is a price comparison for the assumed token counts, not a benchmark of model quality, output length, or task success.
Token Cost Formula
For a simple request using standard input and output rates:
Total Cost = (Input Tokens / 1,000,000 × Input Rate) + (Output Tokens / 1,000,000 × Output Rate)
When cached input is billed separately, use:
Total Cost = (Regular Input Tokens / 1,000,000 × Input Rate) + (Cached Input Tokens / 1,000,000 × Cached Input Rate) + (Output Tokens / 1,000,000 × Output Rate)
These formulas estimate model-token charges. Add separately billed tools, search requests, storage, or other services when applicable. Apply the provider's relevant pricing tier, including any prompt-length thresholds, before calculating the estimate.
How Much Does One Million Tokens Cost?
At the standard paid rates for prompts up to 200,000 tokens, one million input tokens cost $2.00 with Gemini 3.1 Pro Preview and $1.25 with GPT-5. One million output tokens cost $12.00 with Gemini 3.1 Pro Preview and $10.00 with GPT-5.
Input and output are billed separately. Therefore, the cost of processing one million total tokens depends on how many are input tokens and how many are output tokens. A workload with a large generated response can have a substantially different cost from a workload that mainly processes input documents.
Technical Reference: Token Pricing Rules
| Pricing Scenario | Gemini 3.1 Pro Preview | GPT-5 |
|---|---|---|
| Shorter prompts | $2.00 input / $12.00 output per 1M tokens | $1.25 input / $10.00 output per 1M tokens |
| Long prompts | $4.00 input / $18.00 output per 1M tokens when prompts exceed 200K tokens | Published standard input/output rates are $1.25 / $10.00 per 1M tokens |
| Cached input | $0.20 per 1M tokens for prompts up to 200K tokens | $0.125 per 1M tokens |
| Batch processing | Separate Batch pricing applies; consult Google's current pricing table | Separate Batch pricing applies; consult OpenAI's current pricing table |
| Additional API tools | Applicable tool charges may be additional to token costs | Applicable tool charges may be additional to token costs |
These rates refer to the named API models, not every Gemini or GPT model. Gemini's model identifier is gemini-3.1-pro-preview, while GPT-5 uses gpt-5. Verify the model identifier in your integration because similarly named models can have different prices.
How the Comparison Works
This comparison separates the main billable token categories instead of treating all tokens as one undifferentiated quantity. Standard input represents tokens submitted to the model; output represents tokens generated by the model, subject to the provider's billing definitions; cached input represents eligible input tokens processed through supported prompt-caching features.
For a reliable budget, use actual usage counts from API responses or provider billing records where available. Tokenizers differ between model families, so identical text does not necessarily produce identical token counts. Models may also generate different response lengths or reasoning-token volumes for the same task.
Edge Cases and Limitations
- Gemini's 200K threshold: The applicable rate changes when a prompt exceeds 200,000 tokens. Use the correct tier for the request.
- Cached input eligibility: A repeated prompt does not guarantee that all input tokens qualify for discounted cached pricing.
- Reasoning tokens: Output billing may include reasoning or thinking tokens, depending on the model's documented pricing rules.
- Different tokenizers: Matching word counts do not guarantee matching token counts across Gemini and GPT-5.
- Tool usage: Web search, grounding, and other separately priced services may add charges beyond standard model-token costs.
- Batch and service tiers: Alternative processing modes may have different rates or conditions. This comparison focuses on standard paid pricing.
- Pricing changes: Published prices can change. Recheck the official provider pages before using estimates for procurement or production budgets.
Security, Privacy, and Billing Transparency
This page is a pricing reference, not an API proxy or a billing service. It does not claim to inspect your API account, retrieve your actual usage, or calculate a provider invoice. The examples use manually specified token counts and published pricing figures.
Google's pricing documentation distinguishes free and paid usage policies, including different statements about whether content is used to improve its products. OpenAI's API policies also govern data handling. Review each provider's current terms and privacy documentation for the account, service, and data type you use rather than assuming that pricing alone establishes a privacy guarantee.
Frequently Asked Questions
Is GPT-5 cheaper than Gemini 3.1 Pro?
At the cited standard paid rates for prompts up to 200,000 tokens, GPT-5 has lower input and output token prices. The total cost still depends on actual token counts, caching, and additional services.
What is the GPT-5 API input price?
The published standard GPT-5 input price is $1.25 per one million tokens. Eligible cached input is priced at $0.125 per one million tokens.
What is the Gemini 3.1 Pro API output price?
Gemini 3.1 Pro Preview output costs $12.00 per one million tokens for prompts up to 200,000 tokens and $18.00 per one million tokens for prompts above that threshold.
How do I calculate API token costs?
Divide each token count by one million, multiply it by the corresponding per-million-token price, and add input, cached-input, and output charges as applicable. Include separately billed tools when estimating total usage.
Does one million tokens mean one million words?
No. Tokens are units of text or multimodal content processing, not words. Token counts depend on the content, language, encoding, and model tokenizer.
Are cached tokens and batch requests priced the same as standard requests?
No. Eligible cached input can have a discounted rate, and batch processing can use a separate pricing schedule. Confirm eligibility, processing mode, and current provider terms before estimating costs.
Can these prices be used as a guaranteed invoice estimate?
No. They support estimates based on stated assumptions. Actual charges may differ because of token counts, prompt-length thresholds, cached input, tool calls, processing mode, or subsequent pricing changes.
Author
Author Name: Daniel Mercer
Author Description: Technology writer specializing in AI API economics, developer tooling, and cloud software cost analysis.
Technical Review: The pricing methodology distinguishes standard input, cached input, and output rates and accounts for Gemini 3.1 Pro Preview's documented 200,000-token prompt threshold. Verify current prices against the official provider documentation before deployment.