LLM API Pricing

LLM API pricing — 206 models, July 2026

Input price, output price, context window, and cache rates for every major model — GPT-4o mini, Claude Haiku 4.5, Gemini 2.5 Flash, DeepSeek V3.2, Llama 3.3 70B, and more. Prices pulled live from OpenRouter and updated on every deploy.

Most-searched models

openai

OpenAI: GPT-4o-mini

Input
$0.150/1M
Output
$0.600/1M
Context
128K
Compare Gemini 3.5 Flash vs GPT-4o mini →

anthropic

Anthropic: Claude Haiku 4.5

Input
$1.000/1M
Output
$5.000/1M
Context
200K
Compare Gemini 3.5 Flash vs Claude Haiku 4.5 →

google

Google: Gemini 2.5 Flash

Input
$0.300/1M
Output
$2.500/1M
Context
1,048.576K
Compare GPT-4o mini vs Gemini 2.5 Flash →

deepseek

DeepSeek: DeepSeek V3.2

Input
$0.229/1M
Output
$0.343/1M
Context
131.072K
Compare GLM 5.2 vs DeepSeek V3.2 →

Meta

Meta: Llama 3.3 70B Instruct

Input
$0.100/1M
Output
$0.320/1M
Context
131.072K
Compare Llama 3.3 70B vs GPT-4o mini →

openai

OpenAI: GPT-4o

Input
$2.500/1M
Output
$10.000/1M
Context
128K
Compare o1 vs GPT-4o →
206 models
ModelInput $/1MOutput $/1MContextCache Read

Meta: Llama 3.1 8B Instruct

Meta

$0.020$0.030131.072K

Mistral: Mistral Nemo

mistralai

$0.020$0.030131.072K

Meta: Llama 3.2 1B Instruct

Meta

$0.027$0.201131.072K

OpenAI: gpt-oss-20b

openai

$0.029$0.140131.072K

OpenAI: gpt-oss-120b

openai

$0.030$0.150131.072K

Cohere: Command R7B (12-2024)

cohere

$0.037$0.150128K

Qwen: Qwen2.5 7B Instruct

qwen

$0.040$0.100131.072K

Qwen: Qwen3 30B A3B Instruct 2507

qwen

$0.048$0.193131.072K

OpenAI: GPT-5 Nano

openai

$0.050$0.400400K$0.010

Qwen: Qwen3 8B

qwen

$0.050$0.400131.072K$0.050

Google: Gemma 3 4B

google

$0.050$0.100131.072K

Google: Gemma 3 12B

google

$0.050$0.150131.072K

Mistral: Mistral Small 3

mistralai

$0.050$0.08032.768K

Meta: Llama 3.2 3B Instruct

Meta

$0.051$0.335131.072K

Google: Gemma 4 26B A4B

google

$0.060$0.330262.144K

Z.ai: GLM 4.7 Flash

z-ai

$0.060$0.400202.752K$0.010

Google: Gemma 3n 4B

google

$0.060$0.12032.768K

Qwen: Qwen3.5-Flash

qwen

$0.065$0.2601,000K

Qwen: Qwen3 Coder 30B A3B Instruct

qwen

$0.070$0.270160K

OpenAI: gpt-oss-safeguard-20b

openai

$0.075$0.300131.072K$0.037

Mistral: Mistral Small 3.2 24B

mistralai

$0.075$0.200128K

Qwen: Qwen3 VL 8B Instruct

qwen

$0.080$0.500256K

Qwen: Qwen3 30B A3B Thinking 2507

qwen

$0.080$0.400131.072K$0.080

Qwen: Qwen3 32B

qwen

$0.080$0.280131.072K

Google: Gemma 3 27B

google

$0.080$0.160131.072K

DeepSeek: DeepSeek V4 Flash

deepseek

$0.090$0.1801,048.576K$0.020

Qwen: Qwen3 Next 80B A3B Instruct

qwen

$0.090$1.100262.144K

Qwen: Qwen3 235B A22B Instruct 2507

qwen

$0.090$0.100262.144K

Qwen: Qwen3 Next 80B A3B Thinking

qwen

$0.098$0.780262.144K

Qwen: Qwen3.5-9B

qwen

$0.100$0.150262.144K

Mistral: Ministral 3 3B 2512

mistralai

$0.100$0.100131.072K$0.010

Mistral: Voxtral Small 24B 2507

mistralai

$0.100$0.30032K$0.010

Google: Gemini 2.5 Flash Lite Preview 09-2025

google

$0.100$0.4001,048.576K$0.010

Qwen: Qwen3 235B A22B Thinking 2507

qwen

$0.100$0.100262.144K$0.100

Google: Gemini 2.5 Flash Lite

google

$0.100$0.4001,048.576K$0.010

Qwen: Qwen3 14B

qwen

$0.100$0.240131.702K

OpenAI: GPT-4.1 Nano

openai

$0.100$0.4001,047.576K$0.025

Meta: Llama 4 Scout

Meta

$0.100$0.30010,000K

Meta: Llama 3.3 70B Instruct

Meta

$0.100$0.320131.072K

Qwen: Qwen3 VL 32B Instruct

qwen

$0.104$0.416262.144K

Qwen: Qwen3 Coder Next

qwen

$0.110$0.800262.144K$0.070

Qwen: Qwen3 VL 8B Thinking

qwen

$0.117$1.365256K

Google: Gemma 4 31B

google

$0.120$0.350262.144K$0.090

Qwen: Qwen3 30B A3B

qwen

$0.120$0.500131.072K

Qwen: Qwen3 VL 30B A3B Thinking

qwen

$0.130$1.560131.072K

Qwen: Qwen3 VL 30B A3B Instruct

qwen

$0.130$0.520262.144K

Z.ai: GLM 4.5 Air

z-ai

$0.130$0.850131.072K$0.025

Qwen: Qwen3.6 35B A3B

qwen

$0.140$1.000262.144K

Qwen: Qwen3.5-35B-A3B

qwen

$0.140$1.000262.144K$0.050

Meta: Llama 3 8B Instruct

Meta

$0.140$0.1408.192K

Mistral: Mistral Small 4

mistralai

$0.150$0.600262.144K$0.015

Mistral: Ministral 3 8B 2512

mistralai

$0.150$0.150262.144K$0.015

Meta: Llama 4 Maverick

Meta

$0.150$0.6001,048.576K

OpenAI: GPT-4o-mini Search Preview

openai

$0.150$0.600128K

Cohere: Command R (08-2024)

cohere

$0.150$0.600128K

OpenAI: GPT-4o-mini

openai

$0.150$0.600128K$0.075

OpenAI: GPT-4o-mini (2024-07-18)

openai

$0.150$0.600128K$0.075

Meta: Llama Guard 4 12B

Meta

$0.180$0.180163.84K

Qwen: Qwen3.6 Flash

qwen

$0.188$1.1251,000K

Qwen: Qwen3.5-27B

qwen

$0.195$1.560262.144K

Qwen: Qwen3 Coder Flash

qwen

$0.195$0.9751,000K$0.039

OpenAI: GPT-5.4 Nano

openai

$0.200$1.250400K$0.020

Mistral: Ministral 3 14B 2512

mistralai

$0.200$0.200262.144K$0.020

Qwen: Qwen3 VL 235B A22B Instruct

qwen

$0.200$0.880262.144K$0.110

DeepSeek: DeepSeek V3 0324

deepseek

$0.200$0.770163.84K$0.135

Mistral: Saba

mistralai

$0.200$0.60032.768K$0.020

DeepSeek: DeepSeek V3

deepseek

$0.200$0.800131.072K

DeepSeek: DeepSeek V3.1

deepseek

$0.210$0.790163.84K$0.130

Qwen: Qwen3 Coder 480B A35B

qwen

$0.220$1.8001,048.576K

DeepSeek: DeepSeek V3.2

deepseek

$0.229$0.343131.072K$0.023

Google: Gemini 3.1 Flash Lite

google

$0.250$1.5001,048.576K$0.025

Google: Gemini 3.1 Flash Lite Preview

google

$0.250$1.5001,048.576K$0.025

OpenAI: GPT-5.1-Codex-Mini

openai

$0.250$2.000400K$0.025

OpenAI: GPT-5 Mini

openai

$0.250$2.000400K$0.025

Anthropic: Claude 3 Haiku

anthropic

$0.250$1.250200K$0.030

Qwen: Qwen3.5-122B-A10B

qwen

$0.260$2.080262.144K

Qwen: Qwen3.5 Plus 2026-02-15

qwen

$0.260$1.5601,000K

Qwen: Qwen3 VL 235B A22B Thinking

qwen

$0.260$2.600131.072K

Qwen: Qwen Plus 0728 (thinking)

qwen

$0.260$0.7801,000K

Qwen: Qwen Plus 0728

qwen

$0.260$0.7801,000K

Qwen: Qwen-Plus

qwen

$0.260$0.7801,000K$0.052

DeepSeek: DeepSeek V3.2 Exp

deepseek

$0.270$0.410163.84K

DeepSeek: DeepSeek V3.1 Terminus

deepseek

$0.270$0.950163.84K$0.130

Qwen: Qwen3.6 27B

qwen

$0.289$2.650262.144K

Qwen: Qwen3.5 Plus 2026-04-20

qwen

$0.300$1.8001,000K

Z.ai: GLM 4.6V

z-ai

$0.300$0.900131.072K$0.055

Google: Nano Banana (Gemini 2.5 Flash Image)

google

$0.300$2.50032.768K$0.030

Mistral: Codestral 2508

mistralai

$0.300$0.900256K$0.030

Google: Gemini 2.5 Flash

google

$0.300$2.5001,048.576K$0.030

Qwen: Qwen3.7 Plus

qwen

$0.320$1.2801,000K$0.064

Qwen: Qwen3.6 Plus

qwen

$0.325$1.9501,000K

Meta: Llama 3.2 11B Vision Instruct

Meta

$0.345$0.345131.072K

Mistral: Mistral Small 3.1 24B

mistralai

$0.351$0.555128K

Qwen2.5 72B Instruct

qwen

$0.360$0.400131.072K

Qwen: Qwen3.5 397B A17B

qwen

$0.385$2.450256K

Z.ai: GLM 4.7

z-ai

$0.400$1.750202.752K$0.080

Mistral: Devstral 2 2512

mistralai

$0.400$2.000262.144K$0.040

Mistral: Mistral Medium 3.1

mistralai

$0.400$2.000131.072K$0.040

Mistral: Mistral Medium 3

mistralai

$0.400$2.000131.072K$0.040

OpenAI: GPT-4.1 Mini

openai

$0.400$1.6001,047.576K$0.100

Meta: Llama 3.1 70B Instruct

Meta

$0.400$0.400131.072K

Z.ai: GLM 4.6

z-ai

$0.430$1.740202.752K$0.080

DeepSeek: DeepSeek V4 Pro

deepseek

$0.435$0.8701,048.576K$0.004

Qwen: Qwen3 235B A22B

qwen

$0.455$1.820131.072K

Google: Nano Banana 2 (Gemini 3.1 Flash Image)

google

$0.500$3.000131.072K

Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)

google

$0.500$3.000131.072K

Google: Gemini 3 Flash Preview

google

$0.500$3.0001,048.576K$0.050

Mistral: Mistral Large 3 2512

mistralai

$0.500$1.500262.144K$0.050

DeepSeek: R1 0528

deepseek

$0.500$2.150163.84K$0.350

OpenAI: GPT-3.5 Turbo

openai

$0.500$1.50016.385K

Z.ai: GLM 5

z-ai

$0.600$1.920202.752K$0.120

OpenAI: GPT Audio Mini

openai

$0.600$2.400128K

Z.ai: GLM 4.5V

z-ai

$0.600$1.80065.536K$0.110

Z.ai: GLM 4.5

z-ai

$0.600$2.200131.072K$0.110

Qwen: Qwen3 Coder Plus

qwen

$0.650$3.2501,000K$0.130

Google: Gemma 2 27B

google

$0.650$0.6508.192K

Qwen2.5 Coder 32B Instruct

qwen

$0.660$1.000128K

DeepSeek: R1

deepseek

$0.700$2.500163.84K

OpenAI: GPT-5.4 Mini

openai

$0.750$4.500400K$0.075

Qwen: Qwen3 Max Thinking

qwen

$0.780$3.900262.144K

Qwen: Qwen3 Max

qwen

$0.780$3.900262.144K$0.156

Qwen: Qwen2.5 VL 72B Instruct

qwen

$0.800$1.000131.072K$0.400

DeepSeek: R1 Distill Llama 70B

deepseek

$0.800$0.800128K

Z.ai: GLM 5.2

z-ai

$0.950$3.0001,048.576K$0.180

Z.ai: GLM 5.1

z-ai

$0.980$3.080202.752K$0.182

Anthropic: Claude Haiku 4.5

anthropic

$1.000$5.000200K$0.100

Perplexity: Sonar

perplexity

$1.000$1.000127.072K

OpenAI: GPT-3.5 Turbo (older v0613)

openai

$1.000$2.0004.095K

Qwen: Qwen3.6 Max Preview

qwen

$1.040$6.240262.144K

OpenAI: o4 Mini High

openai

$1.100$4.400200K$0.275

OpenAI: o4 Mini

openai

$1.100$4.400200K$0.275

OpenAI: o3 Mini High

openai

$1.100$4.400200K$0.550

OpenAI: o3 Mini

openai

$1.100$4.400200K$0.550

Z.ai: GLM 5V Turbo

z-ai

$1.200$4.000202.752K$0.240

Z.ai: GLM 5 Turbo

z-ai

$1.200$4.000262.144K$0.240

Qwen: Qwen3.7 Max

qwen

$1.250$3.7501,000K$0.250

OpenAI: GPT-5.1-Codex-Max

openai

$1.250$10.000400K$0.125

OpenAI: GPT-5.1

openai

$1.250$10.000400K$0.130

OpenAI: GPT-5.1 Chat

openai

$1.250$10.000128K$0.130

OpenAI: GPT-5.1-Codex

openai

$1.250$10.000400K$0.130

OpenAI: GPT-5 Codex

openai

$1.250$10.000400K$0.125

OpenAI: GPT-5 Chat

openai

$1.250$10.000128K$0.125

OpenAI: GPT-5

openai

$1.250$10.000400K$0.125

Google: Gemini 2.5 Pro

google

$1.250$10.0001,048.576K$0.125

Google: Gemini 2.5 Pro Preview 06-05

google

$1.250$10.0001,048.576K$0.125

Google: Gemini 2.5 Pro Preview 05-06

google

$1.250$10.0001,048.576K$0.125

Google: Gemini 3.5 Flash

google

$1.500$9.0001,048.576K$0.150

Mistral: Mistral Medium 3.5

mistralai

$1.500$7.500262.144K

OpenAI: GPT-3.5 Turbo Instruct

openai

$1.500$2.0004.095K

OpenAI: GPT-5.3 Chat

openai

$1.750$14.000128K$0.175

OpenAI: GPT-5.3-Codex

openai

$1.750$14.000400K$0.175

OpenAI: GPT-5.2-Codex

openai

$1.750$14.000400K$0.175

OpenAI: GPT-5.2 Chat

openai

$1.750$14.000128K$0.175

OpenAI: GPT-5.2

openai

$1.750$14.000400K$0.175

Google: Nano Banana Pro (Gemini 3 Pro Image)

google

$2.000$12.00065.536K$0.200

Google: Gemini 3.1 Pro Preview Custom Tools

google

$2.000$12.0001,048.756K$0.200

Google: Gemini 3.1 Pro Preview

google

$2.000$12.0001,048.576K$0.200

Google: Nano Banana Pro (Gemini 3 Pro Image Preview)

google

$2.000$12.00065.536K$0.200

OpenAI: o4 Mini Deep Research

openai

$2.000$8.000200K$0.500

OpenAI: o3

openai

$2.000$8.000200K$0.500

OpenAI: GPT-4.1

openai

$2.000$8.0001,047.576K$0.500

Perplexity: Sonar Reasoning Pro

perplexity

$2.000$8.000128K

Perplexity: Sonar Deep Research

perplexity

$2.000$8.000128K

Mistral Large 2407

mistralai

$2.000$6.000131.072K$0.200

Mistral: Mixtral 8x22B Instruct

mistralai

$2.000$6.00065.536K$0.200

Mistral Large

mistralai

$2.000$6.000128K$0.200

OpenAI: GPT-5.4

openai

$2.500$15.0001,050K$0.250

OpenAI: GPT Audio

openai

$2.500$10.000128K

OpenAI: GPT-5 Image Mini

openai

$2.500$2.000400K$0.250

Cohere: Command A

cohere

$2.500$10.000256K

OpenAI: GPT-4o Search Preview

openai

$2.500$10.000128K

OpenAI: GPT-4o (2024-11-20)

openai

$2.500$10.000128K$1.250

Cohere: Command R+ (08-2024)

cohere

$2.500$10.000128K

OpenAI: GPT-4o (2024-08-06)

openai

$2.500$10.000128K$1.250

OpenAI: GPT-4o

openai

$2.500$10.000128K

Anthropic: Claude Sonnet 4.6

anthropic

$3.000$15.0001,000K$0.300

Perplexity: Sonar Pro Search

perplexity

$3.000$15.000200K

Anthropic: Claude Sonnet 4.5

anthropic

$3.000$15.0001,000K$0.300

Anthropic: Claude Sonnet 4

anthropic

$3.000$15.0001,000K$0.300

Perplexity: Sonar Pro

perplexity

$3.000$15.000200K

OpenAI: GPT-3.5 Turbo 16k

openai

$3.000$4.00016.385K

Anthropic: Claude Opus 4.8

anthropic

$5.000$25.0001,000K$0.500

OpenAI: GPT Chat Latest

openai

$5.000$30.000400K$0.500

OpenAI: GPT-5.5

openai

$5.000$30.0001,050K$0.500

Anthropic: Claude Opus 4.7

anthropic

$5.000$25.0001,000K$0.500

Anthropic: Claude Opus 4.6

anthropic

$5.000$25.0001,000K$0.500

Anthropic: Claude Opus 4.5

anthropic

$5.000$25.000200K$0.500

OpenAI: GPT-4o (2024-05-13)

openai

$5.000$15.000128K

OpenAI: GPT-5.4 Image 2

openai

$8.000$15.000272K$2.000

Anthropic: Claude Fable 5

anthropic

$10.000$50.0001,000K$1.000

Anthropic: Claude Opus 4.8 (Fast)

anthropic

$10.000$50.0001,000K$1.000

OpenAI: GPT-5 Image

openai

$10.000$10.000400K$1.250

OpenAI: o3 Deep Research

openai

$10.000$40.000200K$2.500

OpenAI: GPT-4 Turbo

openai

$10.000$30.000128K

OpenAI: GPT-4 Turbo Preview

openai

$10.000$30.000128K

OpenAI: GPT-5 Pro

openai

$15.000$120.000400K

Anthropic: Claude Opus 4.1

anthropic

$15.000$75.000200K$1.500

Anthropic: Claude Opus 4

anthropic

$15.000$75.000200K$1.500

OpenAI: o1

openai

$15.000$60.000200K$7.500

OpenAI: o3 Pro

openai

$20.000$80.000200K

OpenAI: GPT-5.2 Pro

openai

$21.000$168.000400K

Anthropic: Claude Opus 4.7 (Fast)

anthropic

$30.000$150.0001,000K$3.000

OpenAI: GPT-5.5 Pro

openai

$30.000$180.0001,050K

Anthropic: Claude Opus 4.6 (Fast)

anthropic

$30.000$150.0001,000K$3.000

OpenAI: GPT-5.4 Pro

openai

$30.000$180.0001,050K

OpenAI: GPT-4

openai

$30.000$60.0008.191K

Turn these prices into a margin forecast

Plug any model's price into the Margin Simulator to see gross margin, cost per user, and breakeven MAU for your product — or compare two models head-to-head.

LLM API pricing — frequently asked questions

How much does GPT-4o mini cost?

GPT-4o mini is priced at $0.15 per million input tokens and $0.60 per million output tokens via the OpenAI API (routed through OpenRouter). It is OpenAI's most cost-efficient hosted model and is optimised for high-volume, latency-sensitive workloads.

What is the cheapest LLM API available in 2026?

Open-source models hosted via OpenRouter — including Llama 3.1 8B, Gemini 2.0 Flash, and DeepSeek V3 — are among the cheapest options, often under $0.05 per million input tokens. For hosted frontier budget models, GPT-4o mini, Claude Haiku 4.5, and Gemini 2.5 Flash lead the pack.

How does Claude Haiku 4.5 pricing compare to GPT-4o mini?

Both sit in the budget tier of their respective providers and are comparably priced. The cheaper model for your workload depends on your input/output token ratio — use the table and the compare calculator to find out which wins for your specific usage.

Why do input and output token prices differ?

Generating (output) tokens is computationally more expensive than reading (input) tokens. Most models price output at 2–4× the input rate. For output-heavy workloads like chat or code generation, the output price dominates your bill.

What is a token in LLM pricing?

A token is roughly 0.75 English words, or approximately 4 characters. A 1,000-word document contains around 1,333 tokens. All prices in this table are quoted per million tokens ($/1M) to allow direct comparison across models.

What is prompt caching and how does it affect cost?

Prompt caching lets you reuse expensive input context across multiple calls at a reduced rate — often 80–90% cheaper than a fresh input read. Models that support it (Claude, Gemini 2.5) charge a slightly higher cache-write rate on first use, then discounted cache-read rates on subsequent calls. The 'Cache Read' column in the table above shows where it's available.

LLM token prices change frequently — OpenAI, Anthropic, and Google all adjust rates as competition intensifies. This table reflects live prices from OpenRouter and is updated on every deploy. For per-model cost forecasting, use the Margin Simulator to model gross margin at your MAU.

Related tools: Compare two LLM models side-by-side · LLM cost per user calculator · How to calculate LLM cost per user