LLM API Pricing
LLM API pricing — 206 models, July 2026
Input price, output price, context window, and cache rates for every major model — GPT-4o mini, Claude Haiku 4.5, Gemini 2.5 Flash, DeepSeek V3.2, Llama 3.3 70B, and more. Prices pulled live from OpenRouter and updated on every deploy.
Most-searched models
openai
OpenAI: GPT-4o-mini
- Input
- $0.150/1M
- Output
- $0.600/1M
- Context
- 128K
anthropic
Anthropic: Claude Haiku 4.5
- Input
- $1.000/1M
- Output
- $5.000/1M
- Context
- 200K
Google: Gemini 2.5 Flash
- Input
- $0.300/1M
- Output
- $2.500/1M
- Context
- 1,048.576K
deepseek
DeepSeek: DeepSeek V3.2
- Input
- $0.229/1M
- Output
- $0.343/1M
- Context
- 131.072K
Meta
Meta: Llama 3.3 70B Instruct
- Input
- $0.100/1M
- Output
- $0.320/1M
- Context
- 131.072K
| Model↕ | Input $/1M↑ | Output $/1M↕ | Context↕ | Cache Read |
|---|---|---|---|---|
Meta: Llama 3.1 8B Instruct Meta | $0.020 | $0.030 | 131.072K | — |
Mistral: Mistral Nemo mistralai | $0.020 | $0.030 | 131.072K | — |
Meta: Llama 3.2 1B Instruct Meta | $0.027 | $0.201 | 131.072K | — |
OpenAI: gpt-oss-20b openai | $0.029 | $0.140 | 131.072K | — |
OpenAI: gpt-oss-120b openai | $0.030 | $0.150 | 131.072K | — |
Cohere: Command R7B (12-2024) cohere | $0.037 | $0.150 | 128K | — |
Qwen: Qwen2.5 7B Instruct qwen | $0.040 | $0.100 | 131.072K | — |
Qwen: Qwen3 30B A3B Instruct 2507 qwen | $0.048 | $0.193 | 131.072K | — |
OpenAI: GPT-5 Nano openai | $0.050 | $0.400 | 400K | $0.010 |
Qwen: Qwen3 8B qwen | $0.050 | $0.400 | 131.072K | $0.050 |
Google: Gemma 3 4B | $0.050 | $0.100 | 131.072K | — |
Google: Gemma 3 12B | $0.050 | $0.150 | 131.072K | — |
Mistral: Mistral Small 3 mistralai | $0.050 | $0.080 | 32.768K | — |
Meta: Llama 3.2 3B Instruct Meta | $0.051 | $0.335 | 131.072K | — |
Google: Gemma 4 26B A4B | $0.060 | $0.330 | 262.144K | — |
Z.ai: GLM 4.7 Flash z-ai | $0.060 | $0.400 | 202.752K | $0.010 |
Google: Gemma 3n 4B | $0.060 | $0.120 | 32.768K | — |
Qwen: Qwen3.5-Flash qwen | $0.065 | $0.260 | 1,000K | — |
Qwen: Qwen3 Coder 30B A3B Instruct qwen | $0.070 | $0.270 | 160K | — |
OpenAI: gpt-oss-safeguard-20b openai | $0.075 | $0.300 | 131.072K | $0.037 |
Mistral: Mistral Small 3.2 24B mistralai | $0.075 | $0.200 | 128K | — |
Qwen: Qwen3 VL 8B Instruct qwen | $0.080 | $0.500 | 256K | — |
Qwen: Qwen3 30B A3B Thinking 2507 qwen | $0.080 | $0.400 | 131.072K | $0.080 |
Qwen: Qwen3 32B qwen | $0.080 | $0.280 | 131.072K | — |
Google: Gemma 3 27B | $0.080 | $0.160 | 131.072K | — |
DeepSeek: DeepSeek V4 Flash deepseek | $0.090 | $0.180 | 1,048.576K | $0.020 |
Qwen: Qwen3 Next 80B A3B Instruct qwen | $0.090 | $1.100 | 262.144K | — |
Qwen: Qwen3 235B A22B Instruct 2507 qwen | $0.090 | $0.100 | 262.144K | — |
Qwen: Qwen3 Next 80B A3B Thinking qwen | $0.098 | $0.780 | 262.144K | — |
Qwen: Qwen3.5-9B qwen | $0.100 | $0.150 | 262.144K | — |
Mistral: Ministral 3 3B 2512 mistralai | $0.100 | $0.100 | 131.072K | $0.010 |
Mistral: Voxtral Small 24B 2507 mistralai | $0.100 | $0.300 | 32K | $0.010 |
Google: Gemini 2.5 Flash Lite Preview 09-2025 | $0.100 | $0.400 | 1,048.576K | $0.010 |
Qwen: Qwen3 235B A22B Thinking 2507 qwen | $0.100 | $0.100 | 262.144K | $0.100 |
Google: Gemini 2.5 Flash Lite | $0.100 | $0.400 | 1,048.576K | $0.010 |
Qwen: Qwen3 14B qwen | $0.100 | $0.240 | 131.702K | — |
OpenAI: GPT-4.1 Nano openai | $0.100 | $0.400 | 1,047.576K | $0.025 |
Meta: Llama 4 Scout Meta | $0.100 | $0.300 | 10,000K | — |
Meta: Llama 3.3 70B Instruct Meta | $0.100 | $0.320 | 131.072K | — |
Qwen: Qwen3 VL 32B Instruct qwen | $0.104 | $0.416 | 262.144K | — |
Qwen: Qwen3 Coder Next qwen | $0.110 | $0.800 | 262.144K | $0.070 |
Qwen: Qwen3 VL 8B Thinking qwen | $0.117 | $1.365 | 256K | — |
Google: Gemma 4 31B | $0.120 | $0.350 | 262.144K | $0.090 |
Qwen: Qwen3 30B A3B qwen | $0.120 | $0.500 | 131.072K | — |
Qwen: Qwen3 VL 30B A3B Thinking qwen | $0.130 | $1.560 | 131.072K | — |
Qwen: Qwen3 VL 30B A3B Instruct qwen | $0.130 | $0.520 | 262.144K | — |
Z.ai: GLM 4.5 Air z-ai | $0.130 | $0.850 | 131.072K | $0.025 |
Qwen: Qwen3.6 35B A3B qwen | $0.140 | $1.000 | 262.144K | — |
Qwen: Qwen3.5-35B-A3B qwen | $0.140 | $1.000 | 262.144K | $0.050 |
Meta: Llama 3 8B Instruct Meta | $0.140 | $0.140 | 8.192K | — |
Mistral: Mistral Small 4 mistralai | $0.150 | $0.600 | 262.144K | $0.015 |
Mistral: Ministral 3 8B 2512 mistralai | $0.150 | $0.150 | 262.144K | $0.015 |
Meta: Llama 4 Maverick Meta | $0.150 | $0.600 | 1,048.576K | — |
OpenAI: GPT-4o-mini Search Preview openai | $0.150 | $0.600 | 128K | — |
Cohere: Command R (08-2024) cohere | $0.150 | $0.600 | 128K | — |
OpenAI: GPT-4o-mini openai | $0.150 | $0.600 | 128K | $0.075 |
OpenAI: GPT-4o-mini (2024-07-18) openai | $0.150 | $0.600 | 128K | $0.075 |
Meta: Llama Guard 4 12B Meta | $0.180 | $0.180 | 163.84K | — |
Qwen: Qwen3.6 Flash qwen | $0.188 | $1.125 | 1,000K | — |
Qwen: Qwen3.5-27B qwen | $0.195 | $1.560 | 262.144K | — |
Qwen: Qwen3 Coder Flash qwen | $0.195 | $0.975 | 1,000K | $0.039 |
OpenAI: GPT-5.4 Nano openai | $0.200 | $1.250 | 400K | $0.020 |
Mistral: Ministral 3 14B 2512 mistralai | $0.200 | $0.200 | 262.144K | $0.020 |
Qwen: Qwen3 VL 235B A22B Instruct qwen | $0.200 | $0.880 | 262.144K | $0.110 |
DeepSeek: DeepSeek V3 0324 deepseek | $0.200 | $0.770 | 163.84K | $0.135 |
Mistral: Saba mistralai | $0.200 | $0.600 | 32.768K | $0.020 |
DeepSeek: DeepSeek V3 deepseek | $0.200 | $0.800 | 131.072K | — |
DeepSeek: DeepSeek V3.1 deepseek | $0.210 | $0.790 | 163.84K | $0.130 |
Qwen: Qwen3 Coder 480B A35B qwen | $0.220 | $1.800 | 1,048.576K | — |
DeepSeek: DeepSeek V3.2 deepseek | $0.229 | $0.343 | 131.072K | $0.023 |
Google: Gemini 3.1 Flash Lite | $0.250 | $1.500 | 1,048.576K | $0.025 |
Google: Gemini 3.1 Flash Lite Preview | $0.250 | $1.500 | 1,048.576K | $0.025 |
OpenAI: GPT-5.1-Codex-Mini openai | $0.250 | $2.000 | 400K | $0.025 |
OpenAI: GPT-5 Mini openai | $0.250 | $2.000 | 400K | $0.025 |
Anthropic: Claude 3 Haiku anthropic | $0.250 | $1.250 | 200K | $0.030 |
Qwen: Qwen3.5-122B-A10B qwen | $0.260 | $2.080 | 262.144K | — |
Qwen: Qwen3.5 Plus 2026-02-15 qwen | $0.260 | $1.560 | 1,000K | — |
Qwen: Qwen3 VL 235B A22B Thinking qwen | $0.260 | $2.600 | 131.072K | — |
Qwen: Qwen Plus 0728 (thinking) qwen | $0.260 | $0.780 | 1,000K | — |
Qwen: Qwen Plus 0728 qwen | $0.260 | $0.780 | 1,000K | — |
Qwen: Qwen-Plus qwen | $0.260 | $0.780 | 1,000K | $0.052 |
DeepSeek: DeepSeek V3.2 Exp deepseek | $0.270 | $0.410 | 163.84K | — |
DeepSeek: DeepSeek V3.1 Terminus deepseek | $0.270 | $0.950 | 163.84K | $0.130 |
Qwen: Qwen3.6 27B qwen | $0.289 | $2.650 | 262.144K | — |
Qwen: Qwen3.5 Plus 2026-04-20 qwen | $0.300 | $1.800 | 1,000K | — |
Z.ai: GLM 4.6V z-ai | $0.300 | $0.900 | 131.072K | $0.055 |
Google: Nano Banana (Gemini 2.5 Flash Image) | $0.300 | $2.500 | 32.768K | $0.030 |
Mistral: Codestral 2508 mistralai | $0.300 | $0.900 | 256K | $0.030 |
Google: Gemini 2.5 Flash | $0.300 | $2.500 | 1,048.576K | $0.030 |
Qwen: Qwen3.7 Plus qwen | $0.320 | $1.280 | 1,000K | $0.064 |
Qwen: Qwen3.6 Plus qwen | $0.325 | $1.950 | 1,000K | — |
Meta: Llama 3.2 11B Vision Instruct Meta | $0.345 | $0.345 | 131.072K | — |
Mistral: Mistral Small 3.1 24B mistralai | $0.351 | $0.555 | 128K | — |
Qwen2.5 72B Instruct qwen | $0.360 | $0.400 | 131.072K | — |
Qwen: Qwen3.5 397B A17B qwen | $0.385 | $2.450 | 256K | — |
Z.ai: GLM 4.7 z-ai | $0.400 | $1.750 | 202.752K | $0.080 |
Mistral: Devstral 2 2512 mistralai | $0.400 | $2.000 | 262.144K | $0.040 |
Mistral: Mistral Medium 3.1 mistralai | $0.400 | $2.000 | 131.072K | $0.040 |
Mistral: Mistral Medium 3 mistralai | $0.400 | $2.000 | 131.072K | $0.040 |
OpenAI: GPT-4.1 Mini openai | $0.400 | $1.600 | 1,047.576K | $0.100 |
Meta: Llama 3.1 70B Instruct Meta | $0.400 | $0.400 | 131.072K | — |
Z.ai: GLM 4.6 z-ai | $0.430 | $1.740 | 202.752K | $0.080 |
DeepSeek: DeepSeek V4 Pro deepseek | $0.435 | $0.870 | 1,048.576K | $0.004 |
Qwen: Qwen3 235B A22B qwen | $0.455 | $1.820 | 131.072K | — |
Google: Nano Banana 2 (Gemini 3.1 Flash Image) | $0.500 | $3.000 | 131.072K | — |
Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview) | $0.500 | $3.000 | 131.072K | — |
Google: Gemini 3 Flash Preview | $0.500 | $3.000 | 1,048.576K | $0.050 |
Mistral: Mistral Large 3 2512 mistralai | $0.500 | $1.500 | 262.144K | $0.050 |
DeepSeek: R1 0528 deepseek | $0.500 | $2.150 | 163.84K | $0.350 |
OpenAI: GPT-3.5 Turbo openai | $0.500 | $1.500 | 16.385K | — |
Z.ai: GLM 5 z-ai | $0.600 | $1.920 | 202.752K | $0.120 |
OpenAI: GPT Audio Mini openai | $0.600 | $2.400 | 128K | — |
Z.ai: GLM 4.5V z-ai | $0.600 | $1.800 | 65.536K | $0.110 |
Z.ai: GLM 4.5 z-ai | $0.600 | $2.200 | 131.072K | $0.110 |
Qwen: Qwen3 Coder Plus qwen | $0.650 | $3.250 | 1,000K | $0.130 |
Google: Gemma 2 27B | $0.650 | $0.650 | 8.192K | — |
Qwen2.5 Coder 32B Instruct qwen | $0.660 | $1.000 | 128K | — |
DeepSeek: R1 deepseek | $0.700 | $2.500 | 163.84K | — |
OpenAI: GPT-5.4 Mini openai | $0.750 | $4.500 | 400K | $0.075 |
Qwen: Qwen3 Max Thinking qwen | $0.780 | $3.900 | 262.144K | — |
Qwen: Qwen3 Max qwen | $0.780 | $3.900 | 262.144K | $0.156 |
Qwen: Qwen2.5 VL 72B Instruct qwen | $0.800 | $1.000 | 131.072K | $0.400 |
DeepSeek: R1 Distill Llama 70B deepseek | $0.800 | $0.800 | 128K | — |
Z.ai: GLM 5.2 z-ai | $0.950 | $3.000 | 1,048.576K | $0.180 |
Z.ai: GLM 5.1 z-ai | $0.980 | $3.080 | 202.752K | $0.182 |
Anthropic: Claude Haiku 4.5 anthropic | $1.000 | $5.000 | 200K | $0.100 |
Perplexity: Sonar perplexity | $1.000 | $1.000 | 127.072K | — |
OpenAI: GPT-3.5 Turbo (older v0613) openai | $1.000 | $2.000 | 4.095K | — |
Qwen: Qwen3.6 Max Preview qwen | $1.040 | $6.240 | 262.144K | — |
OpenAI: o4 Mini High openai | $1.100 | $4.400 | 200K | $0.275 |
OpenAI: o4 Mini openai | $1.100 | $4.400 | 200K | $0.275 |
OpenAI: o3 Mini High openai | $1.100 | $4.400 | 200K | $0.550 |
OpenAI: o3 Mini openai | $1.100 | $4.400 | 200K | $0.550 |
Z.ai: GLM 5V Turbo z-ai | $1.200 | $4.000 | 202.752K | $0.240 |
Z.ai: GLM 5 Turbo z-ai | $1.200 | $4.000 | 262.144K | $0.240 |
Qwen: Qwen3.7 Max qwen | $1.250 | $3.750 | 1,000K | $0.250 |
OpenAI: GPT-5.1-Codex-Max openai | $1.250 | $10.000 | 400K | $0.125 |
OpenAI: GPT-5.1 openai | $1.250 | $10.000 | 400K | $0.130 |
OpenAI: GPT-5.1 Chat openai | $1.250 | $10.000 | 128K | $0.130 |
OpenAI: GPT-5.1-Codex openai | $1.250 | $10.000 | 400K | $0.130 |
OpenAI: GPT-5 Codex openai | $1.250 | $10.000 | 400K | $0.125 |
OpenAI: GPT-5 Chat openai | $1.250 | $10.000 | 128K | $0.125 |
OpenAI: GPT-5 openai | $1.250 | $10.000 | 400K | $0.125 |
Google: Gemini 2.5 Pro | $1.250 | $10.000 | 1,048.576K | $0.125 |
Google: Gemini 2.5 Pro Preview 06-05 | $1.250 | $10.000 | 1,048.576K | $0.125 |
Google: Gemini 2.5 Pro Preview 05-06 | $1.250 | $10.000 | 1,048.576K | $0.125 |
Google: Gemini 3.5 Flash | $1.500 | $9.000 | 1,048.576K | $0.150 |
Mistral: Mistral Medium 3.5 mistralai | $1.500 | $7.500 | 262.144K | — |
OpenAI: GPT-3.5 Turbo Instruct openai | $1.500 | $2.000 | 4.095K | — |
OpenAI: GPT-5.3 Chat openai | $1.750 | $14.000 | 128K | $0.175 |
OpenAI: GPT-5.3-Codex openai | $1.750 | $14.000 | 400K | $0.175 |
OpenAI: GPT-5.2-Codex openai | $1.750 | $14.000 | 400K | $0.175 |
OpenAI: GPT-5.2 Chat openai | $1.750 | $14.000 | 128K | $0.175 |
OpenAI: GPT-5.2 openai | $1.750 | $14.000 | 400K | $0.175 |
Google: Nano Banana Pro (Gemini 3 Pro Image) | $2.000 | $12.000 | 65.536K | $0.200 |
Google: Gemini 3.1 Pro Preview Custom Tools | $2.000 | $12.000 | 1,048.756K | $0.200 |
Google: Gemini 3.1 Pro Preview | $2.000 | $12.000 | 1,048.576K | $0.200 |
Google: Nano Banana Pro (Gemini 3 Pro Image Preview) | $2.000 | $12.000 | 65.536K | $0.200 |
OpenAI: o4 Mini Deep Research openai | $2.000 | $8.000 | 200K | $0.500 |
OpenAI: o3 openai | $2.000 | $8.000 | 200K | $0.500 |
OpenAI: GPT-4.1 openai | $2.000 | $8.000 | 1,047.576K | $0.500 |
Perplexity: Sonar Reasoning Pro perplexity | $2.000 | $8.000 | 128K | — |
Perplexity: Sonar Deep Research perplexity | $2.000 | $8.000 | 128K | — |
Mistral Large 2407 mistralai | $2.000 | $6.000 | 131.072K | $0.200 |
Mistral: Mixtral 8x22B Instruct mistralai | $2.000 | $6.000 | 65.536K | $0.200 |
Mistral Large mistralai | $2.000 | $6.000 | 128K | $0.200 |
OpenAI: GPT-5.4 openai | $2.500 | $15.000 | 1,050K | $0.250 |
OpenAI: GPT Audio openai | $2.500 | $10.000 | 128K | — |
OpenAI: GPT-5 Image Mini openai | $2.500 | $2.000 | 400K | $0.250 |
Cohere: Command A cohere | $2.500 | $10.000 | 256K | — |
OpenAI: GPT-4o Search Preview openai | $2.500 | $10.000 | 128K | — |
OpenAI: GPT-4o (2024-11-20) openai | $2.500 | $10.000 | 128K | $1.250 |
Cohere: Command R+ (08-2024) cohere | $2.500 | $10.000 | 128K | — |
OpenAI: GPT-4o (2024-08-06) openai | $2.500 | $10.000 | 128K | $1.250 |
OpenAI: GPT-4o openai | $2.500 | $10.000 | 128K | — |
Anthropic: Claude Sonnet 4.6 anthropic | $3.000 | $15.000 | 1,000K | $0.300 |
Perplexity: Sonar Pro Search perplexity | $3.000 | $15.000 | 200K | — |
Anthropic: Claude Sonnet 4.5 anthropic | $3.000 | $15.000 | 1,000K | $0.300 |
Anthropic: Claude Sonnet 4 anthropic | $3.000 | $15.000 | 1,000K | $0.300 |
Perplexity: Sonar Pro perplexity | $3.000 | $15.000 | 200K | — |
OpenAI: GPT-3.5 Turbo 16k openai | $3.000 | $4.000 | 16.385K | — |
Anthropic: Claude Opus 4.8 anthropic | $5.000 | $25.000 | 1,000K | $0.500 |
OpenAI: GPT Chat Latest openai | $5.000 | $30.000 | 400K | $0.500 |
OpenAI: GPT-5.5 openai | $5.000 | $30.000 | 1,050K | $0.500 |
Anthropic: Claude Opus 4.7 anthropic | $5.000 | $25.000 | 1,000K | $0.500 |
Anthropic: Claude Opus 4.6 anthropic | $5.000 | $25.000 | 1,000K | $0.500 |
Anthropic: Claude Opus 4.5 anthropic | $5.000 | $25.000 | 200K | $0.500 |
OpenAI: GPT-4o (2024-05-13) openai | $5.000 | $15.000 | 128K | — |
OpenAI: GPT-5.4 Image 2 openai | $8.000 | $15.000 | 272K | $2.000 |
Anthropic: Claude Fable 5 anthropic | $10.000 | $50.000 | 1,000K | $1.000 |
Anthropic: Claude Opus 4.8 (Fast) anthropic | $10.000 | $50.000 | 1,000K | $1.000 |
OpenAI: GPT-5 Image openai | $10.000 | $10.000 | 400K | $1.250 |
OpenAI: o3 Deep Research openai | $10.000 | $40.000 | 200K | $2.500 |
OpenAI: GPT-4 Turbo openai | $10.000 | $30.000 | 128K | — |
OpenAI: GPT-4 Turbo Preview openai | $10.000 | $30.000 | 128K | — |
OpenAI: GPT-5 Pro openai | $15.000 | $120.000 | 400K | — |
Anthropic: Claude Opus 4.1 anthropic | $15.000 | $75.000 | 200K | $1.500 |
Anthropic: Claude Opus 4 anthropic | $15.000 | $75.000 | 200K | $1.500 |
OpenAI: o1 openai | $15.000 | $60.000 | 200K | $7.500 |
OpenAI: o3 Pro openai | $20.000 | $80.000 | 200K | — |
OpenAI: GPT-5.2 Pro openai | $21.000 | $168.000 | 400K | — |
Anthropic: Claude Opus 4.7 (Fast) anthropic | $30.000 | $150.000 | 1,000K | $3.000 |
OpenAI: GPT-5.5 Pro openai | $30.000 | $180.000 | 1,050K | — |
Anthropic: Claude Opus 4.6 (Fast) anthropic | $30.000 | $150.000 | 1,000K | $3.000 |
OpenAI: GPT-5.4 Pro openai | $30.000 | $180.000 | 1,050K | — |
OpenAI: GPT-4 openai | $30.000 | $60.000 | 8.191K | — |
Turn these prices into a margin forecast
Plug any model's price into the Margin Simulator to see gross margin, cost per user, and breakeven MAU for your product — or compare two models head-to-head.
LLM API pricing — frequently asked questions
How much does GPT-4o mini cost?
GPT-4o mini is priced at $0.15 per million input tokens and $0.60 per million output tokens via the OpenAI API (routed through OpenRouter). It is OpenAI's most cost-efficient hosted model and is optimised for high-volume, latency-sensitive workloads.
What is the cheapest LLM API available in 2026?
Open-source models hosted via OpenRouter — including Llama 3.1 8B, Gemini 2.0 Flash, and DeepSeek V3 — are among the cheapest options, often under $0.05 per million input tokens. For hosted frontier budget models, GPT-4o mini, Claude Haiku 4.5, and Gemini 2.5 Flash lead the pack.
How does Claude Haiku 4.5 pricing compare to GPT-4o mini?
Both sit in the budget tier of their respective providers and are comparably priced. The cheaper model for your workload depends on your input/output token ratio — use the table and the compare calculator to find out which wins for your specific usage.
Why do input and output token prices differ?
Generating (output) tokens is computationally more expensive than reading (input) tokens. Most models price output at 2–4× the input rate. For output-heavy workloads like chat or code generation, the output price dominates your bill.
What is a token in LLM pricing?
A token is roughly 0.75 English words, or approximately 4 characters. A 1,000-word document contains around 1,333 tokens. All prices in this table are quoted per million tokens ($/1M) to allow direct comparison across models.
What is prompt caching and how does it affect cost?
Prompt caching lets you reuse expensive input context across multiple calls at a reduced rate — often 80–90% cheaper than a fresh input read. Models that support it (Claude, Gemini 2.5) charge a slightly higher cache-write rate on first use, then discounted cache-read rates on subsequent calls. The 'Cache Read' column in the table above shows where it's available.
LLM token prices change frequently — OpenAI, Anthropic, and Google all adjust rates as competition intensifies. This table reflects live prices from OpenRouter and is updated on every deploy. For per-model cost forecasting, use the Margin Simulator to model gross margin at your MAU.
Related tools: Compare two LLM models side-by-side · LLM cost per user calculator · How to calculate LLM cost per user