Skip to content
ModelPriceLab

Search model prices

Find source-linked rates by model, vendor, or platform.

NVIDIA API Pricing

Complete pricing for all NVIDIA models. Input/output costs per 1M tokens, context windows, and rate limits.

Billing terms: priced in USD per 1M tokens unless a row states otherwise; official direct rates are listed separately from third-party channels, and regional prices are kept distinct. Record updated 2026-10-11.

NVIDIA API Pricing Calculator

Select a model to estimate monthly cost using its listed rates on each channel.

Quick examples:
Input mode
Cache hit rate (share of input tokens served from cache)50%

Workloads reusing system prompts, long documents, or conversation history hit more; channels without a cached rate are billed at the full input rate.

Estimated monthly cost at current inputs
ChannelEffective inputOutput /1MEst. / month
official$0.0000$0.0000$0.0000

Formula: monthly cost = (input tokens × effective input rate + output tokens × output rate) ÷ 1,000,000 × calls. The effective input rate blends list and cached rates at your cache-hit share. Rates share the selected model and billing basis; your vendor bill is authoritative.

Currently available rates

Active Speaker Detection

LLMnvidia-nvidia-active-speaker-detection
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
—
4K
Sources: Official

BGE M3

Embednvidia-baai-bge-m3
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
8K
1K
Sources: Official

ByteDance-Seed/Seed-OSS-36B-Instruct

LLMnvidia-bytedance-seed-oss-36b-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
262K
262K
Sources: Official

Cosmos Reason2 8B

LLMnvidia-nvidia-cosmos-reason2-8b
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
131K
16K
Sources: Official

DeepSeek V4 Flash

LLMnvidia-deepseek-ai-deepseek-v4-flash
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.14
$0.28
$0.0028
1049K
393K
Sources: Official

DeepSeek V4 Flash 0731

LLMnvidia-deepseek-ai-deepseek-v4-flash-0731
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
1000K
384K
Sources: Official

DeepSeek V4 Pro

LLMnvidia-deepseek-ai-deepseek-v4-pro
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.44
$0.87
$0.0036
1049K
393K
Sources: Official

DeepSeek V4 Pro 0813

LLMnvidia-deepseek-ai-deepseek-v4-pro-0813
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
1000K
384K
Sources: Official

DeepSeek V4.1 Flash

LLMnvidia-deepseek-ai-deepseek-v4.1-flash
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
1000K
384K
Sources: Official

DiffusionGemma 26B A4B IT

LLMnvidia-google-diffusiongemma-26b-a4b-it
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
250K
33K
Sources: Official

FLUX.1-Kontext-dev

Image Gennvidia-black-forest-labs-flux_1-kontext-dev
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
41K
41K
Sources: Official

FLUX.1-dev

Image Gennvidia-black-forest-labs-flux.1-dev
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
4K
—
Sources: Official

FLUX.1-schnell

Image Gennvidia-black-forest-labs-flux_1-schnell
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
0K
—
Sources: Official

FLUX.2 Klein 4B

Image Gennvidia-black-forest-labs-flux_2-klein-4b
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
41K
41K
Sources: Official

GLM-5.2

LLMnvidia-z-ai-glm-5.2
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
1000K
131K
Sources: Official

GLM-5.3

LLMnvidia-z-ai-glm-5.3
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
1000K
131K
Sources: Official

GLM-5.3-Flash

LLMnvidia-z-ai-glm-5.3-flash
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
1000K
131K
Sources: Official

GPT OSS 20B

LLMnvidia-openai-gpt-oss-20b
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
131K
33K
Sources: Official

GPT-OSS-120B

LLMnvidia-openai-gpt-oss-120b
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K

Gemma 2 2b It

LLMnvidia-google-gemma-2-2b-it
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

Gemma 3 12B IT

LLMnvidia-google-gemma-3-12b-it
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
131K
16K
Sources: Official

Gemma 3 4B IT

LLMnvidia-google-gemma-3-4b-it
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
131K
16K
Sources: Official

Gemma 3n E2b It

LLMnvidia-google-gemma-3n-e2b-it
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

Gemma 3n E4b It

LLMnvidia-google-gemma-3n-e4b-it
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

Gemma-4-31B-IT

LLMnvidia-google-gemma-4-31b-it
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
256K
16K
Sources: Official

Inkling

LLMnvidia-thinkingmachines-inkling
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
1049K
16K
Sources: Official

Kimi K2 0905

LLMnvidia-moonshotai-kimi-k2-instruct-0905
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
262K
262K

Kimi K2.6

LLMnvidia-moonshotai-kimi-k2.6
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
262K
262K
Sources: Official

Kimi K3

LLMnvidia-moonshotai-kimi-k3
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
1049K
131K
Sources: Official

Laguna XS 2.1

LLMnvidia-poolside-laguna-xs-2.1
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
262K
16K
Sources: Official

Llama 3.1 70b Instruct

LLMnvidia-meta-llama-3.1-70b-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

Llama 3.1 8B Instruct

LLMnvidia-meta-llama-3.1-8b-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
16K
4K
Sources: Official

Llama 3.1 Nemotron 70B Instruct

LLMnvidia-nvidia-llama-3.1-nemotron-70b-instruct
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

Llama 3.1 Nemotron Nano 8B v1

LLMnvidia-nvidia-llama-3.1-nemotron-nano-8b-v1
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
131K
16K
Sources: Official

Llama 3.1 Nemotron Nano VL 8B v1

LLMnvidia-nvidia-llama-3.1-nemotron-nano-vl-8b-v1
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
33K
16K
Sources: Official

Llama 3.1 Nemotron Ultra 253B

LLMnvidia-nvidia-llama-3.1-nemotron-ultra-253b-v1
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
16K
Sources: Official

Llama 3.2 11b Vision Instruct

LLMnvidia-meta-llama-3.2-11b-vision-instruct
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

Llama 3.2 1b Instruct

LLMnvidia-meta-llama-3.2-1b-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

Llama 3.2 3B Instruct

LLMnvidia-meta-llama-3.2-3b-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
33K
32K
Sources: Official

Llama 3.3 70b Instruct

LLMnvidia-meta-llama-3.3-70b-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

Llama 3.3 Nemotron Super 49B v1

LLMnvidia-nvidia-llama-3.3-nemotron-super-49b-v1
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
131K
66K
Sources: Official

Llama 3.3 Nemotron Super 49B v1.5

LLMnvidia-nvidia-llama-3.3-nemotron-super-49b-v1.5
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
131K
66K
Sources: Official

Llama 4 Maverick 17b 128e Instruct

LLMnvidia-meta-llama-4-maverick-17b-128e-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

Llama Guard 4 12B

LLMnvidia-meta-llama-guard-4-12b
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
16K
Sources: Official

Llama-3.2-90B-Vision-Instruct

LLMnvidia-meta-llama-3.2-90b-vision-instruct
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

Magistral Small 2506

LLMnvidia-mistralai-magistral-small-2506
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
33K
33K
Sources: Official

MiniMax-M2.7

LLMnvidia-minimaxai-minimax-m2.7
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
205K
131K
Sources: Official

MiniMax-M3

LLMnvidia-minimaxai-minimax-m3
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
1000K
16K
Sources: Official

Ministral 3 14B Instruct 2512

LLMnvidia-mistralai-ministral-14b-instruct-2512
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
262K
16K
Sources: Official

Mistral Large 3 675B Instruct 2512

LLMnvidia-mistralai-mistral-large-3-675b-instruct-2512
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
262K
262K
Sources: Official

Mistral Medium 3

LLMnvidia-mistralai-mistral-medium-3-instruct
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
131K
33K
Sources: Official

Mistral Medium 3.5

LLMnvidia-mistralai-mistral-medium-3.5-128b
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
262K
33K
Sources: Official

Mistral-7B-Instruct-v0.3

LLMnvidia-mistralai-mistral-7b-instruct-v0.3
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
66K
66K
Sources: Official

Mistral: Mixtral 8x22B Instruct

LLMnvidia-mistralai-mixtral-8x22b-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
66K
13K
Sources: Official

Mistral: Mixtral 8x7B Instruct

LLMnvidia-mistralai-mixtral-8x7b-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
33K
16K
Sources: Official

Muse Glimmer 30B

LLMnvidia-meta-muse-glimmer-30b
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
131K
131K
Sources: Official

Nemotron 3 Nano Omni

LLMnvidia-nvidia-nemotron-3-nano-omni-30b-a3b-reasoning
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
256K
66K
Sources: Official

Nemotron 3 Nano Omni (free)

LLMnvidia-nemotron-3-nano-omni-30b-a3b-reasoning-free
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
OpenRouter
global
1M Tokens
$0.0000
$0.0000
—
256K
66K
Sources: OpenRouter

Nemotron 3 Super

LLMnvidia-nemotron-3-super-120b-a12b
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
OpenRouter
global
1M Tokens
$0.08
$0.45
—
262K
236K
Official
global
1M Tokens
$0.20
$0.80
—
262K
262K
Sources: OpenRouter · Official

Nemotron 3 Super (free)

LLMnvidia-nemotron-3-super-120b-a12b-free
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
OpenRouter
global
1M Tokens
$0.0000
$0.0000
—
262K
236K
Sources: OpenRouter

Nemotron 3 Ultra (free)

LLMnvidia-nemotron-3-ultra-550b-a55b-free
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
OpenRouter
global
1M Tokens
$0.0000
$0.0000
—
1000K
66K
Sources: OpenRouter

Nemotron 3 Ultra 550B A55B

LLMnvidia-nemotron-3-ultra-550b-a55b
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
OpenRouter
global
1M Tokens
$0.50
$2.20
$0.10
262K
16K
Official
global
1M Tokens
$0.50
$2.50
$0.15
1000K
66K
Sources: OpenRouter · Official

Nemotron 3.5 Content Safety

LLMnvidia-nemotron-3.5-content-safety
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
OpenRouter
global
1M Tokens
$0.20
$0.20
—
131K
118K
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: OpenRouter · Official

Nemotron 3.5 Content Safety (free)

LLMnvidia-nemotron-3.5-content-safety-free
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
OpenRouter
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: OpenRouter

Nemotron 3.5 Lightning

LLMnvidia-nemotron-3.5-lightning
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
OpenRouter
global
1M Tokens
$0.06
$0.16
$0.03
262K
33K
Sources: OpenRouter

Nemotron 3.5 Lightning (free)

LLMnvidia-nemotron-3.5-lightning-free
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
OpenRouter
global
1M Tokens
$0.0000
$0.0000
—
1000K
66K
Sources: OpenRouter

Nemotron 3.5 Lightning 30B A3B

LLMnvidia-nvidia-nemotron-3.5-lightning-30b-a3b
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
262K
262K
Sources: Official

Nemotron Nano 12B v2 VL

LLMnvidia-nvidia-nemotron-nano-12b-v2-vl
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
128K
Sources: Official

Phi 4 Multimodal

LLMnvidia-microsoft-phi-4-multimodal-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
16K
Sources: Official

Phi-4-Mini

LLMnvidia-microsoft-phi-4-mini-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
131K
8K
Sources: Official

Qwen Image

Image Gennvidia-qwen-qwen-image
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
—
—
Sources: Official

Qwen Image Edit

Image Gennvidia-qwen-qwen-image-edit
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
—
—
Sources: Official

Qwen2.5 Coder 32b Instruct

LLMnvidia-qwen-qwen2.5-coder-32b-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

Qwen3 Coder 480B A35B Instruct

LLMnvidia-qwen-qwen3-coder-480b-a35b-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
262K
67K
Sources: Official

Qwen3-Next-80B-A3B-Instruct

LLMnvidia-qwen-qwen3-next-80b-a3b-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
262K
16K
Sources: Official

Qwen3.5 122B-A10B

LLMnvidia-qwen-qwen3.5-122b-a10b
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
262K
66K
Sources: Official

Qwen3.5-397B-A17B

LLMnvidia-qwen-qwen3.5-397b-a17b
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
262K
8K
Sources: Official

Step 3.5 Flash

LLMnvidia-stepfun-ai-step-3.5-flash
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
256K
16K
Sources: Official

Step 3.7 Flash

LLMnvidia-stepfun-ai-step-3.7-flash
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
256K
16K
Sources: Official

Switchyard

LLMnvidia-switchyard
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
OpenRouter
global
1M Tokens
-
-
—
1000K
—
Sources: OpenRouter

Whisper Large v3

LLMnvidia-openai-whisper-large-v3
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
—
4K
Sources: Official

bevformer

Specializednvidia-nvidia-bevformer
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

cosmos-predict1-5b

Videonvidia-nvidia-cosmos-predict1-5b
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
—
4K
Sources: Official

cosmos-transfer1-7b

Videonvidia-nvidia-cosmos-transfer1-7b
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
—
4K
Sources: Official

cosmos-transfer2.5-2b

Videonvidia-nvidia-cosmos-transfer2_5-2b
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
—
4K
Sources: Official

dracarys-llama-3.1-70b-instruct

LLMnvidia-abacusai-dracarys-llama-3.1-70b-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

esm2-650m

Specializednvidia-meta-esm2-650m
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

esmfold

Specializednvidia-meta-esmfold
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

gliner-pii

Specializednvidia-nvidia-gliner-pii
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

llama-3.1-nemotron-safety-guard-8b-v3

LLMnvidia-nvidia-llama-3.1-nemotron-safety-guard-8b-v3
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

llama-3_2-nemoretriever-300m-embed-v1

Embednvidia-nvidia-llama-3_2-nemoretriever-300m-embed-v1
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
33K
2K
Sources: Official

llama-nemotron-embed-vl-1b-v2

Embednvidia-nvidia-llama-nemotron-embed-vl-1b-v2
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
33K
2K
Sources: Official

llama-nemotron-rerank-vl-1b-v2

Reranknvidia-nvidia-llama-nemotron-rerank-vl-1b-v2
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

magpie-tts-zeroshot

Audionvidia-nvidia-magpie-tts-zeroshot
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
—
4K
Sources: Official

mistral-nemotron

LLMnvidia-mistralai-mistral-nemotron
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

mistral-small-4-119b-2603

LLMnvidia-mistralai-mistral-small-4-119b-2603
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

nemotron-3-content-safety

LLMnvidia-nvidia-nemotron-3-content-safety
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

nemotron-3-nano-30b-a3b

LLMnvidia-nemotron-3-nano-30b-a3b
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
OpenRouter
global
1M Tokens
$0.05
$0.20
$0.03
262K
236K
Official
global
1M Tokens
$0.0000
$0.0000
—
131K
131K
Sources: OpenRouter · Official

nemotron-content-safety-reasoning-4b

LLMnvidia-nvidia-nemotron-content-safety-reasoning-4b
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

nemotron-mini-4b-instruct

LLMnvidia-nvidia-nemotron-mini-4b-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

nemotron-voicechat

LLMnvidia-nvidia-nemotron-voicechat
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

nv-embed-v1

Embednvidia-nvidia-nv-embed-v1
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
33K
2K
Sources: Official

nv-embedcode-7b-v1

Embednvidia-nvidia-nv-embedcode-7b-v1
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
33K
2K
Sources: Official

nvidia-nemotron-nano-9b-v2

LLMnvidia-nvidia-nvidia-nemotron-nano-9b-v2
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
131K
131K
Sources: Official

paligemma

LLMnvidia-google-google-paligemma
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

rerank-qa-mistral-4b

Reranknvidia-nvidia-rerank-qa-mistral-4b
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

riva-translate-4b-instruct-v1_1

LLMnvidia-nvidia-riva-translate-4b-instruct-v1.1
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

sarvam-m

LLMnvidia-sarvamai-sarvam-m
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

solar-10.7b-instruct

LLMnvidia-upstage-solar-10.7b-instruct
View Details
Deprecated

Deprecated endpoints may still be callable. Migrate before retirement.

Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

sparsedrive

Specializednvidia-nvidia-sparsedrive
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

streampetr

Specializednvidia-nvidia-streampetr
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

studiovoice

LLMnvidia-nvidia-studiovoice
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
8K
Sources: Official

synthetic-video-detector

LLMnvidia-nvidia-synthetic-video-detector
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
—
4K
Sources: Official

usdcode

LLMnvidia-nvidia-usdcode
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
128K
4K
Sources: Official

usdvalidate

LLMnvidia-nvidia-usdvalidate
View Details
Platform
Region
Input / 1M
Output / 1M
Cached input
Context
Max Output
Official
global
1M Tokens
$0.0000
$0.0000
—
—
4K
Sources: Official
AI Cost Optimization

Does this change your best option?

Check your models and usage for free. See the relevant official changes, estimated cost, and a practical next step.