NVIDIA API Pricing
Complete pricing for all NVIDIA models. Input/output costs per 1M tokens, context windows, and rate limits.
Billing terms: priced in USD per 1M tokens unless a row states otherwise; official direct rates are listed separately from third-party channels, and regional prices are kept distinct. Record updated 2026-10-11.
NVIDIA API Pricing Calculator
Select a model to estimate monthly cost using its listed rates on each channel.
Workloads reusing system prompts, long documents, or conversation history hit more; channels without a cached rate are billed at the full input rate.
| Channel | Effective input | Output /1M | Est. / month |
|---|---|---|---|
| official | $0.0000 | $0.0000 | $0.0000 |
Formula: monthly cost = (input tokens × effective input rate + output tokens × output rate) ÷ 1,000,000 × calls. The effective input rate blends list and cached rates at your cache-hit share. Rates share the selected model and billing basis; your vendor bill is authoritative.
Currently available rates
Active Speaker Detection
LLMnvidia-nvidia-active-speaker-detectionBGE M3
Embednvidia-baai-bge-m3ByteDance-Seed/Seed-OSS-36B-Instruct
LLMnvidia-bytedance-seed-oss-36b-instructDeprecated endpoints may still be callable. Migrate before retirement.
Cosmos Reason2 8B
LLMnvidia-nvidia-cosmos-reason2-8bDeepSeek V4 Flash
LLMnvidia-deepseek-ai-deepseek-v4-flashDeprecated endpoints may still be callable. Migrate before retirement.
DeepSeek V4 Flash 0731
LLMnvidia-deepseek-ai-deepseek-v4-flash-0731Deprecated endpoints may still be callable. Migrate before retirement.
DeepSeek V4 Pro
LLMnvidia-deepseek-ai-deepseek-v4-proDeprecated endpoints may still be callable. Migrate before retirement.
DeepSeek V4 Pro 0813
LLMnvidia-deepseek-ai-deepseek-v4-pro-0813Deprecated endpoints may still be callable. Migrate before retirement.
DeepSeek V4.1 Flash
LLMnvidia-deepseek-ai-deepseek-v4.1-flashDiffusionGemma 26B A4B IT
LLMnvidia-google-diffusiongemma-26b-a4b-itFLUX.1-Kontext-dev
Image Gennvidia-black-forest-labs-flux_1-kontext-devFLUX.1-dev
Image Gennvidia-black-forest-labs-flux.1-devFLUX.1-schnell
Image Gennvidia-black-forest-labs-flux_1-schnellFLUX.2 Klein 4B
Image Gennvidia-black-forest-labs-flux_2-klein-4bGLM-5.2
LLMnvidia-z-ai-glm-5.2Deprecated endpoints may still be callable. Migrate before retirement.
GLM-5.3
LLMnvidia-z-ai-glm-5.3GLM-5.3-Flash
LLMnvidia-z-ai-glm-5.3-flashGPT OSS 20B
LLMnvidia-openai-gpt-oss-20bGPT-OSS-120B
LLMnvidia-openai-gpt-oss-120bGemma 2 2b It
LLMnvidia-google-gemma-2-2b-itDeprecated endpoints may still be callable. Migrate before retirement.
Gemma 3 12B IT
LLMnvidia-google-gemma-3-12b-itGemma 3 4B IT
LLMnvidia-google-gemma-3-4b-itGemma 3n E2b It
LLMnvidia-google-gemma-3n-e2b-itDeprecated endpoints may still be callable. Migrate before retirement.
Gemma 3n E4b It
LLMnvidia-google-gemma-3n-e4b-itDeprecated endpoints may still be callable. Migrate before retirement.
Gemma-4-31B-IT
LLMnvidia-google-gemma-4-31b-itInkling
LLMnvidia-thinkingmachines-inklingDeprecated endpoints may still be callable. Migrate before retirement.
Kimi K2 0905
LLMnvidia-moonshotai-kimi-k2-instruct-0905Kimi K2.6
LLMnvidia-moonshotai-kimi-k2.6Deprecated endpoints may still be callable. Migrate before retirement.
Kimi K3
LLMnvidia-moonshotai-kimi-k3Laguna XS 2.1
LLMnvidia-poolside-laguna-xs-2.1Llama 3.1 70b Instruct
LLMnvidia-meta-llama-3.1-70b-instructDeprecated endpoints may still be callable. Migrate before retirement.
Llama 3.1 8B Instruct
LLMnvidia-meta-llama-3.1-8b-instructDeprecated endpoints may still be callable. Migrate before retirement.
Llama 3.1 Nemotron 70B Instruct
LLMnvidia-nvidia-llama-3.1-nemotron-70b-instructLlama 3.1 Nemotron Nano 8B v1
LLMnvidia-nvidia-llama-3.1-nemotron-nano-8b-v1Deprecated endpoints may still be callable. Migrate before retirement.
Llama 3.1 Nemotron Nano VL 8B v1
LLMnvidia-nvidia-llama-3.1-nemotron-nano-vl-8b-v1Deprecated endpoints may still be callable. Migrate before retirement.
Llama 3.1 Nemotron Ultra 253B
LLMnvidia-nvidia-llama-3.1-nemotron-ultra-253b-v1Llama 3.2 11b Vision Instruct
LLMnvidia-meta-llama-3.2-11b-vision-instructLlama 3.2 1b Instruct
LLMnvidia-meta-llama-3.2-1b-instructDeprecated endpoints may still be callable. Migrate before retirement.
Llama 3.2 3B Instruct
LLMnvidia-meta-llama-3.2-3b-instructDeprecated endpoints may still be callable. Migrate before retirement.
Llama 3.3 70b Instruct
LLMnvidia-meta-llama-3.3-70b-instructDeprecated endpoints may still be callable. Migrate before retirement.
Llama 3.3 Nemotron Super 49B v1
LLMnvidia-nvidia-llama-3.3-nemotron-super-49b-v1Deprecated endpoints may still be callable. Migrate before retirement.
Llama 3.3 Nemotron Super 49B v1.5
LLMnvidia-nvidia-llama-3.3-nemotron-super-49b-v1.5Deprecated endpoints may still be callable. Migrate before retirement.
Llama 4 Maverick 17b 128e Instruct
LLMnvidia-meta-llama-4-maverick-17b-128e-instructDeprecated endpoints may still be callable. Migrate before retirement.
Llama Guard 4 12B
LLMnvidia-meta-llama-guard-4-12bLlama-3.2-90B-Vision-Instruct
LLMnvidia-meta-llama-3.2-90b-vision-instructMagistral Small 2506
LLMnvidia-mistralai-magistral-small-2506MiniMax-M2.7
LLMnvidia-minimaxai-minimax-m2.7Deprecated endpoints may still be callable. Migrate before retirement.
MiniMax-M3
LLMnvidia-minimaxai-minimax-m3Deprecated endpoints may still be callable. Migrate before retirement.
Ministral 3 14B Instruct 2512
LLMnvidia-mistralai-ministral-14b-instruct-2512Deprecated endpoints may still be callable. Migrate before retirement.
Mistral Large 3 675B Instruct 2512
LLMnvidia-mistralai-mistral-large-3-675b-instruct-2512Deprecated endpoints may still be callable. Migrate before retirement.
Mistral Medium 3
LLMnvidia-mistralai-mistral-medium-3-instructMistral Medium 3.5
LLMnvidia-mistralai-mistral-medium-3.5-128bDeprecated endpoints may still be callable. Migrate before retirement.
Mistral-7B-Instruct-v0.3
LLMnvidia-mistralai-mistral-7b-instruct-v0.3Mistral: Mixtral 8x22B Instruct
LLMnvidia-mistralai-mixtral-8x22b-instructDeprecated endpoints may still be callable. Migrate before retirement.
Mistral: Mixtral 8x7B Instruct
LLMnvidia-mistralai-mixtral-8x7b-instructDeprecated endpoints may still be callable. Migrate before retirement.
Muse Glimmer 30B
LLMnvidia-meta-muse-glimmer-30bNemotron 3 Nano Omni
LLMnvidia-nvidia-nemotron-3-nano-omni-30b-a3b-reasoningNemotron 3 Nano Omni (free)
LLMnvidia-nemotron-3-nano-omni-30b-a3b-reasoning-freeNemotron 3 Super
LLMnvidia-nemotron-3-super-120b-a12bNemotron 3 Super (free)
LLMnvidia-nemotron-3-super-120b-a12b-freeNemotron 3 Ultra (free)
LLMnvidia-nemotron-3-ultra-550b-a55b-freeNemotron 3 Ultra 550B A55B
LLMnvidia-nemotron-3-ultra-550b-a55bNemotron 3.5 Content Safety
LLMnvidia-nemotron-3.5-content-safetyNemotron 3.5 Content Safety (free)
LLMnvidia-nemotron-3.5-content-safety-freeNemotron 3.5 Lightning
LLMnvidia-nemotron-3.5-lightningNemotron 3.5 Lightning (free)
LLMnvidia-nemotron-3.5-lightning-freeNemotron 3.5 Lightning 30B A3B
LLMnvidia-nvidia-nemotron-3.5-lightning-30b-a3bNemotron Nano 12B v2 VL
LLMnvidia-nvidia-nemotron-nano-12b-v2-vlDeprecated endpoints may still be callable. Migrate before retirement.
Phi 4 Multimodal
LLMnvidia-microsoft-phi-4-multimodal-instructDeprecated endpoints may still be callable. Migrate before retirement.
Phi-4-Mini
LLMnvidia-microsoft-phi-4-mini-instructDeprecated endpoints may still be callable. Migrate before retirement.
Qwen Image
Image Gennvidia-qwen-qwen-imageQwen Image Edit
Image Gennvidia-qwen-qwen-image-editQwen2.5 Coder 32b Instruct
LLMnvidia-qwen-qwen2.5-coder-32b-instructDeprecated endpoints may still be callable. Migrate before retirement.
Qwen3 Coder 480B A35B Instruct
LLMnvidia-qwen-qwen3-coder-480b-a35b-instructDeprecated endpoints may still be callable. Migrate before retirement.
Qwen3-Next-80B-A3B-Instruct
LLMnvidia-qwen-qwen3-next-80b-a3b-instructDeprecated endpoints may still be callable. Migrate before retirement.
Qwen3.5 122B-A10B
LLMnvidia-qwen-qwen3.5-122b-a10bDeprecated endpoints may still be callable. Migrate before retirement.
Qwen3.5-397B-A17B
LLMnvidia-qwen-qwen3.5-397b-a17bDeprecated endpoints may still be callable. Migrate before retirement.
Step 3.5 Flash
LLMnvidia-stepfun-ai-step-3.5-flashDeprecated endpoints may still be callable. Migrate before retirement.
Step 3.7 Flash
LLMnvidia-stepfun-ai-step-3.7-flashDeprecated endpoints may still be callable. Migrate before retirement.
Switchyard
LLMnvidia-switchyardWhisper Large v3
LLMnvidia-openai-whisper-large-v3bevformer
Specializednvidia-nvidia-bevformercosmos-predict1-5b
Videonvidia-nvidia-cosmos-predict1-5bcosmos-transfer1-7b
Videonvidia-nvidia-cosmos-transfer1-7bcosmos-transfer2.5-2b
Videonvidia-nvidia-cosmos-transfer2_5-2bdracarys-llama-3.1-70b-instruct
LLMnvidia-abacusai-dracarys-llama-3.1-70b-instructDeprecated endpoints may still be callable. Migrate before retirement.
esm2-650m
Specializednvidia-meta-esm2-650mesmfold
Specializednvidia-meta-esmfoldgliner-pii
Specializednvidia-nvidia-gliner-piiDeprecated endpoints may still be callable. Migrate before retirement.
llama-3.1-nemotron-safety-guard-8b-v3
LLMnvidia-nvidia-llama-3.1-nemotron-safety-guard-8b-v3llama-3_2-nemoretriever-300m-embed-v1
Embednvidia-nvidia-llama-3_2-nemoretriever-300m-embed-v1llama-nemotron-embed-vl-1b-v2
Embednvidia-nvidia-llama-nemotron-embed-vl-1b-v2llama-nemotron-rerank-vl-1b-v2
Reranknvidia-nvidia-llama-nemotron-rerank-vl-1b-v2magpie-tts-zeroshot
Audionvidia-nvidia-magpie-tts-zeroshotmistral-nemotron
LLMnvidia-mistralai-mistral-nemotronDeprecated endpoints may still be callable. Migrate before retirement.
mistral-small-4-119b-2603
LLMnvidia-mistralai-mistral-small-4-119b-2603Deprecated endpoints may still be callable. Migrate before retirement.
nemotron-3-content-safety
LLMnvidia-nvidia-nemotron-3-content-safetyDeprecated endpoints may still be callable. Migrate before retirement.
nemotron-3-nano-30b-a3b
LLMnvidia-nemotron-3-nano-30b-a3bDeprecated endpoints may still be callable. Migrate before retirement.
nemotron-content-safety-reasoning-4b
LLMnvidia-nvidia-nemotron-content-safety-reasoning-4bDeprecated endpoints may still be callable. Migrate before retirement.
nemotron-mini-4b-instruct
LLMnvidia-nvidia-nemotron-mini-4b-instructDeprecated endpoints may still be callable. Migrate before retirement.
nemotron-voicechat
LLMnvidia-nvidia-nemotron-voicechatnv-embed-v1
Embednvidia-nvidia-nv-embed-v1nv-embedcode-7b-v1
Embednvidia-nvidia-nv-embedcode-7b-v1nvidia-nemotron-nano-9b-v2
LLMnvidia-nvidia-nvidia-nemotron-nano-9b-v2Deprecated endpoints may still be callable. Migrate before retirement.
paligemma
LLMnvidia-google-google-paligemmarerank-qa-mistral-4b
Reranknvidia-nvidia-rerank-qa-mistral-4briva-translate-4b-instruct-v1_1
LLMnvidia-nvidia-riva-translate-4b-instruct-v1.1sarvam-m
LLMnvidia-sarvamai-sarvam-mDeprecated endpoints may still be callable. Migrate before retirement.
solar-10.7b-instruct
LLMnvidia-upstage-solar-10.7b-instructDeprecated endpoints may still be callable. Migrate before retirement.
sparsedrive
Specializednvidia-nvidia-sparsedrivestreampetr
Specializednvidia-nvidia-streampetrstudiovoice
LLMnvidia-nvidia-studiovoicesynthetic-video-detector
LLMnvidia-nvidia-synthetic-video-detectorusdcode
LLMnvidia-nvidia-usdcodeusdvalidate
LLMnvidia-nvidia-usdvalidateDoes this change your best option?
Check your models and usage for free. See the relevant official changes, estimated cost, and a practical next step.