Together AI API Pricing
Complete pricing for all Together AI models. Input/output costs per 1M tokens, context windows, and rate limits.
Billing terms: priced in USD per 1M tokens unless a row states otherwise; official direct rates are listed separately from third-party channels, and regional prices are kept distinct. Record updated 2026-10-10.
Together AI API Pricing Calculator
Select a model to estimate monthly cost using its listed rates on each channel.
This model is deprecated. Migrate before retirement.
Workloads reusing system prompts, long documents, or conversation history hit more; channels without a cached rate are billed at the full input rate.
| Channel | Effective input | Output /1M | Est. / month |
|---|---|---|---|
| together | $0.60 | $1.70 | $2.56 |
Formula: monthly cost = (input tokens × effective input rate + output tokens × output rate) ÷ 1,000,000 × calls. The effective input rate blends list and cached rates at your cache-hit share. Rates share the selected model and billing basis; your vendor bill is authoritative.
Currently available rates
Cogito v2.1 671B
LLMtogether-deepcogito-cogito-v2-1-671bDeepSeek R1
LLMtogether-deepseek-r1DeepSeek V3
LLMtogether-deepseek-v3DeepSeek V3.1
LLMtogether-deepseek-ai-deepseek-v3-1Deprecated endpoints may still be callable. Migrate before retirement.
DeepSeek V4 Flash 0731
LLMtogether-deepseek-ai-deepseek-v4-flash-0731DeepSeek V4 Pro
LLMtogether-deepseek-ai-deepseek-v4-proDeepSeek V4 Pro 0813
LLMtogether-deepseek-ai-deepseek-v4-pro-0813DeepSeek V4.1 Flash
LLMtogether-deepseek-ai-deepseek-v4.1-flashDeepSeek-R1
LLMtogether-deepseek-ai-deepseek-r1Deprecated endpoints may still be callable. Migrate before retirement.
DeepSeek-V3
LLMtogether-deepseek-ai-deepseek-v3Deprecated endpoints may still be callable. Migrate before retirement.
FLUX.1 Dev
Image Gentogether-flux-1-devFLUX.1 Schnell
Image Gentogether-flux-1-schnellGLM-5
LLMtogether-zai-org-glm-5Deprecated endpoints may still be callable. Migrate before retirement.
GLM-5.1
LLMtogether-zai-org-glm-5.1Deprecated endpoints may still be callable. Migrate before retirement.
GLM-5.2
LLMtogether-zai-org-glm-5.2GLM-5.3
LLMtogether-zai-org-glm-5.3GLM-5.3-Flash
LLMtogether-zai-org-glm-5.3-flashGPT OSS 120B
LLMtogether-openai-gpt-oss-120bGPT OSS 20B
LLMtogether-openai-gpt-oss-20bGemma 3N E4B Instruct
LLMtogether-google-gemma-3n-e4b-itGemma 4 31B Instruct
LLMtogether-google-gemma-4-31b-itInkling
LLMtogether-thinkingmachines-inklingKimi K2.5
LLMtogether-moonshotai-kimi-k2.5Kimi K2.6
LLMtogether-moonshotai-kimi-k2.6Kimi K2.7 Code
LLMtogether-moonshotai-kimi-k2.7-codeKimi K3
LLMtogether-moonshotai-kimi-k3LFM2-24B-A2B
LLMtogether-liquidai-lfm2-24b-a2bLlama 3.1 405B Turbo
LLMtogether-llama-3.1-405bLlama 3.1 8B Turbo
LLMtogether-llama-3.1-8bLlama 3.3 70B
LLMtogether-meta-llama-llama-3.3-70b-instruct-turboLlama 3.3 70B Turbo
LLMtogether-llama-3.3-70bMeta Llama 3 8B Instruct Lite
LLMtogether-meta-llama-meta-llama-3-8b-instruct-liteMiniMax-M2.5
LLMtogether-minimaxai-minimax-m2.5Deprecated endpoints may still be callable. Migrate before retirement.
MiniMax-M2.7
LLMtogether-minimaxai-minimax-m2.7MiniMax-M3
LLMtogether-minimaxai-minimax-m3Mistral Large 2
LLMtogether-mistral-largeNemotron 3 Ultra 550B A55B
LLMtogether-nvidia-nemotron-3-ultra-550b-a55bPearl AI Gemma 4 31B Instruct
LLMtogether-pearl-ai-gemma-4-31b-itQwen 2.5 72B Turbo
LLMtogether-qwen-2.5-72bQwen 2.5 7B Instruct Turbo
LLMtogether-qwen-qwen2.5-7b-instruct-turboQwen3 235B A22B Instruct 2507 FP8
LLMtogether-qwen-qwen3-235b-a22b-instruct-2507-tputDeprecated endpoints may still be callable. Migrate before retirement.
Qwen3 Coder 480B A35B Instruct
LLMtogether-qwen-qwen3-coder-480b-a35b-instruct-fp8Deprecated endpoints may still be callable. Migrate before retirement.
Qwen3 Coder Next FP8
LLMtogether-qwen-qwen3-coder-next-fp8Deprecated endpoints may still be callable. Migrate before retirement.
Qwen3.5 397B A17B
LLMtogether-qwen-qwen3.5-397b-a17bDeprecated endpoints may still be callable. Migrate before retirement.
Qwen3.5 9B
LLMtogether-qwen-qwen3.5-9bQwen3.6 Plus
LLMtogether-qwen-qwen3.6-plusQwen3.7 Max
LLMtogether-qwen-qwen3.7-maxRnj-1 Instruct
LLMtogether-essentialai-rnj-1-instructDeprecated endpoints may still be callable. Migrate before retirement.
Does this change your best option?
Check your models and usage for free. See the relevant official changes, estimated cost, and a practical next step.