Nemotron API Pricing Calculator & Cost Guide
Calculate NVIDIA Nemotron API costs for 5 models. Compare per-token pricing across Llama Nemotron Ultra, Super, Nano, and earlier Nemotron variants.
Pricing TLDR
- • Pay-per-token with no monthly fees, priced separately for input and output
- • Nano is the cheapest tier; Super and Ultra are the larger reasoning models
- • 5 Nemotron models with live rates from the OpenRouter catalog
Nemotron API Cost Calculator: Monthly Pricing
Calculate by
Input Tokens
Output Tokens
API Calls / Month
Quick Examples:
Sort:
(nvidia/nemotron-3-ultra-550b-a55b)
Context
Quality
Popularity
Per 1M Tokens
In: $0.60
Out: $3.60
Monthly Cost
(nvidia/nemotron-3-ultra-550b-a55b:batch)
Context
Quality
Popularity
Per 1M Tokens
In: $0.60
Out: $3.60
Monthly Cost
(nvidia/nemotron-3.5-lightning)
Context
Quality
Popularity
Per 1M Tokens
In: $0.08
Out: $0.20
Monthly Cost
(nvidia/nemotron-3-nano-30b-a3b)
Context
Quality
Popularity
Per 1M Tokens
In: $0.05
Out: $0.20
Monthly Cost
(nvidia/nemotron-3-super-120b-a12b)
Context
Quality
Popularity
Per 1M Tokens
In: $0.08
Out: $0.40
Monthly Cost
Spending across LLM providers?
Track your AI API costs across all providers in real-time.

About Nemotron
What is Nemotron?
Nemotron is NVIDIA's family of open large language models, many of them Llama derivatives that NVIDIA post-trains for stronger reasoning and agentic work. The lineup spans the compact Llama Nemotron Nano, the mid-size Super, and the flagship Ultra, alongside earlier Nemotron-4 variants. Because the weights are open, you can run Nemotron through NVIDIA NIM microservices, self-host it, or reach it through hosts like OpenRouter, all priced per token.
- Tuned for Reasoning: NVIDIA post-trains Nemotron for math, multi-step reasoning, and tool use, so the Super and Ultra tiers target hard agentic workloads while Nano keeps latency and cost low for simpler tasks.
- Open Weights: Nemotron ships with open weights, so you can self-host at no per-token cost or use a hosted API for convenience. Few reasoning-focused model families offer both paths.
- Runs Anywhere: Serve Nemotron through NVIDIA NIM microservices on your own infrastructure, or call it through aggregators like OpenRouter that route to hosted providers, each billed on the same per-token basis.
When to Use Nemotron
Nemotron is a strong fit when you want open-weight reasoning models with NVIDIA tuning, or when self-hosting on NVIDIA hardware and a hosted API option both matter.
Ideal for
- Reasoning and agentic workloads that benefit from NVIDIA tuning
- Teams that want open weights plus a hosted API option
- Self-hosting on NVIDIA GPUs via NIM microservices
- Cost-sensitive tasks routed to the compact Nano tier
- Projects already standardized on the Llama architecture
Not ideal for
- Tasks that need the absolute top of the quality leaderboard
- Workloads dependent on provider-specific features like prompt caching
- Teams that want a single first-party billing relationship only
Nemotron Pricing Breakdown
How Nemotron API Billing Works
Per-Token Pricing
Each model has separate input (prompt) and output (completion) rates per million tokens. Output is usually priced higher than input. You pay only for tokens processed, with no monthly minimum.
Model Tiers Set the Price
Cost scales with size: Llama Nemotron Nano is the cheapest, then Super, then the flagship Ultra. Choosing the right tier for each task is the single biggest lever on your bill.
Free and Open Options
Nemotron weights are open, so you can self-host at no per-token cost. Some Nemotron models are also free with rate limits on OpenRouter for testing before you pick a paid host.
Host Sets the Rate
Because Nemotron is open weight, the per-token rate depends on which host you use. Compare NVIDIA build, OpenRouter providers, and self-hosting to find the cheapest path for your volume.
Nemotron API Monthly Cost Estimates
Hobby / Testing
$0-15/mo
• Nemotron Nano
• <1K requests/day
• Single project
Light Use
$15-75/mo
• Nano / Super
• 1-5K requests/day
• Mixed tasks
Medium Use
$75-400/mo
• Nemotron Super
• 5-20K requests/day
• Production apps
Heavy Use
$400+/mo
• Nemotron Ultra
• 20K+ requests/day
• Agentic workloads
5 Nemotron Cost Optimization Tips
Match the Model to the Task
Do not run Nemotron Ultra on work that Nano or Super handles well. Reserve the flagship for genuinely hard reasoning, and route classification, extraction, and simple chat to the smaller tiers.
Self-Host the Open Weights
For steady high-volume workloads, serving the open Nemotron weights on your own NVIDIA GPUs via NIM can undercut per-token API pricing once your usage is predictable enough to fill the hardware.
Compare Hosts
Since Nemotron is open weight, the same model can cost different amounts across NVIDIA build, OpenRouter providers, and self-hosting. Check the live rates before you route production traffic.
Trim Input and Cap Output
Input tokens cost money on every call and output usually costs more. Use concise prompts, summarize long context, set max_tokens, and ask for terse responses when a short answer will do.
Track Spend with CostGoat
Watch your Nemotron credit balance and usage in real time with CostGoat. Get desktop or email alerts before you run low, and see which models drive your spend.
Start Tracking Your LLM API Spending
Monitor spending across OpenAI, Anthropic, Google, and other LLM providers from one menubar app.

Nemotron API Pricing FAQ
Common questions about NVIDIA Nemotron API costs and billing
