Llama API Pricing Calculator & Cost Guide
Calculate Llama API costs for 7 models. Compare per-token host pricing across Llama 3.1, Llama 3.2, Llama 3.3, and Llama 4 Scout and Maverick.
Pricing TLDR
- • Llama is open-weight: run it yourself for free compute cost, or pay a host per token
- • Llama 3.1 8B and 3.2 are the cheapest tiers; 405B and Llama 4 Maverick cost the most
- • 7 Llama models with representative host rates from the OpenRouter catalog
Llama API Cost Calculator: Monthly Pricing
Calculate by
Input Tokens
Output Tokens
API Calls / Month
Quick Examples:
Sort:
(meta-llama/llama-4-maverick)
Context
Quality
Popularity
Per 1M Tokens
In: $0.20
Out: $0.80
Monthly Cost
(meta-llama/llama-4-scout)
Context
Quality
Popularity
Per 1M Tokens
In: $0.11
Out: $0.34
Monthly Cost
(meta-llama/llama-3.3-70b-instruct)
Context
Quality
Popularity
Per 1M Tokens
In: $0.71
Out: $0.71
Monthly Cost
(meta-llama/llama-3.1-8b-instruct)
Context
Quality
Popularity
Per 1M Tokens
In: $0.05
Out: $0.08
Monthly Cost
(meta-llama/llama-3.1-70b-instruct)
Context
Quality
Popularity
Per 1M Tokens
In: $0.40
Out: $0.40
Monthly Cost
(meta-llama/llama-3.2-3b-instruct)
Context
Quality
Popularity
Per 1M Tokens
In: $0.05
Out: $0.33
Monthly Cost
(meta-llama/llama-3.2-1b-instruct)
Context
Quality
Popularity
Per 1M Tokens
In: $0.03
Out: $0.20
Monthly Cost
Spending across LLM providers?
Track your AI API costs across all providers in real-time.

About Llama
What is Llama?
Llama is Meta's family of open-weight large language models. Meta releases the weights under an open license, so there is no first-party paid Llama API in the usual sense. Instead, hosts like OpenRouter, Together, Fireworks, and Groq serve the models and bill per token, while teams that prefer control can download the weights and self-host. The lineup spans Llama 3.1 (8B, 70B, 405B), Llama 3.2, Llama 3.3 70B, and Llama 4 Scout and Maverick.
- Open Weights, Run Anywhere: Because the weights are freely available, you can self-host Llama on your own hardware, rent a GPU, or pick from many hosted providers. That flexibility is Llama's defining feature and keeps any single vendor from locking you in.
- A Wide Size Range: From the tiny Llama 3.2 and 8B models up to 405B and Llama 4 Maverick, the family covers a broad span of cost and capability. Small models handle high-volume simple tasks cheaply; the largest tackle complex reasoning.
- Many Hosts, Many Prices: The same Llama model can be served by several providers at different rates and speeds. You can shop for the best price or latency for a given model rather than accepting one vendor's number.
When to Use Llama
Llama is a strong fit when open weights, portability, or the freedom to self-host matter, and when you want a wide range of model sizes to match cost to task.
Ideal for
- Teams that want to avoid vendor lock-in
- Steady high-volume workloads suited to self-hosting
- Projects needing on-premise or air-gapped deployment
- Comparing hosts to shop for the best rate per model
- High-volume classification, extraction, and routing on small models
Not ideal for
- Tasks that need the absolute top of the quality leaderboard
- Teams wanting a single official vendor and support contract
- Workloads dependent on provider-specific features like prompt caching
Llama Pricing Breakdown
How Llama Pricing Works
Host Pricing vs Self-Hosting
Llama pricing takes two forms. Hosted providers charge per token with separate input and output rates and no infrastructure to manage. Self-hosting means downloading the free weights and paying only for the compute you run, which favors steady high-volume traffic.
Model Size Sets the Price
Cost scales with model size. Llama 3.1 8B and Llama 3.2 are the cheapest, 70B and Llama 4 Scout sit in the middle, and 405B and Llama 4 Maverick cost the most because they need far more compute to serve.
Rates Vary by Provider
The same model can be served by several hosts at different prices and speeds. It pays to compare providers, since one host may be cheaper for a given model while another is faster or serves a longer context.
Usage Tracking
Whichever host you pick, monitor spend per model and per key in that provider's console. Set alerts so a runaway job or a switch to a larger model does not surprise you at the end of the month.
Llama API Monthly Cost Estimates
Hobby / Testing
$0-15/mo
• Llama 3.1 8B / 3.2
• <1K requests/day
• Single project
Light Use
$15-75/mo
• Llama 3.3 70B
• 1-5K requests/day
• Mixed tasks
Medium Use
$75-400/mo
• Llama 4 Scout
• 5-20K requests/day
• Production apps
Heavy Use
$400+/mo
• 405B / Maverick
• 20K+ requests/day
• Agentic workloads
5 Llama Cost Optimization Tips
Match the Model Size to the Task
Do not run 405B or Llama 4 Maverick on work that an 8B or 70B model handles well. Reserve the largest models for genuinely hard reasoning, and route classification, extraction, and simple chat to the small tiers.
Self-Host for Steady Volume
Since the weights are free, self-hosting can undercut per-token host pricing once your usage is predictable enough to keep a GPU busy. For steady high-volume workloads this is often the cheapest path.
Compare Hosts
The same Llama model is served by many providers at different rates. Check a few hosts before committing, since prices and speeds vary and can shift as providers compete.
Trim Input and Cap Output
Input tokens cost money on every call, and output usually costs more. Use concise prompts, summarize long context instead of pasting whole documents, and set a max output length so generations stay short.
Track Spend with CostGoat
Watch your Llama host credit balance and usage in real time with CostGoat. Get desktop or email alerts before you run low, and see which models drive your spend.
Start Tracking Your LLM API Spending
Monitor spending across OpenAI, Anthropic, Google, and other LLM providers from one menubar app.

Llama API Pricing FAQ
Common questions about Llama model costs and billing
