Cohere API Pricing Calculator & Cost Guide
Calculate Cohere Command API costs for 4 chat models. Compare per-token pricing across Command A, Command R+, Command R, and Command R7B.
Pricing TLDR
- • Pay-per-token with no monthly fees, priced separately for input and output
- • Command R7B is the cheapest tier; Command A is the enterprise flagship
- • 4 Cohere Command chat models with live rates from the OpenRouter catalog
Cohere API Cost Calculator: Monthly Pricing
Calculate by
Input Tokens
Output Tokens
API Calls / Month
Quick Examples:
Sort:
(cohere/command-a)
Context
Quality
Popularity
Per 1M Tokens
In: $2.50
Out: $10.00
Monthly Cost
(cohere/command-r7b-12-2024)
Context
Quality
Popularity
Per 1M Tokens
In: $0.04
Out: $0.15
Monthly Cost
(cohere/command-r-08-2024)
Context
Quality
Popularity
Per 1M Tokens
In: $0.15
Out: $0.60
Monthly Cost
(cohere/command-r-plus-08-2024)
Context
Quality
Popularity
Per 1M Tokens
In: $2.50
Out: $10.00
Monthly Cost
Spending across LLM providers?
Track your AI API costs across all providers in real-time.

About Cohere
What is Cohere?
Cohere is an enterprise-focused AI lab whose Command family of chat models targets retrieval-augmented generation, tool use, and agentic workflows for businesses. The Command models run from the small Command R7B up to the flagship Command A, with Command R+ and Command R in the middle. Pricing is per token, billed separately for input and output. Cohere also sells Embed and Rerank models that power search and RAG, and those are priced separately on its own pricing page.
- Built for Enterprise RAG: The Command models are tuned for retrieval-augmented generation with citations, so they pair naturally with Cohere's Embed and Rerank stack for grounded, source-backed answers over your own documents.
- A Model for Every Cost Tier: Command R7B handles high-volume simple tasks cheaply. Command R and Command R+ cover most production work, and Command A takes on the hardest reasoning and tool-use jobs. Pick the smallest model that clears your quality bar.
- Strong Tool Use and Retrieval: Beyond chat, the Command models are designed for multi-step tool calling and agentic tasks, which is where Cohere positions itself for enterprise automation.
When to Use Cohere
Cohere is a strong fit when retrieval quality and grounded answers matter more than chasing the top of the general chat leaderboard, especially in enterprise settings.
Ideal for
- Enterprise RAG over private document collections
- Grounded answers with citations and source tracking
- Search and relevance pipelines using Embed and Rerank
- Multi-step tool use and agentic automation
- High-volume classification and extraction on Command R7B
Not ideal for
- Tasks that need the absolute top of the general quality leaderboard
- Teams that only need a cheap general-purpose chat model
- Workloads dependent on provider-specific features like prompt caching
Cohere Pricing Breakdown
How Cohere API Billing Works
Per-Token Pricing
Each Command model has separate input (prompt) and output (completion) rates per million tokens. Output is usually priced higher than input. You pay only for tokens processed, with no monthly minimum.
Model Tiers Set the Price
Cost scales with capability: Command R7B is the cheapest, then Command R and Command R+, up to the flagship Command A. Choosing the right tier for each task is the single biggest lever on your chat bill.
Embed and Rerank Billed Separately
The live table here covers the Command chat models. Cohere's Embed and Rerank models, which drive its retrieval and search features, are priced on their own units and listed on Cohere's pricing page rather than the token table above.
Usage Tracking
Monitor spend per model and per key in the Cohere dashboard. Set alerts so a runaway job or a switch to a pricier model does not surprise you at the end of the month.
Cohere API Monthly Cost Estimates
Hobby / Testing
$0-15/mo
• Command R7B
• <1K requests/day
• Single project
Light Use
$15-75/mo
• Command R
• 1-5K requests/day
• Basic RAG
Medium Use
$75-400/mo
• Command R+
• 5-20K requests/day
• Production RAG
Heavy Use
$400+/mo
• Command A
• 20K+ requests/day
• Agentic workloads
5 Cohere Cost Optimization Tips
Match the Model to the Task
Do not run Command A on work that Command R7B or Command R handles well. Reserve the flagship for genuinely hard reasoning and tool use, and route classification, extraction, and simple chat to the cheaper tiers.
Lean on Rerank Before Generation
In a RAG pipeline, using Embed and Rerank to feed only the most relevant chunks into a Command model cuts input tokens and often lets a smaller, cheaper model do the job.
Trim Input Tokens
Input tokens cost money on every call. Use concise system prompts, summarize long context instead of pasting whole documents, and drop retrieved passages you do not need.
Cap Output Length
Output tokens usually cost more than input. Set a max token limit, ask for structured or terse responses, and avoid open-ended generations when a short answer will do.
Track Spend with CostGoat
Watch your Cohere credit balance and usage in real time with CostGoat. Get desktop or email alerts before you run low, and see which models drive your spend.
Start Tracking Your LLM API Spending
Monitor spending across OpenAI, Anthropic, Google, and other LLM providers from one menubar app.

Cohere API Pricing FAQ
Common questions about Cohere API costs and billing
