NEW: 70+ real-time integrations - track Claude, OpenAI, AWS, OpenRouter & more. Try it free →

CostGoat Logo

CostGoat

LAST UPDATED: AUGUST 27, 2026

GLM API Pricing Calculator & Cost Guide

Calculate GLM API costs for 14 models from Zhipu AI (Z.ai). Compare per-token pricing across GLM-4.6, GLM-4.5, GLM-4.5-Air, and GLM-Z1.

CalculatorPricing GuideSave MoneyFAQ

Pricing TLDR

  • Pay-per-token with no monthly fees, priced separately for input and output
  • GLM-4.5-Air is the cheapest tier; GLM-4.6 is the flagship for coding and agents
  • 14 GLM models with live rates from the OpenRouter catalog

Official pricing:

Z.ai

Live rates: OpenRouter

Quality Scores: Theozard

GLM API Cost Calculator: Monthly Pricing

Calculate by

Input Tokens

Output Tokens

API Calls / Month

Quick Examples:

Sort:

(z-ai/glm-5.3)

Context

1.0M

Quality

94

Popularity

#19

Per 1M Tokens

In: $1.40

Out: $4.40

Monthly Cost

$3.60

(z-ai/glm-5.3-flash)

Context

1.3M

Quality

91

Popularity

#37

Per 1M Tokens

In: $0.08

Out: $0.25

Monthly Cost

$0.20

(z-ai/glm-5.2)

Context

1.0M

Quality

83

Popularity

#9

Per 1M Tokens

In: $1.19

Out: $3.74

Monthly Cost

$3.06

(z-ai/glm-5.1)

Context

205K

Quality

65

Popularity

#81

Per 1M Tokens

In: $1.26

Out: $3.96

Monthly Cost

$3.24

(z-ai/glm-5)

Context

205K

Quality

64

Popularity

#100

Per 1M Tokens

In: $0.60

Out: $1.92

Monthly Cost

$1.56

(z-ai/glm-5-turbo)

Context

203K

Quality

62

Popularity

#178

Per 1M Tokens

In: $1.20

Out: $4.00

Monthly Cost

$3.20

(z-ai/glm-5v-turbo)

Context

203K

Quality

56

Popularity

#114

Per 1M Tokens

In: $1.20

Out: $4.00

Monthly Cost

$3.20

(z-ai/glm-4.7)

Context

205K

Quality

55

Popularity

#94

Per 1M Tokens

In: $0.40

Out: $1.75

Monthly Cost

$1.28

(z-ai/glm-4.6)

Context

205K

Quality

46

Popularity

#152

Per 1M Tokens

In: $0.43

Out: $1.75

Monthly Cost

$1.31

(z-ai/glm-4.7-flash)

Context

203K

Quality

37

Popularity

#130

Per 1M Tokens

In: $0.06

Out: $0.40

Monthly Cost

$0.26

(z-ai/glm-4.5)

Context

131K

Quality

31

Popularity

#282

Per 1M Tokens

In: $0.60

Out: $2.20

Monthly Cost

$1.70

(z-ai/glm-4.6v)

Context

131K

Quality

27

Popularity

Per 1M Tokens

In: $0.30

Out: $0.90

Monthly Cost

$0.75

(z-ai/glm-4.5-air)

Context

131K

Quality

26

Popularity

#201

Per 1M Tokens

In: $0.13

Out: $0.85

Monthly Cost

$0.55

(z-ai/glm-4.5v)

Context

66K

Quality

14

Popularity

#351

Per 1M Tokens

In: $0.60

Out: $1.80

Monthly Cost

$1.50

Spending across LLM providers?

Track your AI API costs across all providers in real-time.

7-day free trial, no credit card required

CostGoat desktop app showing AI agent quotas, usage costs, credit balances, and subscriptions

About GLM

What is GLM?

GLM is the model family from Zhipu AI, also branded Z.ai, a Chinese lab building both open-weight and commercial large language models. Its API offers a range of models from the low-cost GLM-4.5-Air up to the flagship GLM-4.6, plus the GLM-Z1 reasoning line and smaller GLM-4 variants. Pricing is per token, billed separately for input and output.

  • Strong Coding at Low Cost: GLM models are known for strong coding and agentic performance relative to their price. GLM-4.6 in particular targets developer and agent workloads while staying well under the cost of Western flagship models.
  • A Model for Every Cost Tier: GLM-4.5-Air and the smaller GLM-4 variants handle high-volume, simple tasks cheaply. GLM-4.5 and GLM-4.6 take on complex reasoning and agentic work. GLM-Z1 adds dedicated reasoning. Pick the smallest model that clears your quality bar.
  • Open Weights and a Coding Plan: Zhipu AI ships open-weight GLM releases you can self-host, and Z.ai offers a subscription GLM Coding Plan with a flat monthly fee for heavy coding agents as an alternative to metered per-token billing.

When to Use GLM

GLM is a strong fit when you want frontier-adjacent quality, especially for coding and agents, at a lower price than GPT or Claude, or when open weights matter.

Ideal for

  • Cost-sensitive coding and agentic workloads
  • Teams that want open weights plus a hosted API option
  • High-volume classification, extraction, and routing
  • Developers weighing a flat-fee GLM Coding Plan
  • Budget-conscious production apps at scale

Not ideal for

  • Tasks that need the absolute top of the quality leaderboard
  • Workloads dependent on provider-specific features like prompt caching
  • Projects with strict data-residency needs outside the provider's regions

GLM Pricing Breakdown

How GLM API Billing Works

Per-Token Pricing

Each GLM model has separate input (prompt) and output (completion) rates per million tokens. Output is usually priced higher than input. You pay only for tokens processed, with no monthly minimum.

Model Tiers Set the Price

Cost scales with capability: GLM-4.5-Air is the cheapest, then GLM-4.5, the GLM-Z1 reasoning models, and the flagship GLM-4.6. Choosing the right tier for each task is the single biggest lever on your bill.

Coding Plan and Open Options

Z.ai offers a flat-fee GLM Coding Plan for heavy coding agents, and open-weight GLM releases can be self-hosted at no per-token cost. Some GLM models are also free with rate limits on OpenRouter.

Usage Tracking

Monitor spend per model and per key in the Z.ai console. Set alerts so a runaway job or a switch to a pricier model does not surprise you at the end of the month.

GLM API Monthly Cost Estimates

Hobby / Testing

$0-15/mo

GLM-4.5-Air

<1K requests/day

Single project

Light Use

$15-75/mo

GLM-4.5

1-5K requests/day

Mixed tasks

Medium Use

$75-400/mo

GLM-4.6

5-20K requests/day

Production apps

Heavy Use

$400+/mo

GLM-4.6 agents

20K+ requests/day

Coding workloads

5 GLM Cost Optimization Tips

1

Match the Model to the Task

Do not run GLM-4.6 on work that GLM-4.5-Air handles well. Reserve the flagship for genuinely hard reasoning and agentic coding, and route classification, extraction, and simple chat to the cheap tiers.

2

Consider the GLM Coding Plan

If you run heavy, steady coding-agent workloads, the flat-fee GLM Coding Plan from Z.ai can undercut metered per-token billing. Compare its allowance against your expected API cost before subscribing.

3

Trim Input Tokens

Input tokens cost money on every call. Use concise system prompts, summarize long context instead of pasting whole documents, and drop history you do not need.

4

Cap Output Length

Output tokens usually cost more than input. Set max output tokens, ask for structured or terse responses, and avoid open-ended generations when a short answer will do.

5

Track Spend with CostGoat

Watch your GLM credit balance and usage in real time with CostGoat. Get desktop or email alerts before you run low, and see which models drive your spend.

Start Tracking Your LLM API Spending

Monitor spending across OpenAI, Anthropic, Google, and other LLM providers from one menubar app.

7-day free trial, no credit card required

CostGoat desktop app showing AI agent quotas, usage costs, credit balances, and subscriptions

GLM API Pricing FAQ

Common questions about GLM (Zhipu AI / Z.ai) API costs and billing

AI Pricing

Gemini API PricingClaude API PricingGoogle Veo PricingAI Cost CalculatorsReplicate API PricingOpenRouter API PricingOpenRouter Free Models
DownloadsPricingDealsAccountContactIssuesAffiliatesTermsPrivacy

© 2026 CostGoat. All rights reserved.

Made by Functioncraft: Redis GUI Client · SSH GUI Client

Affiliate disclosure: Some links earn CostGoat a commission or credit when you sign up — no extra cost to you.