All Z.ai GLM API Plans & Pricing

Plan Monthly Annual Best For
View all features by plan (compare side-by-side)

GLM-5.2 ($1.40 in / $4.40 out per 1M tokens)

  • $1.40 per 1M input tokens
  • $0.26 cached input
  • $4.40 per 1M output tokens

GLM-5 ($1.00 in / $3.20 out per 1M tokens)

  • $1.00 per 1M input tokens
  • $0.20 cached input
  • $3.20 per 1M output tokens

GLM-4.7 ($0.60 in / $2.20 out per 1M tokens)

  • $0.60 per 1M input tokens
  • $0.11 cached input
  • $2.20 per 1M output tokens

GLM-4.7-Flash / GLM-4.5-Flash (free)

  • Free model tier
  • GLM-4.6V-Flash vision also free
Pricing Alerts

Track Z.ai GLM API pricing

Get an email when Z.ai GLM API's pricing changes — plus the weekly SaaS Price Watch: verified price changes and deals across 3,000+ products. One-click unsubscribe.

Compare Z.ai GLM API with alternativesAdjust seats, lock a tier, add up to 2 more products side-by-side. Shareable URL.
Quick Answer
Last verified:
High confidence

Z.ai GLM API costs Free to $4.40 per million tokens as of September 2026, with 4 plans available including a free tier. Plans: GLM-5.2 ($1.40 in / $4.40 out per 1M tokens) (free), GLM-5 ($1.00 in / $3.20 out per 1M tokens) (free), GLM-4.7 ($0.60 in / $2.20 out per 1M tokens) (free), and GLM-4.7-Flash / GLM-4.5-Flash (free) (free). Pricing depends on your chosen tier, contract length, and negotiated discounts.

Use the interactive pricing calculator to estimate your exact cost based on team size and requirements.

  • Free tier: Yes

Z.ai GLM API offers 4 pricing tiers: GLM-5.2 ($1.40 in / $4.40 out per 1M tokens), GLM-5 ($1.00 in / $3.20 out per 1M tokens), GLM-4.7 ($0.60 in / $2.20 out per 1M tokens), GLM-4.7-Flash / GLM-4.5-Flash (free). The GLM-5 ($1.00 in / $3.20 out per 1M tokens) plan is strong general-purpose workloads.

Compared to other llm api providers software, Z.ai GLM API is positioned at the budget-friendly price point.

  • 20 documented hidden costs beyond list price

How much does Z.ai GLM API cost?

Z.ai GLM API offers a free plan with 3 paid tiers from $Infinity to $4.40 per million tokens. Plans include GLM-5.2 ($1.40 in / $4.40 out per 1M tokens) (free), GLM-5 ($1.00 in / $3.20 out per 1M tokens) (free), GLM-4.7 ($0.60 in / $2.20 out per 1M tokens) (free), GLM-4.7-Flash / GLM-4.5-Flash (free) (free).

Z.ai GLM API Pricing Overview

Z.ai GLM API has 4 pricing plans, including a free tier. Paid plans range from $0 to $4.40/per million tokens. The GLM-5.2 ($1.40 in / $4.40 out per 1M tokens) plan is free and is best for frontier-quality glm workloads. The GLM-5 ($1.00 in / $3.20 out per 1M tokens) plan is free and is best for strong general-purpose workloads. The GLM-4.7 ($0.60 in / $2.20 out per 1M tokens) plan is free and is best for cost-efficient production workloads. The GLM-4.7-Flash / GLM-4.5-Flash (free) plan is free and is best for prototyping and light workloads at zero cost.

Z.ai GLM API with a $10,000-$50,000 for volume discounts minimum commitment.

There are at least 20 documented hidden costs beyond Z.ai GLM API's list price, including implementation, training, and add-on fees.

This pricing was last verified in August 17, 2026 from 1 independent source.

How Z.ai GLM API Pricing Compares

Compare Z.ai GLM API pricing against top alternatives in LLM API Providers.

Compare Z.ai GLM API vs Alternatives

Before committing to Z.ai GLM API, compare pricing with these 3 alternatives in the same category.

All Z.ai GLM API alternatives & migration guides

What Companies Actually Pay for Z.ai GLM API

Review scores
Trustpilot 2.2out of 5 (31)
Third-party review aggregates, as of Aug 2026
Top pricing complaints
Unreliability and Slowness"Peak Hours" and Service UnavailabilityAPI Errors and Vanished QuotaPoor Customer Support and Refund Issues

How Z.ai GLM API Pricing Compares

Software Starting Price Top Price
Z.ai GLM API Free $4.4 per million tokens
Amazon Bedrock $0.07 per million tokens $75 per million tokens
Anyscale Free $4.9591 per million tokens
Baidu ERNIE API $0.1 per million tokens $10 per million tokens
Cerebras Inference API $0.1 per million tokens $6 per million tokens
Cohere API Free $15 per million tokens

20 Z.ai GLM API Hidden Costs Beyond the List Price

Beyond the listed price, Z.ai GLM API has at least 20 documented hidden costs that can significantly increase total cost of ownership.

Watch for 20 hidden costs
  • Cached Input Storage $0.26 per 1 million tokens
    medium 1 source
    industry "Cached Input Storage: Z.ai currently lists cached input storage as free for a limited time, implying it could become a paid feature in the future"
  • Premium Model Tiers 64% higher for input and 81% higher for output
    high 1 source
    industry "Its high-throughput tier, GLM-5.2-Fast, is priced higher at $2.29 per 1 million input tokens and $8.00 per 1 million output tokens"
  • Conversation History Reprocessing 10x cost multiplier
    critical 1 source
    industry "A 50-message thread can send the equivalent of a short document, leading to a "10x cost multiplier" if not managed through prompt caching"
  • Suboptimal Model Selection
    medium 1 source
    industry "For instance, GLM-5.2, a flagship model, costs $1.40 per 1 million input tokens and $4.40 per 1 million output tokens"
  • Data Training Opt-ins
    high 1 source
    industry "Cheaper models like GLM-4.7-FlashX are available at $0.07 input and $0.40 output per 1 million tokens, with some "Flash" models being free"
  • System Prompt Overhead $400 monthly
    high 1 source
    industry "A 2,000-token system prompt with 100,000 requests per month could cost around $400 monthly at GPT-4.1 rates before any user tokens are counted."
  • Uncapped Output Generation
    critical 1 source
    industry "Uncapped Output Generation: Models can produce variable-length responses, potentially exceeding projected output tokens by 10-25 times if max_tokens limits are not explicitly set."
  • Web Search Tool Usage $0.01 per use
    low 1 source
    industry "ai's built-in Web Search tool costs an additional $0.01 per use, on top of the token charges for the request."
  • Latency and Data Residency
    medium 1 source
    industry "ai's infrastructure is primarily located in China, which could affect latency for international users and raise data residency concerns, potentially leading to additional compliance or infrastructure costs for some buyers."
  • Model Verbosity
    medium 1 source
    industry "Model Verbosity: Even with identical per-token rates, a more verbose model can lead to higher overall costs for a completed workload due to generating more output tokens."
  • Token Pricing Nuances
    medium 1 source
    industry "Token-based Pricing Nuances: LLM API pricing is primarily based on the number of tokens processed, with output tokens typically costing 4 to 10 times more than input tokens due to higher compute requirements"
  • Context Window Billing
    medium 1 source
    industry "Context Window Costs: Every token in the context window, including system prompts, conversation history, and retrieved documents, is billed as input on each API call"
  • Missing Prompt Caching 80-90%
    high 1 source
    industry "Context Window Costs: Every token in the context window, including system prompts, conversation history, and retrieved documents, is billed as input on each API call"
  • Unbatched Requests 40-50%
    high 1 source
    industry "Not Batching Requests: Not batching non-urgent requests can lead to missed discounts of 40-50%"
  • Over-spec'd Model Usage 80-95%
    high 1 source
    industry "Switching from a flagship model like GPT-5.2 to an equivalent open-weight model can reduce monthly API spend by 80-95%"
  • Vendor Lock-in $15,000-$30,000
    critical 1 source
    industry "Vendor Lock-in: This is a significant hidden cost, often creeping in through provider-specific SDKs, prompts tuned to one model's quirks, and evaluations trapped in one vendor's platform"
  • Egress Charges 5-10%
    high 1 source
    industry "Egress Charges: For applications with large token volumes and long generation lengths, data egress charges (data leaving the provider's network) can add 5-10% to per-token costs, as seen with services like AWS Bedrock charging $0."
  • Vendor Lock-in and Switching Costs
    high 1 source
    industry "Vendor Lock-in and Switching Costs: Changing LLM API providers can incur costs related to code changes and potential modifications to business logic if providers differ in inference speeds or output consistency."
  • Account Minimum Commitments $10,000-$50,000
    high 1 source
    industry "Account Minimum Commitments: Enterprise providers may require annual commitments, such as $10,000-$50,000 for volume discounts, which can be sunk costs for early-stage companies with lower usage."
  • Lack of Observability
    high 1 source
    industry "Only 43% track spending by customer, and just 22% track costs by individual transaction, leading to suboptimal optimization decisions."
Tip

Ask your Z.ai GLM API sales rep about these costs upfront. Getting them in writing before signing can save you from surprise charges later.

Full hidden costs breakdown →

Intelligence sourced from 2 independent sources
industry Trustpilot Consumer reviews
Key claims include inline source attribution. Data verified against multiple independent sources. 20 source citations total.

Z.ai GLM API Contract Terms

Z.ai GLM API contracts do not auto-renew. Changes require advance notice. These terms are sourced from verified buyer experiences.

Contract Terms
Minimum Commitment $10,000-$50,000 for volume discounts

Is this pricing incorrect? — we'll verify and update it.