Pricing

Simple per-token pricing.

Prepaid credits · billed in USD · pay by card

You buy credits up front, allocate them to API keys, and every request draws down the key by its exact token count, billed to $0.000001. One key works with every model we serve. No subscriptions, no minimum commitment beyond the first top-up.

ModelInput · cache missInput · cache hitOutput
glm-5.3availablefp8$1.40$0.26$4.40
glm-5.3-fastavailablefp4$2.80$0.52$8.80
glm-5.2availablefp8$1.40$0.26$4.40

glm-5.3 and glm-5.2 are served in fp8 at the base rate; glm-5.3-fast serves GLM-5.3 in fp4 on a faster lane with priority scheduling, priced at twice the base rate. Pick the model per request with the model field ("model": "glm-5.3"), or ANTHROPIC_MODEL=glm-5.3 for Claude Code — details in the docs. Prices are USD per million tokens. A cache hit is an input token whose prefix was already processed in a recent request; it is billed at a fraction of a miss, so multi-turn conversations stay cheap. Reasoning tokens are billed as output.

How credits work

Credits are prepaid and non-refundable except as set out in our Refund Policy. Need higher limits, invoicing, or a private deployment? Email contact@supernova-labs.ai.