AI inference infrastructure

The best inference experience.

A better inference system. The inference cloud for open models — fast, affordable, and built to scale with you.

SOTA
open benchmark
world champion · global inference competition
Today
runs on the GPUs you rent now

Provided LLM API

● Available now — early access Software

The Supernova inference cloud

A drop-in inference service for open models. Point your workload at our API and serve the same models far below standard cost — we keep the KV cache resident on cheap NAND and optimise the whole memory hierarchy for you. Strongest exactly where it hurts most: long context, high batch, and agent workloads.

  • Lower cost per token on open models — GLM-5.3, GLM-5.3 Fast and GLM-5.2 today, more on the way

Get an API key

Create an account with your email, add credit, and create a key. It takes about a minute.

Sign upOpen console

Already have an account? Log in · Docs

You are signed in · Models · Docs

Deployment

Nova Cloud API

Available now — shared endpoint, pay as you go.

Private deployment

The same stack in your cloud or on-prem, tuned to your workload. Contact us →

Models

Explore our featured models.

See pricing →

zai-org/available

GLM-5.3

A frontier open MoE model from the GLM family — strong on reasoning, coding and agentic tool use across a 1M-token context window.

PriorityFlexfp81M
zai-org/available

GLM-5.3 Fast

GLM-5.3 in fp4 on a faster lane — higher throughput and priority scheduling for latency-sensitive agents and interactive coding.

FastPriorityfp41M
zai-org/available

GLM-5.2

The previous GLM generation, served in fp8 at the same rate as GLM-5.3 — for workloads already tuned to 5.2 or that need a stable target.

PriorityFlexfp81M

Team

Built by a world-class memory-systems & HPC team.

Supernova is a team of systems engineers and researchers from the world's leading memory, high-performance-computing, and machine-learning groups, with degrees from top Ivy League universities and several international inference competitions won between us.

Champion · global inference competition Top supercomputing team World-ranked system team

Get an API key.

Run inference on the Supernova cloud — far below HBM-bound cost. Sign up with your email, add credit, and you're calling the model in a minute.

Create an accountOpen console

Already have one? Log in · Questions? contact@supernova-labs.ai

Questions? contact@supernova-labs.ai