GLM-5.3
A frontier open MoE model from the GLM family — strong on reasoning, coding and agentic tool use across a 1M-token context window.
AI inference infrastructure
A better inference system. The inference cloud for open models — fast, affordable, and built to scale with you.
A drop-in inference service for open models. Point your workload at our API and serve the same models far below standard cost — we keep the KV cache resident on cheap NAND and optimise the whole memory hierarchy for you. Strongest exactly where it hurts most: long context, high batch, and agent workloads.
Create an account with your email, add credit, and create a key. It takes about a minute.
Sign upOpen consoleDeployment
Available now — shared endpoint, pay as you go.
The same stack in your cloud or on-prem, tuned to your workload. Contact us →
Models
A frontier open MoE model from the GLM family — strong on reasoning, coding and agentic tool use across a 1M-token context window.
GLM-5.3 in fp4 on a faster lane — higher throughput and priority scheduling for latency-sensitive agents and interactive coding.
The previous GLM generation, served in fp8 at the same rate as GLM-5.3 — for workloads already tuned to 5.2 or that need a stable target.
Team
Supernova is a team of systems engineers and researchers from the world's leading memory, high-performance-computing, and machine-learning groups, with degrees from top Ivy League universities and several international inference competitions won between us.
Run inference on the Supernova cloud — far below HBM-bound cost. Sign up with your email, add credit, and you're calling the model in a minute.
Already have one? Log in · Questions? contact@supernova-labs.ai
Questions? contact@supernova-labs.ai