Inference Cloud

Inference Cloud

Overview

Cerebras-operated cloud that serves open and dedicated models through an OpenAI-compatible API, plus a training cloud for fine-tunes and from-scratch runs. Production customers include OpenAI (GPT-5.6 Sol Ultrafast), Meta's Llama API, Mistral, Perplexity, Cognition, Lovable, CrowdStrike, AlphaSense, GSK, Block, and Figma. Partner distribution includes AWS Marketplace, OpenRouter, Hugging Face, Vercel, and IBM watsonx.

Core cloud revenue $127.7M in Q2 2026 (+287% YoY). AMD Helios disaggregated inference is due in production on the cloud in Q4 2026; AWS Bedrock is targeted for Q1 2027. New 165 MW Mikkeli capacity is under construction.

Key Specs

Q2 2026 core cloud revenue
$127.7M (+287% YoY)
Frontier decode speed
GPT-5.6 Sol Ultrafast up to 750 tok/s; open models up to ~3,000 tok/s (GPT-OSS 120B)
Contracted capacity
600+ MW by end-2027 (live and under contract)