Celeris-1 - Fastest general-purpose AI, delivering 2,000+ tokens per second
Celeris-1 is a new language model built on a diffusion-based architecture that achieves 1,664 tokens per second while maintaining comparable accuracy to frontier models. It features an OpenAI-compatible API and delivers up to 15x lower latency than GPT-5.
How does Celeris-1 differ from traditional autoregressive language models?
Celeris-1 uses a new diffusion-based inference architecture instead of traditional autoregressive generation. Conventional models generate one token at a time where each token depends on the previous one, creating inherently sequential latency. Celeris's diffusion architecture enables parallel token generation, resulting in significantly lower latency.
What are the main benefits of switching from OpenAI API to Celeris?
Celeris provides an OpenAI-compatible API, so you only need to change the endpoint to inference.celeris.ai and the rest of your code stays the same. It also delivers lower latency and higher throughput—Celeris generates up to 24x more tokens per second than GPT-5 while maintaining comparable accuracy.
How does Celeris-1 compare to competitors in accuracy?
Celeris-1 achieves 75.9% accuracy on the MMLU-Pro benchmark, placing it near the frontier—only a few points below GPT-5 (81.9%) and other top models. It achieves this with a very short latency of 158 ms, while GPT-5 takes 2 seconds.
How is Celeris-1 priced?
Celeris-1 is billed per token generated. Users pay only for the tokens the model outputs, so increased speed has no impact on pricing—generating 1,000 tokens costs the same whether it takes 100 ms or 1 second.
- OpenAI and Cerebras accelerate GPT-5.6 Sol with Ultrafast Mode — cerebras.ai 74 % match
- Cerebras launches CS-4 AI accelerator claiming up to 30x faster performance — cerebras.ai 72 % match
- Runway Solaris could transform how websites and applications are built — runway.com 70 % match