Prime Inference: Fast, Reliable Serving for Frontier Open Models

Prime's mission is to build frontier open models and the open superintelligence stack for continuously improving agents. We already provide end-to-end post-training infrastructure, from prime-rl and verifiers to sandboxes and RL environments. But the continual learning loop is not complete until a trained model can serve real users, generate new experience, and feed those production traces back into training.

We're excited to release Prime Inference today. It covers both serverless endpoints and reserved capacity and offers resilient serving of frontier open-source models on our GPU infrastructure across multiple datacenters.

Prime Inference began as the serving platform we needed ourselves. Long before public release, it powered large-scale RL rollouts, synthetic data generation, evaluations, and long-running coding agents, processing nearly a trillion tokens every day just internally. Beyond our own workloads, we've also been serving large-scale customer deployments in production since January. This scale pushed us to optimize for sustained performance, quality, and reliability, rather than benchmark speed alone.