
Prime Intellect has launched Prime Inference, a production inference service for open AI models that offers developers serverless endpoints and reserved capacity across multiple datacenters.
The service, announced on October 2, 2026, provides an OpenAI-compatible API, allowing developers to use familiar OpenAI-style clients while sending requests to models hosted through Prime Intellect’s infrastructure. The launch is focused on running long-context and agent workloads in production rather than introducing a new model.
Prime Inference currently lists Z.ai’s GLM-5.3 as its first named hosted model. Prime previously made the model available through OpenRouter on September 22, with the company saying that the endpoint has maintained 100% uptime since launch.
Prime’s API is available at api.pinference.ai/api/v1, while its inference documentation provides details on the available serving options and model configurations.
The company is targeting workloads where large prompts, persistent context and repeated tool calls put more pressure on inference infrastructure than conventional chatbot requests. Prime says its internal deployments process hundreds of billions of tokens each day, with the launch post describing the workload as approaching one trillion tokens per day across reinforcement-learning rollouts, evaluations, synthetic data generation and long-running coding agents.
That workload has influenced the architecture behind Prime Inference. The serving stack combines NVIDIA Dynamo, vLLM, Mooncake and FlashInfer, with Prime working alongside NVIDIA and Inferact on parts of the system.
One of the main changes is the separation of prefill and decode workloads. Prompt processing and token generation run on different GPU groups, with the system transferring KV-cache state between them. Prime reported that this arrangement reduced p90 inter-token latency by nearly 40% in its testing.
Long contexts also prompted changes to how cached model state is stored. Prime uses Mooncake to retain KV-cache data in host DRAM and has implemented an NVFP4-based KV-cache format for GLM-5.3. The company reports that the compressed representation reduced the cache size per row from 576 bytes to 352 bytes, increasing cached-token capacity from about 1.09 million to 1.63 million tokens at the same memory budget.
Prime also changed how KV-cache data is laid out for transfer between GPUs. In one test involving tensor parallelism across eight GPUs, it reported reducing the number of transfer descriptors from 19,559 to about 1,940 and lowering mean transfer time from 146 milliseconds to 78 milliseconds.
Scheduling was another source of latency. Prime said its testing found that requests could spend substantial time waiting to enter an active batch even when the required cached context was already available. Reducing the prefill token budget from 8,000 to 4,000 tokens per GPU per step cut median queue wait from 550 milliseconds to 110 milliseconds and reduced median time to first token by about 20% in the workload it tested.
The platform also places significant emphasis on tool-call reliability, reflecting Prime’s focus on agent workloads. The company said its testing uncovered issues involving missing arguments, incorrect argument types, malformed tool calls and problems with more complicated tool schemas. Prime uses structured output constraints, including vLLM’s xgrammar, to restrict generated tool calls to valid schemas.
For infrastructure failures, Prime says the service uses shared circuit breakers, lease-based admission control, automatic capacity recovery, datacenter failover and hardware-level health checks extending to NVLink and InfiniBand. The company also says it maintains spare capacity to allow traffic to move between deployments when problems occur.
Prime Inference supports two main capacity models. Serverless endpoints are intended for workloads with changing demand, while reserved capacity is aimed at customers that need more predictable access to inference resources.
The platform also includes a gateway layer for models served by third-party providers. That distinction matters because not every model accessible through Prime’s API is necessarily running on Prime’s own GPU fleet.
Prime is additionally supporting deployment of trained LoRA adapters, allowing customized model adapters to be served through the same API infrastructure. Its documentation describes an adapter deployment format that combines a base model with an adapter identifier.
OpenRouter currently lists Prime Intellect as a provider for GLM-5.3 at $1.40 per million input tokens and $4.40 per million output tokens. Prime’s own documentation says inference pricing varies by model and directs users to its model catalog for current rates and serving limits.
Prime says it plans to add batch and asynchronous inference for large offline workloads and more direct dedicated deployment options for customers running reserved capacity and fine-tuned models.
Discover more from Aree Blog
Subscribe now to keep reading and get access to the full archive.



