PACINFRAX · PRODUCT
Choose how your model serves requests.
Compare shared requests, dedicated endpoints and asynchronous batches.
Your decision
Match the serving mode to latency, demand and operating responsibility.
- Shared API: variable request demand with model-level input and output units. Compare model revision and compatibility before integration.
- Dedicated endpoint: a controlled serving configuration with an accepted isolation, scaling and drain policy.
- Batch: queued work with a completion window, result retention and cancellation policy rather than an interactive latency promise.
Public service activation is not available here. Compare requirements before choosing a deployment; no resource, price or capacity is reserved.