ModelService Custom Resource
This document applies to the Modelplane main branch and not to the latest release v0.4.
A ModelService is one model as a caller sees it: a stable name that resolves to whichever ModelEndpoint should serve the next request. The endpoints behind it can be replicas Modelplane runs, models bought from a provider, or both, in more than one region. A caller reaches it by naming it as the model in an ordinary OpenAI or Anthropic request to any InferenceGateway that serves it.
Concept guide: Expose a Model →
#Metadata
#Example
Manifest
apiVersion: modelplane.ai/v1alpha1
kind: ModelService
metadata:
name: qwen-72b
namespace: ml-team
labels:
# Matched by an InferenceGateway's serviceSelector. Your label, under your
# own prefix.
example.org/region: eu
spec:
endpoints:
# Entries at the same priority share traffic by weight, so this pair is a
# 90/10 canary across two deployments. name is a stable handle, unique
# within the service.
- name: stable
priority: 0
weight: 90
selector:
matchLabels:
modelplane.ai/deployment: qwen-72b
- name: canary
priority: 0
weight: 10
selector:
matchLabels:
modelplane.ai/deployment: qwen-72b-next
# Priority 1 takes traffic as priority 0 loses healthy endpoints, which
# makes this provider a failover for the deployments above.
- name: together
priority: 1
selector:
matchLabels:
modelplane.ai/endpoint: together-qwen-72b
#Spec
A stable name for this entry, unique within the service. For example stable, canary, or a provider’s name.
Lower is preferred. Entries at the same priority share traffic by weight. A higher-numbered entry takes a growing share of traffic as lower-numbered ones lose healthy endpoints, and takes over entirely once they have none. A failed request is retried, on another endpoint at the same priority if there is one and then at the next, and each attempt gets that endpoint’s own model name, credential and path. Retrying is only possible until the first byte reaches the caller, because after that the tokens are already sent, so a backend that dies mid-stream truncates the response instead.
Selects ModelEndpoints in this ModelService’s namespace. Scope a service to a region by selecting only endpoints in it; Modelplane stamps an InferenceCluster’s labels onto every endpoint composed there, so the region is declared once on the cluster.
Share of traffic for this entry relative to the other entries at the same priority, spread as evenly as possible across the endpoints it matches. A pair of entries weighted 90 and 10 is a canary. At least 1. A weight of 0 doesn’t deprioritise a backend, it drops it from the gateway’s load assignment entirely, which is indistinguishable from removing the entry and easy to mistake for parking it. Remove the entry instead.
How long a gateway waits on this service’s endpoints. The right values depend on the model and on whether callers stream, so set them from its observed response times.
How long an endpoint may send nothing. Before its first byte, the gateway abandons it, counts it as failed, and retries the request, on another endpoint if there is one. After that the stream is cut short. A streamed response sends its first byte after prefill, so for streaming callers this bounds time to first token and every gap between chunks. A non-streamed response sends nothing until it’s complete, so if any caller doesn’t stream, set this at least as long as request, or to 0s to disable it.
How long a request may take end to end, including retries. Set it above the longest response you expect: prefill plus the maximum output tokens at the model’s decode rate.
#Status
The name a caller passes as the request’s model. Namespaced, so two services can’t collide.
Counts of the ModelRoutes this service composes, one per gateway that serves it. ready is how many are carrying traffic; total is how many gateways serve the service. Per-gateway detail, including each gateway’s address, is on the ModelRoutes themselves: kubectl get modelroutes -l modelplane.ai/service=