Modelplane Modelplane docs

ModelRoute Custom Resource

This document is for an unreleased version of Modelplane.

This document applies to the Modelplane main branch and not to the latest release v0.4.

A ModelRoute is one ModelService’s routing on one InferenceGateway: the AIGatewayRoute matching the service’s model name, plus, per endpoint, the Backend, credential and policy the gateway needs to reach it. The ModelService composes one per gateway that serves it, pinned to that gateway, and this renders onto the gateway’s cluster. It is machine-generated. kubectl get modelroutes -l modelplane.ai/service=<name> is the per-gateway view of where a service is served and whether each gateway is carrying it.

#Metadata

API version
modelplane.ai/v1alpha1
Kind
ModelRoute
Scope
Namespaced
Short names
mrt

#Spec

# endpoints required object[] 1–32 items
# name required string 1–63 chars

A stable name for this entry, unique within the service.

pattern: ^[a-z0-9]([-a-z0-9]*[a-z0-9])?$

# priority optional integer ≤ 63

Lower is preferred. Entries at the same priority share traffic by weight. A higher-numbered entry takes a growing share of traffic as lower-numbered ones lose healthy endpoints, and takes over entirely once they have none.

format: int32

# selector required object

Selects ModelEndpoints in this route’s namespace.

# matchLabels required map[string]string
# weight optional integer 1–1000000 default: 1

Share of traffic for this entry relative to the other entries at the same priority, spread as evenly as possible across the endpoints it matches.

format: int32

# gatewayName required string 1–253 chars

Name of the InferenceGateway this route is pinned to. The function resolves the gateway’s cluster, its ProviderConfig and its address from here, and renders onto that cluster.

# serviceName required string 1–253 chars

Name of the ModelService this route belongs to. With the ModelRoute’s own namespace it derives the composed object names and the model name a caller passes (/), so the service and its routes agree on both.

# timeouts required object

How long the gateway waits on this service’s endpoints. A verbatim copy of the ModelService’s spec.timeouts.

# idle required string

How long an endpoint may send nothing. Before its first byte the gateway retries the request, on another endpoint if there is one. After it the stream is cut short. 0s disables it.

pattern: ^([0-9]{1,5}(h|m|s|ms)){1,4}$

# request required string

How long a request may take end to end, including retries.

pattern: ^([0-9]{1,5}(h|m|s|ms)){1,4}$

#Status

# address optional string

External address of the gateway this route is on.

# conditions optional object[]
# endpoints optional object

Observed endpoint counts for this route, across all priorities.

# ready optional integer

format: int32

# total optional integer

format: int32

# model optional string

The name a caller passes as the request’s model, /.