ModelEndpoint Custom Resource
This document applies to the Modelplane main branch and not to the latest release v0.4.
A ModelEndpoint is somewhere a request can be served: one replica of a ModelDeployment, or a model at a provider like Together or Groq. It describes a backend well enough for a gateway to talk to it without knowing where it came from, so a ModelService can fan over endpoints Modelplane runs and endpoints from a provider. Modelplane composes one per replica. You write them by hand for anything it doesn’t run.
Concept guide: Route to External Providers →
#Metadata
#Example
Manifest
# ModelDeployment composes a ModelEndpoint per replica. Write one by hand only
# to register a model Modelplane doesn't run, like this one at Together.
apiVersion: modelplane.ai/v1alpha1
kind: ModelEndpoint
metadata:
name: together-qwen-72b
namespace: ml-team
labels:
modelplane.ai/endpoint: together-qwen-72b
spec:
# Scheme and host, no path. An https origin gets TLS originated to it. The
# host must be a name; an address stops the gateway applying the model
# rewrite, the credential and priority failover.
origin: https://api.together.xyz
api:
# OpenAI (the default) or Anthropic. The gateway translates an Anthropic
# caller's request for an OpenAI backend, but an Anthropic backend serves
# only Anthropic callers.
schema: OpenAI
# The path this backend serves that API under. /v1 for most, /openai/v1 for
# Groq, a per-replica path for a Modelplane-composed endpoint.
prefix: /v1
# The name this backend knows the model by. Unset, the caller's model name
# passes through unchanged.
model: Qwen/Qwen2.5-72B-Instruct-Turbo
# This backend's credential, attached by the gateway on the way out. APIKey
# sends the key in the Secret as a bearer token, or in x-api-key to an
# Anthropic backend.
credential:
method: APIKey
apiKey:
secretRef:
name: together-api-key
#Spec
The API this backend speaks, and where it serves it. Defaults to the OpenAI API under /v1, which most providers serve.
The path the backend serves that API under: /v1 for most, /openai/v1 for Groq, and a per-replica path for a Modelplane-composed endpoint, whose cluster gateway distinguishes replicas by path.
The API the backend speaks. A gateway translates an Anthropic request for an OpenAI backend, but not the reverse: an Anthropic backend serves only Anthropic callers, and an OpenAI request routed to it fails.
This backend’s credential, which the gateway attaches on the way out. When the gateway authenticates callers, the caller’s own key never reaches the backend. An endpoint whose credential is missing carries no traffic and reports EndpointReady=False.
Authenticates to the backend with an API key. Required when method is APIKey.
How the gateway authenticates to this backend. APIKey sends a key held in a Secret, in the x-api-key header to a backend whose api.schema is Anthropic, and as a bearer token in the Authorization header otherwise.
The name this backend knows the model by, which a gateway rewrites the request’s model to on the way out. Unset, the caller’s model name passes through unchanged. A caller names a ModelService and gets back whichever model actually served: ask for ml-team/assistant and the response names the model that answered, such as Qwen/Qwen3-8B.
Scheme and host of the backend, with no path: an https origin gets TLS originated to it. A port is only needed for a non-default one. The host must be a name, never an address.