Modelplane Modelplane docs

InferenceGateway Custom Resource

This document is for an unreleased version of Modelplane.

This document applies to the Modelplane main branch and not to the latest release v0.4.

An InferenceGateway is the front door for inference requests: the only address a caller sees. It speaks the OpenAI and Anthropic APIs, authenticates callers, resolves the model a request names to a ModelService, and forwards to whichever of that service’s endpoints should serve it, translating the request for the chosen backend. A Modelplane can run several, each on an InferenceCluster of its own, one per region to keep a caller’s traffic in its jurisdiction or two in a region to survive losing a cluster. Distributing callers across them is yours to configure.

Concept guide: Set Up the Gateway →

#Metadata

API version
modelplane.ai/v1alpha1
Kind
InferenceGateway
Scope
Cluster
Short names
ig

#Example

Manifest
apiVersion: modelplane.ai/v1alpha1
kind: InferenceGateway
metadata:
  name: eu
spec:
  # The InferenceCluster this gateway runs on, which decides its region and its
  # address. The cluster needs no GPU pools.
  clusterName: gw-gcp-eu
  # Certificates for the DNS names callers use. The names are yours: point your
  # DNS at status.address.
  tls:
    certificateRefs:
      - name: eu-example-com-tls
  auth:
    # How the gateway authenticates callers.
    method: APIKey
    apiKey:
      # Each key in a selected Secret is one caller: the entry's name is the
      # caller's identity, its value is the key.
      secretSelector:
        matchLabels:
          modelplane.ai/inference-keys: "true"
  # The ModelServices this gateway serves. Absent, it serves every one. Scoped
  # here to a region, which is how residency is expressed.
  serviceSelector:
    matchLabels:
      example.org/region: eu

#Spec

# auth optional object

Authenticates callers. Omit it and the gateway authenticates nobody, so anything that can reach the address can invoke any ModelService it serves, and set its own x-modelplane-caller identity on every usage record. Omitting auth is only appropriate behind something that has already established who is calling and that the gateway is reachable only through it; with auth set, the gateway derives the caller header itself and overwrites any a caller sent. Modelplane authenticates callers; it does not authorize them. Every authenticated caller can reach every ModelService this gateway serves, and /v1/models lists them all regardless of caller. To narrow what a caller can reach, narrow the gateway with serviceSelector or run a separate gateway for them.

# apiKey optional object

Authenticates callers by API key. Required when method is APIKey.

# secretSelector required object

Selects Secrets holding caller API keys. Each key in a selected Secret is one caller: the entry’s name is the caller’s identity and its value is the key. The gateway stamps the resolved identity onto every request and every usage record, and never forwards the caller’s key.

# matchLabels required map[string]string
# method required enum: APIKey

How the gateway authenticates callers. APIKey matches the key a caller presents against keys held in Secrets.

# clusterName required string 1–253 chars

The InferenceCluster this gateway runs on, which decides its region and its address. A cluster hosts at most one gateway. The cluster needs no GPU pools: one with none can host a gateway and nothing else. A cluster that serves models can host a gateway too. Immutable. To move a gateway, create one on the new cluster and move callers to its address.

# serviceSelector optional object

Selects the ModelServices this gateway serves, by their labels. Absent, it serves every one. This is how a gateway is scoped: to a region, so an EU service is only reachable through EU gateways; to your public services on an internet-facing front door; or to a named set on a dedicated gateway. These are your labels, under your own prefix.

# matchLabels required map[string]string
# tls optional object

Serves callers over HTTPS, on the DNS names you point at status.address and issue its certificates for. Without it the caller’s hop is unencrypted.

# certificateRefs required object[] 1–8 items
# name required string 1–253 chars

#Status

# address optional string

The address this gateway answers on, and what its DNS names should point at. It is also the target to health check, at /healthz on port 80 over plain HTTP even when the gateway terminates TLS, to decide whether this gateway is in rotation. /healthz answers 200 whenever this gateway’s proxy is running and serving. It says nothing about whether any ModelService is reachable through it, so a gateway with no healthy backend stays in rotation and answers requests with a 503. Read each ModelService’s RoutingReady for that.

# clientCACertificate optional string ≤ 16384 chars

PEM certificate of the CA that signs this gateway’s client certificate. Every InferenceCluster accepts client certificates from it.

# endpoints optional object

The URLs this gateway serves each API at, built from its address. Published only without tls: a gateway serving HTTPS is reached on its DNS names, at https:///v1 for the OpenAI API and https:///anthropic/v1 for Anthropic’s.

# anthropic optional string

Base URL for Anthropic’s Messages API.

# openAI optional string

Base URL for the OpenAI API. A caller sets its SDK’s base_url to this and names a ModelService as the model.