# InferenceGateway

Source: /reference/inferencegateways/

An InferenceGateway is the front door for inference requests: the only address a caller sees. It speaks the OpenAI and Anthropic APIs, authenticates callers, resolves the model a request names to a ModelService, and forwards to whichever of that service's endpoints should serve it, translating the request for the chosen backend.
A Modelplane can run several, each on an InferenceCluster of its own, one per region to keep a caller's traffic in its jurisdiction or two in a region to survive losing a cluster. Distributing callers across them is yours to configure.

Apply instances as `apiVersion: modelplane.ai/v1alpha1`, `kind: InferenceGateway`.

[Concept guide: Set Up the Gateway](/platform/inference-gateway/index.md)

## Example

```yaml
apiVersion: modelplane.ai/v1alpha1
kind: InferenceGateway
metadata:
  name: eu
spec:
  # The InferenceCluster this gateway runs on, which decides its region and its
  # address. The cluster needs no GPU pools.
  clusterName: gw-gcp-eu
  # Certificates for the DNS names callers use. The names are yours: point your
  # DNS at status.address.
  tls:
    certificateRefs:
      - name: eu-example-com-tls
  auth:
    # How the gateway authenticates callers.
    method: APIKey
    apiKey:
      # Each key in a selected Secret is one caller: the entry's name is the
      # caller's identity, its value is the key.
      secretSelector:
        matchLabels:
          modelplane.ai/inference-keys: "true"
  # The ModelServices this gateway serves. Absent, it serves every one. Scoped
  # here to a region, which is how residency is expressed.
  serviceSelector:
    matchLabels:
      example.org/region: eu
```

## Definition

The CompositeResourceDefinition this reference is generated from, with the complete OpenAPI schema, validation rules, and defaults:

```yaml
apiVersion: apiextensions.crossplane.io/v2
kind: CompositeResourceDefinition
metadata:
  name: inferencegateways.modelplane.ai
spec:
  group: modelplane.ai
  names:
    categories: [crossplane, modelplane, platform]
    kind: InferenceGateway
    plural: inferencegateways
    shortNames: [ig]
  scope: Cluster
  versions:
  - name: v1alpha1
    served: true
    referenceable: true
    additionalPrinterColumns:
    - name: CLUSTER
      type: string
      jsonPath: .spec.clusterName
    - name: ADDRESS
      type: string
      jsonPath: .status.address
    schema:
      openAPIV3Schema:
        description: >-
          An InferenceGateway is the front door for inference requests: the only
          address a caller sees. It speaks the OpenAI and Anthropic APIs,
          authenticates callers, resolves the model a request names to a
          ModelService, and forwards to whichever of that service's endpoints
          should serve it, translating the request for the chosen backend.

          A Modelplane can run several, each on an InferenceCluster of its own,
          one per region to keep a caller's traffic in its jurisdiction or two in
          a region to survive losing a cluster. Distributing callers across them
          is yours to configure.
        type: object
        required: [spec]
        properties:
          spec:
            type: object
            required: [clusterName]
            properties:
              clusterName:
                type: string
                description: >-
                  The InferenceCluster this gateway runs on, which decides its
                  region and its address. A cluster hosts at most one gateway.

                  The cluster needs no GPU pools: one with none can host a gateway
                  and nothing else. A cluster that serves models can host a gateway
                  too.

                  Immutable. To move a gateway, create one on the new cluster and
                  move callers to its address.
                minLength: 1
                maxLength: 253
                x-kubernetes-validations:
                # Changing the cluster would move the gateway's address.
                - rule: "self == oldSelf"
                  message: spec.clusterName is immutable.
              tls:
                type: object
                description: >-
                  Serves callers over HTTPS, on the DNS names you point at
                  status.address and issue its certificates for. Without it the
                  caller's hop is unencrypted.
                required: [certificateRefs]
                properties:
                  certificateRefs:
                    type: array
                    description: >-
                      Secrets holding the gateway's certificates, of type
                      kubernetes.io/tls, in the same namespace as this
                      Modelplane's other gateway Secrets. Modelplane copies them
                      to the gateway's cluster.

                      The gateway presents whichever certificate matches the
                      name a caller asked for, so one certificate can cover
                      several names or each can have its own.
                    minItems: 1
                    maxItems: 8
                    x-kubernetes-list-type: map
                    x-kubernetes-list-map-keys: [name]
                    items:
                      type: object
                      required: [name]
                      properties:
                        name:
                          type: string
                          minLength: 1
                          maxLength: 253
              auth:
                type: object
                description: >-
                  Authenticates callers. Omit it and the gateway authenticates
                  nobody, so anything that can reach the address can invoke any
                  ModelService it serves, and set its own x-modelplane-caller
                  identity on every usage record. Omitting auth is only
                  appropriate behind something that has already established who
                  is calling and that the gateway is reachable only through it;
                  with auth set, the gateway derives the caller header itself and
                  overwrites any a caller sent.

                  Modelplane authenticates callers; it does not authorize them.
                  Every authenticated caller can reach every ModelService this
                  gateway serves, and /v1/models lists them all regardless of
                  caller. To narrow what a caller can reach, narrow the gateway
                  with serviceSelector or run a separate gateway for them.
                required: [method]
                x-kubernetes-validations:
                - rule: "has(self.apiKey) == (self.method == 'APIKey')"
                  message: spec.auth.apiKey must be set when spec.auth.method is APIKey, and only then.
                properties:
                  method:
                    type: string
                    description: >-
                      How the gateway authenticates callers. APIKey matches the
                      key a caller presents against keys held in Secrets.
                    enum: [APIKey]
                  apiKey:
                    type: object
                    description: >-
                      Authenticates callers by API key. Required when method is
                      APIKey.
                    required: [secretSelector]
                    properties:
                      secretSelector:
                        type: object
                        description: >-
                          Selects Secrets holding caller API keys. Each key in a
                          selected Secret is one caller: the entry's name is the
                          caller's identity and its value is the key.

                          The gateway stamps the resolved identity onto every
                          request and every usage record, and never forwards the
                          caller's key.
                        required: [matchLabels]
                        properties:
                          matchLabels:
                            type: object
                            additionalProperties:
                              type: string
                              maxLength: 63
                            minProperties: 1
                            maxProperties: 16
              serviceSelector:
                type: object
                description: >-
                  Selects the ModelServices this gateway serves, by their
                  labels. Absent, it serves every one.

                  This is how a gateway is scoped: to a region, so an EU service
                  is only reachable through EU gateways; to your public services
                  on an internet-facing front door; or to a named set on a
                  dedicated gateway. These are your labels, under your own
                  prefix.
                required: [matchLabels]
                properties:
                  matchLabels:
                    type: object
                    additionalProperties:
                      type: string
                      maxLength: 63
                    minProperties: 1
                    maxProperties: 16
          status:
            type: object
            properties:
              address:
                type: string
                description: >-
                  The address this gateway answers on, and what its DNS names
                  should point at. It is also the target to health check, at
                  /healthz on port 80 over plain HTTP even when the gateway
                  terminates TLS, to decide whether this gateway is in rotation.

                  /healthz answers 200 whenever this gateway's proxy is running
                  and serving. It says nothing about whether any ModelService is
                  reachable through it, so a gateway with no healthy backend
                  stays in rotation and answers requests with a 503. Read each
                  ModelService's RoutingReady for that.
              clientCACertificate:
                type: string
                maxLength: 16384
                description: >-
                  PEM certificate of the CA that signs this gateway's client
                  certificate. Every InferenceCluster accepts client
                  certificates from it.
              endpoints:
                type: object
                description: >-
                  The URLs this gateway serves each API at, built from its
                  address. Published only without tls: a gateway serving HTTPS is
                  reached on its DNS names, at https://<name>/v1 for the OpenAI
                  API and https://<name>/anthropic/v1 for Anthropic's.
                properties:
                  openAI:
                    type: string
                    description: >-
                      Base URL for the OpenAI API. A caller sets its SDK's
                      base_url to this and names a ModelService as the model.
                  anthropic:
                    type: string
                    description: Base URL for Anthropic's Messages API.
```
