Mock AI Model responses without calling a provider

Incompatible with
on-prem
Minimum Version
AI Gateway - 2.0
TL;DR

Attach the Mocking Policy to an AI Model and give it an OpenAPI spec with example responses for the Model’s route. Requests to the Model return the examples, and no tokens are spent.

Prerequisites

This is a Konnect tutorial and requires a Konnect personal access token.

  1. Create a new personal access token by opening the Konnect PAT page and selecting Generate Token.

  2. Export your token to an environment variable:

    export KONNECT_TOKEN='YOUR_KONNECT_PAT'
  3. Run the AI Gateway quickstart script to automatically provision a control plane and data plane in Kong Konnect, and configure your environment:

    curl -Ls https://get.konghq.com/ai | bash -s -- -k $KONNECT_TOKEN 

This sets up a AI Gateway control plane named ai-quickstart, provisions a local data plane, and prints out the following environment variables export:

export AI_GATEWAY_ID=your-gateway-id
export KONNECT_TOKEN=$KONNECT_TOKEN
export KONNECT_CONTROL_PLANE_NAME=ai-quickstart
export KONNECT_CONTROL_PLANE_URL=https://us.api.konghq.com
export KONNECT_PROXY_URL='http://localhost:8000'

Copy and paste these into your terminal to configure your session.

This tutorial uses kongctl to manage Konnect resources programmatically. We recommend keeping kongctl up to date with the latest version (1.16.0).

  1. Install kongctl from developer.konghq.com/kongctl.
  2. Verify the installation:

    kongctl version

The Mocking Policy answers requests before they reach the AI Model Provider, so the credential is never sent upstream. Export a placeholder value:

export FAKE_AUTH_HEADER="Bearer not-a-real-key"

Create the AI Model Provider, AI Model, and Mocking Policy

Create an AI Model Provider, an AI Model, and a Mocking Policy that returns one of two example responses for POST /v1/chat/completions.

The Policy matches the full request path, so the spec path /v1/chat/completions includes the AI Model’s route path /v1.

kongctl apply -f - --auto-approve --pat "$KONNECT_TOKEN" << 'EOF'
ai_gateway_policies:
  - ref: mock-chat-completions
    ai_gateway: !lookup name:ai-quickstart
    name: mock-chat-completions
    display_name: "Mock chat completions"
    type: mocking
    enabled: true
    global: false
    config:
      api_specification: |
        openapi: 3.0.0
        info:
          title: Mock chat completions
          version: 1.0.0
        paths:
          /v1/chat/completions:
            post:
              responses:
                '200':
                  description: A chat completion.
                  content:
                    application/json:
                      examples:
                        Hello:
                          value:
                            id: chatcmpl-mock-1
                            object: chat.completion
                            model: my-gpt-4o
                            choices:
                              - index: 0
                                finish_reason: stop
                                message:
                                  role: assistant
                                  content: "This is a mocked response."
                            usage:
                              prompt_tokens: 9
                              completion_tokens: 6
                              total_tokens: 15
                        Goodbye:
                          value:
                            id: chatcmpl-mock-2
                            object: chat.completion
                            model: my-gpt-4o
                            choices:
                              - index: 0
                                finish_reason: stop
                                message:
                                  role: assistant
                                  content: "Goodbye from the mock."
ai_gateway_model_providers:
  - ref: fake-openai
    ai_gateway: !lookup name:ai-quickstart
    name: fake-openai
    display_name: "fake-openai"
    type: openai
    config:
      auth:
        type: basic
        headers:
          - name: Authorization
            value: !env FAKE_AUTH_HEADER
ai_gateway_models:
  - ref: my-gpt-4o
    ai_gateway: !lookup name:ai-quickstart
    name: my-gpt-4o
    display_name: "my-gpt-4o"
    type: model
    formats:
      - type: openai
    config:
      route:
        paths:
          - /v1
        model:
          body_param: model
          values:
            - my-gpt-4o
    targets:
      - name: gpt-4o
        provider: fake-openai
        config:
          type: openai
    policies:
      - !ref mock-chat-completions#name
    capabilities:
      - generate
EOF

Validate

Send requests to the configured AI Model and check the X-Kong-Mocking-Plugin: true response header, which confirms the Policy answered instead of an AI Model Provider.

Cleanup

To clean up all AI Gateway resources created in this guide, run:

curl -Ls https://get.konghq.com/ai | bash -s -- -d

Help us make these docs great!

Kong Developer docs are open source. If you find these useful and want to make them better, contribute today!