Compress agent traffic with Headroom using Kong Operator

Incompatible with
on-prem
Related Documentation
Minimum Version
Kong Operator - 2.4 AI Gateway - 2.2
TL;DR

Run a single Headroom instance in the cluster with remote compression enabled and a proxy token. Then create an AIGatewayPolicy of type ai-prompt-compressor with provider: headroom, pointing compressor_url at Headroom’s /v1/compress endpoint. Keep the AI Policy configuration in a Secret so the proxy token never appears in the custom resource.

Prerequisites

If you don’t have a Konnect account, you can get started quickly with our onboarding wizard.

  1. The following Konnect items are required to complete this tutorial:
    • Personal access token (PAT): Create a new personal access token by opening the Konnect PAT page and selecting Generate Token.
  2. Set the personal access token as an environment variable:

    export KONNECT_TOKEN='YOUR KONNECT TOKEN'

Kong AI Gateway features ship in Kong Operator rapid releases. Use Kong Operator 2.4.0-rapid.2.0, chart version 1.5.0-rapid.2.0, or later. Rapid releases are fully supported production releases. For more information, see Rapid releases.

  1. Add the Kong Helm charts:

    helm repo add kong https://charts.konghq.com
    helm repo update
  2. Install Kong Operator using Helm:

    helm upgrade --install kong-operator kong/kong-operator -n kong-system \
      --create-namespace \
      --version 1.5.0-rapid.2.0 \
      --set env.ENABLE_CONTROLLER_KONNECT=true \
      --set env.ENABLE_CONTROLLER_AIGATEWAYDATAPLANE=true

    If you want cert-manager to issue and rotate the admission and conversion webhook certificates, install cert-manager to your cluster and enable cert-manager integration by passing the following argument while installing, in the next step:

    --set global.webhooks.options.certManager.enabled=true

    If you do not enable this, the chart will generate and inject self-signed certificates automatically. We recommend enabling cert-manager to manage the lifecycle of these certificates. Kong Operator needs a certificate authority to sign the certificate for mTLS communication between the control plane and the data plane. This is handled automatically by the Helm chart. If you need to provide a custom CA certificate, refer to the certificateAuthority section in the values.yaml of the Helm chart to learn how to create and reference your own CA certificate.

kubectl create namespace kong --dry-run=client -o yaml | kubectl apply -f -
echo '
kind: KonnectAPIAuthConfiguration
apiVersion: konnect.konghq.com/v1alpha1
metadata:
  name: konnect-api-auth
  namespace: kong
spec:
  type: token
  token: "'$KONNECT_TOKEN'"
  serverURL: us.api.konghq.com
' | kubectl apply -f -

This tutorial uses OpenAI:

  1. Create an OpenAI account.
  2. Get an API key.
  3. Export your API key as an environment variable:

    export OPENAI_API_KEY='YOUR_OPENAI_API_KEY'

Create an AI Gateway control plane, an OpenAI AI Model Provider, a gpt-4o-mini AI Model, and a data plane.

  1. Store your OpenAI API key in a Secret:

    kubectl create secret generic openai-credentials -n kong \
      --from-literal=token="Bearer ${OPENAI_API_KEY}"
    kubectl label secret openai-credentials -n kong konghq.com/secret=true
  2. Create the AI Gateway resources:

    echo '
    apiVersion: konnect.konghq.com/v1alpha1
    kind: KonnectAIGateway
    metadata:
      name: my-ai-gateway-cp
      namespace: kong
    spec:
      apiSpec:
        name: my-ai-gateway-cp
        displayName: My AI Gateway
      konnect:
        authRef:
          name: konnect-api-auth
    ---
    apiVersion: aiconfiguration.konghq.com/v1alpha1
    kind: AIGatewayModelProvider
    metadata:
      name: openai-provider
      namespace: kong
    spec:
      aiGatewayRef:
        type: namespacedRef
        namespacedRef:
          name: my-ai-gateway-cp
      apiSpec:
        type: openai
        openai:
          name: openai-provider
          displayName: OpenAI
          config:
            auth:
              headers:
                - name: Authorization
                  value:
                    type: secretRef
                    secretRef:
                      name: openai-credentials
                      key: token
    ---
    apiVersion: aiconfiguration.konghq.com/v1alpha1
    kind: AIGatewayModel
    metadata:
      name: gpt-4o-mini
      namespace: kong
    spec:
      aiGatewayRef:
        type: namespacedRef
        namespacedRef:
          name: my-ai-gateway-cp
      apiSpec:
        type: model
        model:
          name: gpt-4o-mini
          displayName: GPT-4o Mini
          enabled: Enabled
          formats:
            - type: openai
          capabilities:
            - generate
          config:
            route:
              paths:
                - /v1
          targets:
            - name: gpt-4o-mini
              provider:
                name: openai-provider
              config:
                type: openai
                openai:
                  upstreamURL: https://api.openai.com/v1/chat/completions
    ---
    apiVersion: aigateway.konghq.com/v1alpha1
    kind: AIGatewayDataPlane
    metadata:
      name: my-ai-gateway-dp
      namespace: kong
    spec:
      controlPlaneRef:
        type: konnectNamespacedRef
        konnectNamespacedRef:
          name: my-ai-gateway-cp
      deployment:
        replicas: 1
      network:
        services:
          ingress:
            type: LoadBalancer
            ports:
              - name: http
                port: 8000
                targetPort: 8000
    ' | kubectl apply -f -
  3. Wait for the resources to be ready:

    kubectl wait konnectaigateway/my-ai-gateway-cp aigatewaymodelprovider/openai-provider aigatewaymodel/gpt-4o-mini -n kong \
      --for=condition=Programmed=True \
      --timeout=10m
    kubectl wait aigatewaydataplane/my-ai-gateway-dp -n kong \
      --for=condition=Ready=True \
      --timeout=10m
  4. Export the data plane address:

    export AIGW_HOST=$(kubectl get service my-ai-gateway-dp-ingress -n kong \
      -o jsonpath='{.status.loadBalancer.ingress[0].ip}')
    echo $AIGW_HOST

Headroom compresses agent traffic: tool results, large JSON payloads, search output, and logs. With the AI Prompt Compressor Policy set to provider: headroom, the AI Gateway data plane sends each request’s messages to Headroom and forwards the compressed messages to the LLM. Headroom decides how much to compress.

This guide uses the following recommended layout:

  • One Headroom Deployment with exactly one replica in the same namespace as the data plane. Headroom keeps its sessions in memory in a single process, so the data plane must always reach the same instance.
  • Headroom’s /v1/compress endpoint opened to in-cluster callers and protected by a proxy token.
  • The AI Prompt Compressor Policy configuration, including the token, stored in a Kubernetes Secret and referenced from the AIGatewayPolicy.

Deploy Headroom

  1. Create a proxy token. The data plane sends it to Headroom on every call:

    export HEADROOM_TOKEN=$(openssl rand -hex 16)
    kubectl create secret generic headroom-token -n kong \
      --from-literal=token="$HEADROOM_TOKEN"
  2. Deploy Headroom and its Service:

    echo '
    apiVersion: apps/v1
    kind: Deployment
    metadata:
      name: headroom
      namespace: kong
    spec:
      replicas: 1
      selector:
        matchLabels:
          app: headroom
      template:
        metadata:
          labels:
            app: headroom
        spec:
          containers:
            - name: headroom
              image: ghcr.io/headroomlabs-ai/headroom:latest
              args: ["--host", "0.0.0.0", "--port", "8787"]
              env:
                - name: HEADROOM_COMPRESS_ALLOW_REMOTE
                  value: "1"
                - name: HEADROOM_PROXY_TOKEN
                  valueFrom:
                    secretKeyRef:
                      name: headroom-token
                      key: token
              ports:
                - containerPort: 8787
    ---
    apiVersion: v1
    kind: Service
    metadata:
      name: headroom
      namespace: kong
    spec:
      selector:
        app: headroom
      ports:
        - port: 8787
          targetPort: 8787
    ' | kubectl apply -f -

    By default, Headroom only answers /v1/compress for loopback callers and returns 404 to everyone else. HEADROOM_COMPRESS_ALLOW_REMOTE=1 lets the data plane pod call it, and HEADROOM_PROXY_TOKEN rejects any caller without the token with 401.

  3. Wait for Headroom to be ready:

    kubectl rollout status deployment/headroom -n kong --timeout=5m

Record a baseline

Headroom compresses tool output, JSON, code, and logs. It doesn’t compress plain user prose by default, and it leaves blocks under about 500 tokens alone. A short chat message will show no change, so the test request carries a large tool result: 400 structured log entries returned by a fetch_logs tool call.

  1. Build the request:

    jq -n '[range(400) | {id: ., level: (if . % 7 == 0 then "ERROR" else "INFO" end),
      service: "checkout", latency_ms: ((. * 37) % 900),
      msg: (if . % 7 == 0 then "upstream timeout" else "request completed" end)}]' > logs.json
    
    jq -n --rawfile logs logs.json '{
      model: "gpt-4o-mini",
      messages: [
        {role: "user", content: "Which service produced errors in these logs?"},
        {role: "assistant", content: null, tool_calls: [{id: "call_1", type: "function",
          function: {name: "fetch_logs", arguments: "{\"window\":\"1h\"}"}}]},
        {role: "tool", tool_call_id: "call_1", content: $logs}
      ]}' > request.json
  2. Send the request before you create the AI Prompt Compressor Policy and note the prompt tokens OpenAI reports:

    curl -s http://$AIGW_HOST:8000/v1/chat/completions \
      -H "Content-Type: application/json" \
      -d @request.json | jq '.usage.prompt_tokens'

    In this example, this request should use about 16,900 prompt tokens.

Create the AI Prompt Compressor Policy

  1. Store the AI Policy configuration in a Secret. Kong Operator reads the whole AI Policy config from the Secret key, as YAML or JSON, so the proxy token never appears in the AIGatewayPolicy. The konghq.com/secret=true label lets Kong Operator read the Secret:

    cat > headroom-config.yaml <<EOF
    provider: headroom
    compressor_url: http://headroom.kong.svc.cluster.local:8787/v1/compress
    timeout: 45000
    keepalive_timeout: 60000
    stop_on_error: true
    log_text_data: false
    headroom:
      proxy_token: ${HEADROOM_TOKEN}
      ssl_verify: false
      session_id_headers:
        - x-claude-code-session-id
        - x-claude-code-agent-id
        - thread-id
        - session-id
    EOF
    
    kubectl create secret generic headroom-policy-config -n kong \
      --from-file=config.yaml=headroom-config.yaml
    kubectl label secret headroom-policy-config -n kong konghq.com/secret=true
  2. Create the AI Policy and apply it to every AI Model:

    echo '
    apiVersion: aiconfiguration.konghq.com/v1alpha1
    kind: AIGatewayPolicy
    metadata:
      name: headroom-compressor
      namespace: kong
    spec:
      aiGatewayRef:
        type: namespacedRef
        namespacedRef:
          name: my-ai-gateway-cp
      apiSpec:
        name: headroom-compressor
        displayName: Headroom compressor
        type: ai-prompt-compressor
        enabled: Enabled
        global: Enabled
        config:
          type: secretRef
          secretRef:
            name: headroom-policy-config
            key: config.yaml
    ' | kubectl apply -f -
  3. Wait for the AI Policy to be reconciled:

    kubectl wait aigatewaypolicy/headroom-compressor -n kong \
      --for=condition=Programmed=True \
      --timeout=5m

Validate

  1. Send the same request again, this time with a session header. Use a fresh session ID, because Headroom compresses a repeated tool result more aggressively within the same session:

    curl -s http://$AIGW_HOST:8000/v1/chat/completions \
      -H "Content-Type: application/json" \
      -H "x-claude-code-session-id: $(uuidgen)" \
      -d @request.json | jq '.usage.prompt_tokens'

    The prompt token count drops from about 16,900 in the baseline to about 4,500.

  2. To confirm the savings from Headroom’s side, forward Headroom’s port and read its statistics:

    kubectl port-forward -n kong deployment/headroom 8787:8787 > /dev/null 2>&1 &
    curl -s http://localhost:8787/stats -H "Authorization: Bearer $HEADROOM_TOKEN" \
      | jq '{compress_calls: .requests.by_provider.compress, tokens_saved: .tokens.saved}'

FAQs

The data plane sends every call to Headroom with a session ID, and Headroom compresses differently depending on what it has already seen in that session. In the example in this guide:

  • Without the AI Prompt Compressor Policy: about 16,900 prompt tokens. The LLM received the raw JSON tool result.
  • First request in a session: about 4,500 prompt tokens. The LLM received all 400 rows as a compact table.
  • Same tool result again in the same session: about 700 prompt tokens. The LLM received a heavily reduced summary with most rows removed.

Compression can remove information the model needs. For requests where the model must see every value exactly, apply the AI Prompt Compressor Policy to specific AI Models instead of globally, and leave out the AI Models that need the full data.

stop_on_error decides what happens when Headroom can’t be reached, rejects the token, or times out:

  • true fails the request with 500 and failed to compress prompts: Failure when calling the Headroom compression service. Use it while you roll out, so configuration mistakes are visible immediately.
  • false forwards the original, uncompressed messages. Switch to it once compression is working, because with a global AI Policy a Headroom outage would otherwise fail every request through the gateway.

Help us make these docs great!

Kong Developer docs are open source. If you find these useful and want to make them better, contribute today!