helm upgrade --install kong-operator kong/kong-operator -n kong-system \
--create-namespace \
--version 1.5.0-rapid.2.0 \
--set env.ENABLE_CONTROLLER_KONNECT=true \
--set env.ENABLE_CONTROLLER_AIGATEWAYDATAPLANE=trueCompress agent traffic with Headroom using Kong Operator
Run a single Headroom instance in the cluster with remote compression enabled and a proxy token.
Then create an AIGatewayPolicy of type ai-prompt-compressor with provider: headroom, pointing compressor_url at Headroom’s /v1/compress endpoint.
Keep the AI Policy configuration in a Secret so the proxy token never appears in the custom resource.
Prerequisites
Kong Konnect
If you don’t have a Konnect account, you can get started quickly with our onboarding wizard.
- The following Konnect items are required to complete this tutorial:
- Personal access token (PAT): Create a new personal access token by opening the Konnect PAT page and selecting Generate Token.
-
Set the personal access token as an environment variable:
export KONNECT_TOKEN='YOUR KONNECT TOKEN'
Kong Operator running
Kong AI Gateway features ship in Kong Operator rapid releases. Use Kong Operator 2.4.0-rapid.2.0, chart version
1.5.0-rapid.2.0, or later. Rapid releases are fully supported production releases. For more information, see Rapid releases.
-
Add the Kong Helm charts:
helm repo add kong https://charts.konghq.com helm repo update -
Install Kong Operator using Helm:
If you want cert-manager to issue and rotate the admission and conversion webhook certificates, install cert-manager to your cluster and enable cert-manager integration by passing the following argument while installing, in the next step:
--set global.webhooks.options.certManager.enabled=trueIf you do not enable this, the chart will generate and inject self-signed certificates automatically. We recommend enabling cert-manager to manage the lifecycle of these certificates. Kong Operator needs a certificate authority to sign the certificate for mTLS communication between the control plane and the data plane. This is handled automatically by the Helm chart. If you need to provide a custom CA certificate, refer to the
certificateAuthoritysection in thevalues.yamlof the Helm chart to learn how to create and reference your own CA certificate.
Create a KonnectAPIAuthConfiguration resource
kubectl create namespace kong --dry-run=client -o yaml | kubectl apply -f -
echo '
kind: KonnectAPIAuthConfiguration
apiVersion: konnect.konghq.com/v1alpha1
metadata:
name: konnect-api-auth
namespace: kong
spec:
type: token
token: "'$KONNECT_TOKEN'"
serverURL: us.api.konghq.com
' | kubectl apply -f -OpenAI API key
This tutorial uses OpenAI:
- Create an OpenAI account.
- Get an API key.
-
Export your API key as an environment variable:
export OPENAI_API_KEY='YOUR_OPENAI_API_KEY'
Kong AI Gateway resources
Create an AI Gateway control plane, an OpenAI AI Model Provider, a gpt-4o-mini AI Model, and a data plane.
-
Store your OpenAI API key in a Secret:
kubectl create secret generic openai-credentials -n kong \ --from-literal=token="Bearer ${OPENAI_API_KEY}" kubectl label secret openai-credentials -n kong konghq.com/secret=true -
Create the AI Gateway resources:
echo ' apiVersion: konnect.konghq.com/v1alpha1 kind: KonnectAIGateway metadata: name: my-ai-gateway-cp namespace: kong spec: apiSpec: name: my-ai-gateway-cp displayName: My AI Gateway konnect: authRef: name: konnect-api-auth --- apiVersion: aiconfiguration.konghq.com/v1alpha1 kind: AIGatewayModelProvider metadata: name: openai-provider namespace: kong spec: aiGatewayRef: type: namespacedRef namespacedRef: name: my-ai-gateway-cp apiSpec: type: openai openai: name: openai-provider displayName: OpenAI config: auth: headers: - name: Authorization value: type: secretRef secretRef: name: openai-credentials key: token --- apiVersion: aiconfiguration.konghq.com/v1alpha1 kind: AIGatewayModel metadata: name: gpt-4o-mini namespace: kong spec: aiGatewayRef: type: namespacedRef namespacedRef: name: my-ai-gateway-cp apiSpec: type: model model: name: gpt-4o-mini displayName: GPT-4o Mini enabled: Enabled formats: - type: openai capabilities: - generate config: route: paths: - /v1 targets: - name: gpt-4o-mini provider: name: openai-provider config: type: openai openai: upstreamURL: https://api.openai.com/v1/chat/completions --- apiVersion: aigateway.konghq.com/v1alpha1 kind: AIGatewayDataPlane metadata: name: my-ai-gateway-dp namespace: kong spec: controlPlaneRef: type: konnectNamespacedRef konnectNamespacedRef: name: my-ai-gateway-cp deployment: replicas: 1 network: services: ingress: type: LoadBalancer ports: - name: http port: 8000 targetPort: 8000 ' | kubectl apply -f - -
Wait for the resources to be ready:
kubectl wait konnectaigateway/my-ai-gateway-cp aigatewaymodelprovider/openai-provider aigatewaymodel/gpt-4o-mini -n kong \ --for=condition=Programmed=True \ --timeout=10m kubectl wait aigatewaydataplane/my-ai-gateway-dp -n kong \ --for=condition=Ready=True \ --timeout=10m -
Export the data plane address:
export AIGW_HOST=$(kubectl get service my-ai-gateway-dp-ingress -n kong \ -o jsonpath='{.status.loadBalancer.ingress[0].ip}') echo $AIGW_HOST
Headroom compresses agent traffic: tool results, large JSON payloads, search output, and logs. With the AI Prompt Compressor Policy set to provider: headroom, the AI Gateway data plane sends each request’s messages to Headroom and forwards the compressed messages to the LLM. Headroom decides how much to compress.
This guide uses the following recommended layout:
- One Headroom
Deploymentwith exactly one replica in the same namespace as the data plane. Headroom keeps its sessions in memory in a single process, so the data plane must always reach the same instance. - Headroom’s
/v1/compressendpoint opened to in-cluster callers and protected by a proxy token. - The AI Prompt Compressor Policy configuration, including the token, stored in a Kubernetes Secret and referenced from the
AIGatewayPolicy.
Deploy Headroom
-
Create a proxy token. The data plane sends it to Headroom on every call:
export HEADROOM_TOKEN=$(openssl rand -hex 16) kubectl create secret generic headroom-token -n kong \ --from-literal=token="$HEADROOM_TOKEN" -
Deploy Headroom and its Service:
echo ' apiVersion: apps/v1 kind: Deployment metadata: name: headroom namespace: kong spec: replicas: 1 selector: matchLabels: app: headroom template: metadata: labels: app: headroom spec: containers: - name: headroom image: ghcr.io/headroomlabs-ai/headroom:latest args: ["--host", "0.0.0.0", "--port", "8787"] env: - name: HEADROOM_COMPRESS_ALLOW_REMOTE value: "1" - name: HEADROOM_PROXY_TOKEN valueFrom: secretKeyRef: name: headroom-token key: token ports: - containerPort: 8787 --- apiVersion: v1 kind: Service metadata: name: headroom namespace: kong spec: selector: app: headroom ports: - port: 8787 targetPort: 8787 ' | kubectl apply -f -By default, Headroom only answers
/v1/compressfor loopback callers and returns404to everyone else.HEADROOM_COMPRESS_ALLOW_REMOTE=1lets the data plane pod call it, andHEADROOM_PROXY_TOKENrejects any caller without the token with401. -
Wait for Headroom to be ready:
kubectl rollout status deployment/headroom -n kong --timeout=5m
Record a baseline
Headroom compresses tool output, JSON, code, and logs. It doesn’t compress plain user prose by default, and it leaves blocks under about 500 tokens alone. A short chat message will show no change, so the test request carries a large tool result: 400 structured log entries returned by a fetch_logs tool call.
-
Build the request:
jq -n '[range(400) | {id: ., level: (if . % 7 == 0 then "ERROR" else "INFO" end), service: "checkout", latency_ms: ((. * 37) % 900), msg: (if . % 7 == 0 then "upstream timeout" else "request completed" end)}]' > logs.json jq -n --rawfile logs logs.json '{ model: "gpt-4o-mini", messages: [ {role: "user", content: "Which service produced errors in these logs?"}, {role: "assistant", content: null, tool_calls: [{id: "call_1", type: "function", function: {name: "fetch_logs", arguments: "{\"window\":\"1h\"}"}}]}, {role: "tool", tool_call_id: "call_1", content: $logs} ]}' > request.json -
Send the request before you create the AI Prompt Compressor Policy and note the prompt tokens OpenAI reports:
curl -s http://$AIGW_HOST:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d @request.json | jq '.usage.prompt_tokens'In this example, this request should use about 16,900 prompt tokens.
Create the AI Prompt Compressor Policy
-
Store the AI Policy configuration in a Secret. Kong Operator reads the whole AI Policy
configfrom the Secret key, as YAML or JSON, so the proxy token never appears in theAIGatewayPolicy. Thekonghq.com/secret=truelabel lets Kong Operator read the Secret:cat > headroom-config.yaml <<EOF provider: headroom compressor_url: http://headroom.kong.svc.cluster.local:8787/v1/compress timeout: 45000 keepalive_timeout: 60000 stop_on_error: true log_text_data: false headroom: proxy_token: ${HEADROOM_TOKEN} ssl_verify: false session_id_headers: - x-claude-code-session-id - x-claude-code-agent-id - thread-id - session-id EOF kubectl create secret generic headroom-policy-config -n kong \ --from-file=config.yaml=headroom-config.yaml kubectl label secret headroom-policy-config -n kong konghq.com/secret=true -
Create the AI Policy and apply it to every AI Model:
echo ' apiVersion: aiconfiguration.konghq.com/v1alpha1 kind: AIGatewayPolicy metadata: name: headroom-compressor namespace: kong spec: aiGatewayRef: type: namespacedRef namespacedRef: name: my-ai-gateway-cp apiSpec: name: headroom-compressor displayName: Headroom compressor type: ai-prompt-compressor enabled: Enabled global: Enabled config: type: secretRef secretRef: name: headroom-policy-config key: config.yaml ' | kubectl apply -f - -
Wait for the AI Policy to be reconciled:
kubectl wait aigatewaypolicy/headroom-compressor -n kong \ --for=condition=Programmed=True \ --timeout=5m
Validate
-
Send the same request again, this time with a session header. Use a fresh session ID, because Headroom compresses a repeated tool result more aggressively within the same session:
curl -s http://$AIGW_HOST:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -H "x-claude-code-session-id: $(uuidgen)" \ -d @request.json | jq '.usage.prompt_tokens'The prompt token count drops from about 16,900 in the baseline to about 4,500.
-
To confirm the savings from Headroom’s side, forward Headroom’s port and read its statistics:
kubectl port-forward -n kong deployment/headroom 8787:8787 > /dev/null 2>&1 & curl -s http://localhost:8787/stats -H "Authorization: Bearer $HEADROOM_TOKEN" \ | jq '{compress_calls: .requests.by_provider.compress, tokens_saved: .tokens.saved}'
FAQs
How do Headroom sessions affect compression?
The data plane sends every call to Headroom with a session ID, and Headroom compresses differently depending on what it has already seen in that session. In the example in this guide:
- Without the AI Prompt Compressor Policy: about 16,900 prompt tokens. The LLM received the raw JSON tool result.
- First request in a session: about 4,500 prompt tokens. The LLM received all 400 rows as a compact table.
- Same tool result again in the same session: about 700 prompt tokens. The LLM received a heavily reduced summary with most rows removed.
Compression can remove information the model needs. For requests where the model must see every value exactly, apply the AI Prompt Compressor Policy to specific AI Models instead of globally, and leave out the AI Models that need the full data.
What happens when Headroom can’t be reached?
stop_on_error decides what happens when Headroom can’t be reached, rejects the token, or times out:
truefails the request with500andfailed to compress prompts: Failure when calling the Headroom compression service. Use it while you roll out, so configuration mistakes are visible immediately.falseforwards the original, uncompressed messages. Switch to it once compression is working, because with a global AI Policy a Headroom outage would otherwise fail every request through the gateway.