Route Claude CLI traffic through AI Gateway and Hugging Face

Incompatible with
on-prem
Minimum Version
AI Gateway - 2.0
Previous Versions of this page
TL;DR

Create an AI Provider entity to store your Hugging Face token, create an AI Policy that strips fields Claude CLI sends that Hugging Face’s API rejects, create an AI Model entity with an Anthropic-compatible format that routes to Hugging Face through that provider, then point Claude CLI’s ANTHROPIC_BASE_URL at your local AI Gateway endpoint so all LLM requests pass through the gateway for monitoring and control.

Prerequisites

This is a Konnect tutorial and requires a Konnect personal access token.

  1. Create a new personal access token by opening the Konnect PAT page and selecting Generate Token.

  2. Export your token to an environment variable:

    export KONNECT_TOKEN='YOUR_KONNECT_PAT'
  3. Run the AI Gateway quickstart script to automatically provision a control plane and data plane in Kong Konnect, and configure your environment:

    curl -Ls https://get.konghq.com/ai | bash -s -- -k $KONNECT_TOKEN 

This sets up a AI Gateway control plane named ai-quickstart, provisions a local data plane, and prints out the following environment variables export:

export AI_GATEWAY_ID=your-gateway-id
export KONNECT_TOKEN=$KONNECT_TOKEN
export KONNECT_CONTROL_PLANE_NAME=ai-quickstart
export KONNECT_CONTROL_PLANE_URL=https://us.api.konghq.com
export KONNECT_PROXY_URL='http://localhost:8000'

Copy and paste these into your terminal to configure your session.

This tutorial uses kongctl to manage Konnect resources programmatically. We recommend keeping kongctl up to date with the latest version (1.13.0).

  1. Install kongctl from developer.konghq.com/kongctl.
  2. Verify the installation:

    kongctl version
  1. Create a Hugging Face access token with inference permissions.
  2. Export the token as a bearer header value:

    export HUGGINGFACE_AUTH_HEADER='Bearer YOUR_HUGGINGFACE_TOKEN'
  1. Install Claude:
     curl -fsSL https://claude.ai/install.sh | bash
  2. Verify the installation:
     claude --version

Create an AI Model Provider, and AI Model, and an AI Policy

Create an AI Model Provider, and AI Model, and an AI Policy with all of your configuration

  • The AI Model Provider entity defines your connection and stores your authentication credentials.
  • The AI Request Transformer Policy removes extra fields that Hugging Face’s API does not support.
  • The AI Model entity declares which upstream models are available, configures how client requests are routed, and specifies which AI Model Provider to use.

formats: [type: anthropic] accepts requests in Anthropic format even though the upstream is Hugging Face.

kongctl apply -f - --auto-approve --pat "$KONNECT_TOKEN" << 'EOF'
ai_gateway_model_providers:
  - ref: my-huggingface-account
    ai_gateway: !lookup {id: !env AI_GATEWAY_ID}
    name: my-huggingface-account
    display_name: "Hugging Face"
    type: huggingface
    config:
      auth:
        type: basic
        headers:
          - name: Authorization
            value: !secret {source: !env HUGGINGFACE_AUTH_HEADER}
ai_gateway_policies:
  - ref: strip-claude-beta-info
    ai_gateway: !lookup {id: !env AI_GATEWAY_ID}
    name: strip-claude-beta-info
    display_name: "Strip Claude beta info"
    type: request-transformer-advanced
    config:
      remove:
        headers:
          - anthropic-beta
        querystring:
          - beta
        body:
          - output_config
          - context_management
          - mcp_servers
          - container
          - service_tier
          - reasoning_effort
ai_gateway_models:
  - ref: my-huggingface
    ai_gateway: !lookup {id: !env AI_GATEWAY_ID}
    name: my-huggingface
    display_name: "my-huggingface"
    type: model
    formats:
      - type: anthropic
    config:
      route:
        paths:
          - /
        model:
          body_param: model
          values:
            - my-huggingface
    targets:
      - name: deepseek-ai/DeepSeek-V4-Pro-0813
        provider: !ref my-huggingface-account#name
        config:
          type: huggingface
    policies:
      - !ref strip-claude-beta-info#name
    capabilities:
      - generate
EOF

The AI Model Provider uses the following settings:

  • type: huggingface: Specifies that this provider speaks Hugging Face’s Messages API format.
  • config.auth.headers[0].value: !secret {source: !env HUGGINGFACE_AUTH_HEADER}: Loads the API key from your environment at apply time so it is not embedded in the config, and kongctl redacts it in plan and diff output.

The AI Policy uses the following settings:

  • type: request-transformer-advanced: Modifies requests before AI Gateway forwards them upstream.
  • config.remove.headers: Removes the anthropic-beta header.
  • config.remove.querystring: Removes the beta query string parameter.
  • config.remove.body: Removes the output_config, context_management, mcp_servers, container, service_tier, and reasoning_effort body fields.

    Claude Code beta features vary by version and may add other incompatible fields over time. If you still see a 400 error mentioning an unexpected field after applying this Policy, add that field to the appropriate remove list and re-apply.

The AI Model uses the following settings:

  • name/display_name: my-huggingface: The identifier you pass to claude --model. Claude Code uses this, not the upstream target name, to select the model.
  • formats: [type: anthropic]: Declares that this model accepts requests in Anthropic-compatible format, matching what Claude Code sends natively.
  • config.route.paths: [/]: Configures the base path where this model’s Routes are accessible.
  • capabilities: [generate]: Enables text generation. For a model using the anthropic format, generate creates a /messages endpoint matching Anthropic’s native Messages API, so combined with your base path, clients send requests to /v1/messages.
  • policies: Attaches the strip-claude-beta-info policy created in the previous step, so its header and body transformations apply to every request sent through this model.

Run Claude Code

Claude Code’s experimental beta features send fields that Hugging Face rejects even with the AI Policy in place. Disable them, then start a session pointed at your local AI Gateway endpoint:

export CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1
export ANTHROPIC_BASE_URL=http://localhost:8000/

CLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENT=1 \
  claude --allowedTools "WebSearch,Read" --model "my-huggingface"

And ask a question to confirm that requests reach AI Gateway.

Tell me about the Madrid Skylitzes manuscript.

Claude Code might prompt you approve its web search for answering the question. When you select Yes, Claude will produce a full-length response to your request:

The Madrid Skylitzes is a remarkable 12th-century illuminated Byzantine
manuscript that represents one of the most important surviving examples
of medieval historical documentation. Here are the key details:

What it is

The Madrid Skylitzes is the only surviving illustrated manuscript of John
Skylitzes' "Synopsis of Histories" (Σύνοψις Ἱστοριῶν), which chronicles
Byzantine history from 811 to 1057 CE - covering the period from the death
of Emperor Nicephorus I to the deposition of Michael VI.

Artistic Significance

- 574 miniature paintings (with about 100 lost over time)
- Lavishly decorated with gold leaf, vibrant pigments, and intricate
detailing
- Depicts everything from imperial coronations and battles to daily life
in Byzantium
- The only surviving Byzantine illuminated chronicle written in Greek

Unique Collaboration

The manuscript is believed to be the work of 7 different artists from
various backgrounds:
- 4 Italian artists
- 1 English or French artist
- 2 Byzantine artists

Help us make these docs great!

Kong Developer docs are open source. If you find these useful and want to make them better, contribute today!