Passthrough format in AI Gateway

Related Documentation
Minimum Version
AI Gateway - 2.2
Incompatible with
on-prem
Tags
#ai

What is passthrough?

In an AI Model, the passthrough format forwards traffic to an upstream AI service without transforming it. AI Gateway doesn’t transform the request or response body, doesn’t enforce any Content-Type, and doesn’t validate the payload against any schema. Request and response bodies reach their destination byte-for-byte. This allows you to put AI Gateway in front of any AI service and keep all the capabilities that don’t depend on the payload shape:

  • Upstream provider authentication, including AWS Signature Version 4, OAuth 2.0, and API keys
  • AI Consumer authentication, identity, and access control lists
  • Request-count rate limiting
  • Logging and metrics

Use cases

Passthrough is useful for AI endpoints that fall outside the providers that AI Gateway supports natively. For example:

  • Self-hosted model servers with proprietary request and response schemas.
  • Vendor preview APIs whose schema is still changing ahead of general availability.
  • Non-LLM AI endpoints, such as custom computer vision or speech APIs, that don’t map to any capability.

If you’re migrating from the preserve route type in the AI Proxy Advanced plugin, passthrough is its replacement. See Migrate from the preserve route type.

Configure a passthrough AI Model

Set formats[].type to passthrough on the AI Model. The following example proxies a self-hosted Triton Inference Server instance that exposes a schema AI Gateway doesn’t recognize:

kongctl apply -f - --auto-approve --pat "$KONNECT_TOKEN" << 'EOF'
ai_gateway_models:
  - ref: custom-inference
    ai_gateway: !lookup {id: !env AI_GATEWAY_ID}
    name: custom-inference
    display_name: "custom-inference"
    type: model
    formats:
      - type: passthrough
    config:
      route:
        paths:
          - /custom-inference
    targets:
      - name: my-model
        provider: my-triton-account
        config:
          type: openai
          upstream_url: http://my-triton-server.internal:8000/v2/models/my-model/infer
    policies: []
    capabilities:
      - generate
EOF

In this example:

  • formats: [type: passthrough] disables format translation and body parsing for the whole AI Model.
  • config.route.paths: [/custom-inference] sets the base path clients send requests to. Passthrough adds no capability-specific suffix, so clients can call any path under this prefix.
  • targets[].config.upstream_url sets the destination, including the path. See Upstream paths.
  • targets[].config.type selects openai here for its API-key authentication behavior, not because the upstream is OpenAI. Passthrough never checks this value against the actual payload shape, so pick whichever supported type’s auth mechanism matches your upstream, such as openai for a plain API key or bedrock for AWS Signature Version 4.
  • provider: my-triton-account references an AI Model Provider that contains the upstream connection and credentials, the same as any other AI Model.

Credentials still come from the AI Model Provider, and AI Gateway signs each upstream request as it normally would. Passthrough skips format transformation, not authentication, so AWS Signature Version 4, OAuth 2.0, and API key auth all work. Set targets[].allow_auth_override to true if you want request-level credentials to take precedence instead.

Upstream paths

The upstream_url on a target carries both the host and the path. The way AI Gateway builds the upstream request depends on whether that URL includes a path. The following examples assume an AI Model with config.route.paths: [/custom-inference], a target whose provider defaults to https://api.openai.com, and a client request sent to /custom-inference/v1/chat/completions:

upstream_url

Upstream request path

Example

Set, with a path other than / AI Gateway uses the configured path as-is and discards the client’s request path. upstream_url: http://my-server.internal:8000/v2/models/my-model/infer sends every request to http://my-server.internal:8000/v2/models/my-model/infer, no matter what path the client used.
Set, with no path or only / AI Gateway forwards the client’s request path, stripping the AI Model’s base path. upstream_url: http://my-server.internal:8000/ sends the request to http://my-server.internal:8000/v1/chat/completions, the client’s path with /custom-inference stripped.
Not set The provider’s default host supplies the host, and AI Gateway forwards the client’s request path. No upstream_url sends the request to https://api.openai.com/v1/chat/completions, the provider’s default host plus the client’s path with /custom-inference stripped.

Observability in passthrough mode

The log phase still runs, so HTTP Log, File Log, OpenTelemetry, and Prometheus receive metadata for every passthrough request. The following is available regardless of payload shape:

  • Request and response latency, in milliseconds
  • HTTP status code
  • Request and response byte counts
  • AI Consumer identity, when an authentication Policy is attached
  • AI Model and AI Model Provider identifiers

Token usage and cost

Whether AI Gateway can report token counts and cost for a passthrough request depends on your target’s provider:

  • anthropic, bedrock, cohere, gemini, or huggingface token counts and cost are extracted reliably, for both streaming and buffered responses.
  • With an OpenAI-compatible custom or self-hosted server, such as vLLM or Ollama, that return an OpenAI-shaped usage object, token counts can populate. However, this isn’t guaranteed, since passthrough forwards whatever shape your upstream actually returns.
  • With anything else, for example Amazon SageMaker, Llama 2, and other custom or self-hosted servers with their own response shape, AI Gateway has no way to locate token counts, and usage stays empty.

Cost calculation and token-based rate limiting only work when token counts are available. The input_cost and output_cost settings on a target apply only to requests whose token counts AI Gateway could read, and AI Rate Limiting Advanced has nothing to meter when extraction fails.

Before you rely on passthrough usage data for billing or quota enforcement, confirm that your target’s provider has a native adapter, or that its response includes an OpenAI-shaped usage object, and that token counts appear for real traffic, streaming and buffered.

Streaming detection

AI Gateway detects a streaming response from the request itself through:

  • A "stream": true field in the body
  • An Accept: text/event-stream header
  • A known streaming path for providers that signal streaming through the URL instead of the body

If none of these match, AI Gateway buffers the response instead of streaming it. The client still gets the full, unmodified response, but all at once instead of incrementally. If your upstream signals streaming some other way, send "stream": true in the request body or set the Accept header yourself.

Migrate from the preserve route type

kongctl convert ai-gateway doesn’t have any special handling for route_type: preserve. To migrate, recreate the AI Model manually using Configure a passthrough AI Model as a starting point, and use the following sections as a manual field-by-field reference for what preserve used to do.

Field mapping

AI Proxy Advanced setting

AI Model setting

config.targets[].route_type: preserve formats[].type: passthrough
model.options.upstream_url targets[].config.upstream_url
model.options.upstream_path No equivalent. Merge the path into targets[].config.upstream_url.
config.targets[].auth.allow_override targets[].allow_auth_override

Client path forwarding

Under preserve, path fallback used the raw incoming request path and ignored the Route’s strip_path setting. If you relied on the full raw path reaching your upstream under preserve, set an explicit upstream_url with a fixed path on the new AI Model so client-path forwarding doesn’t apply at all. Otherwise, verify the path arriving upstream is still what your backend expects once you’ve recreated the target.

Migration steps

  • Don’t rely on kongctl convert ai-gateway for preserve targets: it warns and skips them, without producing an output for them. Recreate each one as a new AI Model manually.
  • Set formats[].type to passthrough on the new AI Model.
  • Merge any upstream_path value into targets[].config.upstream_url.
  • Set upstream_url explicitly on Azure and Databricks targets, which don’t have a default host. Azure fails at request time without it; Databricks fails configuration validation instead.
  • Verify the path arriving upstream, now that client-path forwarding strips the AI Model’s base path.
  • Split any plugin instance that mixed preserve with other route types into separate AI Models.
  • Remove any dependency on model aliasing, semantic load balancing, the realtime capability, and generation parameters.

Limitations

  • AI Models using the passthrough format don’t support format normalization or anything that rewrites the request.
  • Token accounting depends on whether AI Gateway can make sense of your upstream’s actual response shape; see Token usage and cost.
  • Guardrail Policies (AI Prompt Guard, AI Semantic Prompt/Response Guard, AI AWS/Azure/GCP/Lakera/Custom Guardrails) aren’t supported under passthrough, since there’s no parsed body for them to read.
  • Policies that write into a specific body field (AI Prompt Decorator, AI Prompt Compressor, AI RAG Injector, AI Prompt Template, AI Semantic Cache) aren’t supported. Passthrough doesn’t parse the body into a shape these Policies can read or write. MCP and agent-to-agent traffic isn’t affected either way.
  • The following capabilities are not compatible with passthrough mode: * Format normalization: No translation between the OpenAI shape and a provider shape, in either direction. * Semantic load balancing: An AI Model can’t combine the semantic algorithm with a passthrough target. Configuration is rejected at validation time for every provider, because the semantic algorithm needs a parsed body to match against each target’s semantic_description. Other load balancing algorithms, including round-robin, priority, and lowest-latency, work normally. * Realtime: Passthrough is explicitly rejected for the realtime/generation category. Every other capability category is allowed under passthrough.

FAQs

No. If you need both passthrough and a native format, declare two AI Models.

Policies that operate on raw bytes (AI Request Transformer, AI Response Transformer, AI Sanitizer in PII mode) and request-count rate limiting always work. Guardrails that read prompt or completion content (AI Prompt Guard, AI Semantic Prompt/Response Guard, AI AWS/Azure/GCP/Lakera/Custom Guardrails) and Policies that write into a specific body field (AI Prompt Decorator, AI Prompt Compressor, AI RAG Injector, AI Prompt Template, AI Semantic Cache) aren’t supported: passthrough doesn’t parse the body into a shape these Policies can read or write. MCP and agent-to-agent traffic isn’t affected either way.

Help us make these docs great!

Kong Developer docs are open source. If you find these useful and want to make them better, contribute today!