DISCOVER & TEST KONNECT APIS IN REAL TIME WITH INSOMNIA 13 MIGRATE 50% FASTER WITH KONG MIGRATION SERVICES DON'T MISS OUT ON API + AI SUMMIT 2026 | PRICES INCREASE SEPTEMBER 1
  • [Why Kong ](/company/why-kong)Why Kong
  • _API & AI CONNECTIVITY TECHNOLOGIES_
    The Unified API and AI Platform
    []
    API ManagementAI ManagementEvent ManagementMonetization
    Migration Services
    API Advisory Services + Forward Deployed EngineersNEW
    • RUNTIMES
    • [API Gateway ](/products/kong-gateway)API Gateway
    • [AI Gateway HOT](/products/kong-ai-gateway)AI Gateway HOT
    • [Event Gateway ](/products/event-gateway)Event Gateway
    • [Service Mesh ](/products/kong-mesh)Service Mesh
    • [Context Mesh ](/products/kong-konnect/features/context-mesh)Context Mesh
    • [Ingress Controller ](/products/kong-ingress-controller)Ingress Controller
    • [Kong Operator ](/products/kong-operator)Kong Operator
    • CORE SERVICES
    • [MCP Registry NEW](/products/mcp-registry)MCP Registry NEW
    • [API Service Catalog ](/products/kong-konnect/features/api-service-catalog)API Service Catalog
    • [Runtime Management ](/products/kong-konnect/features/runtime-management)Runtime Management
    • [APIOps & Automation ](/products/apiops-automation)APIOps & Automation
    • APPS & AI AGENTS
    • [Developer Portal ](/products/kong-konnect/features/developer-portal)Developer Portal
    • [Usage Billing & Metering $](/products/kong-konnect/features/usage-based-metering-and-billing)Usage Billing & Metering $
    • [Observability ](/products/kong-konnect/features/api-observability)Observability
    • [KAi Agent ](/products/kong-konnect/features/kai-ai-agent)KAi Agent
    DEVELOPER TOOLS
    [Insomnia ](https://insomnia.rest/)Insomnia [Plugins ](https://developer.konghq.com/plugins/)Plugins [Volcano ](https://volcano.dev/)Volcano [Kong MCP ](https://developer.konghq.com/konnect-platform/konnect-mcp/)Kong MCP [Documentation ](https://docs.konghq.com/)Documentation [Open Source ](/community)Open Source
      • FOR PLATFORM TEAMS
      • [Developer Platform ](/solutions/building-developer-platform)Developer Platform
      • [Kubernetes and Microservices ](/solutions/build-on-kubernetes)Kubernetes and Microservices
      • [Observability ](/solutions/observability)Observability
      • [Service Mesh Connectivity ](/solutions/service-mesh-connectivity)Service Mesh Connectivity
      • [Kafka Event Streaming ](/solutions/kafka-stream-api-management)Kafka Event Streaming
      • FOR EXECUTIVES
      • [AI Connectivity ](/ai-connectivity)AI Connectivity
      • [Open Banking ](/solutions/open-banking)Open Banking
      • [Legacy Migration ](/solutions/legacy-api-management-migration)Legacy Migration
      • [Platform Cost Reduction ](/solutions/api-platform-consolidation)Platform Cost Reduction
      • [Kafka Cost Optimization ](/solutions/reduce-kafka-cost)Kafka Cost Optimization
      • [API Monetization ](/solutions/api-monetization)API Monetization
      • [AI Monetization ](/solutions/ai-monetization)AI Monetization
      • [AI FinOps ](/solutions/ai-cost-governance-finops)AI FinOps
      • FOR AI TEAMS
      • [Agent Gateway ](/agent-gateway)Agent Gateway
      • [AI Governance ](/solutions/ai-governance)AI Governance
      • [AI Security ](/solutions/ai-security)AI Security
      • [Token Cost Management ](/solutions/ai-cost-optimization-management)Token Cost Management
      • [Agentic Infrastructure ](/solutions/agentic-ai-workflows)Agentic Infrastructure
      • [MCP Production ](/solutions/mcp-production-and-consumption)MCP Production
      • [MCP Traffic Gateway ](/solutions/mcp-governance)MCP Traffic Gateway
      • FOR DEVELOPERS
      • [Mobile App API Development ](/solutions/mobile-application-api-development)Mobile App API Development
      • [GenAI App Development ](/solutions/power-openai-applications)GenAI App Development
      • [API Gateway for Istio ](/solutions/istio-gateway)API Gateway for Istio
      • [Decentralized Load Balancing ](/solutions/decentralized-load-balancing)Decentralized Load Balancing
      • BY INDUSTRY
      • [Financial Services ](/solutions/financial-services-industry)Financial Services
      • [Healthcare ](/solutions/healthcare)Healthcare
      • [Higher Education ](/solutions/api-platform-for-education-services)Higher Education
      • [Insurance ](/solutions/insurance)Insurance
      • [Manufacturing ](/solutions/manufacturing)Manufacturing
      • [Retail ](/solutions/retail)Retail
      • [Software & Technology ](/solutions/software-and-technology)Software & Technology
      • [Transportation ](/solutions/transportation-and-logistics)Transportation
    NEW
    Kong for Startups
    Apply for $100k credits & 50% off AI Gateway
  • [Pricing ](/pricing)Pricing
      • DOCUMENTATION
      • [Kong Konnect ](https://developer.konghq.com/konnect/)Kong Konnect
      • [Kong Gateway ](https://developer.konghq.com/gateway/)Kong Gateway
      • [Kong Mesh ](https://developer.konghq.com/mesh/)Kong Mesh
      • [Kong AI Gateway ](https://developer.konghq.com/ai-gateway/)Kong AI Gateway
      • [Kong Event Gateway ](https://developer.konghq.com/event-gateway/)Kong Event Gateway
      • [Kong Insomnia ](https://developer.konghq.com/insomnia/)Kong Insomnia
      • [Plugin Hub ](https://developer.konghq.com/plugins/)Plugin Hub
      • EXPLORE
      • [Blog ](/blog)Blog
      • [Learning Center ](/blog/learning-center)Learning Center
      • [eBooks ](/resources/e-book)eBooks
      • [Reports ](/resources/reports)Reports
      • [Demos ](/resources/demos)Demos
      • [Customer Stories ](/customer-stories)Customer Stories
      • [Videos ](/resources/videos)Videos
      • EVENTS
      • [API + AI Summit ](/events/conferences/api-ai-summit)API + AI Summit
      • [Webinars ](/events/webinars)Webinars
      • [User Calls ](/events/user-calls)User Calls
      • [Workshops ](/events/workshops)Workshops
      • [Meetups ](/events/meetups)Meetups
      • [See All Events ](/events)See All Events
      • FOR DEVELOPERS
      • [Get Started ](https://developer.konghq.com/)Get Started
      • [Community ](/community)Community
      • [Certification ](/academy/certification)Certification
      • [Training ](https://education.konghq.com)Training
      • COMPANY
      • [About Us ](/company/about-us)About Us
      • [We're Hiring! ](/company/careers)We're Hiring!
      • [Press Room ](/company/press-room)Press Room
      • [Contact Us ](/company/contact-us)Contact Us
      • [Kong Partner Program ](/partners)Kong Partner Program
      • [Enterprise Support Portal ](https://support.konghq.com/s/)Enterprise Support Portal
      • [Documentation ](https://developer.konghq.com)Documentation
  • [](/search)
  • [Login](https://cloud.konghq.com/login)Login
  • [Book Demo](/contact-sales)Book Demo
  • [Get Started](/products/kong-konnect/register)Get Started
[Blog](/blog)Blog
  • [AI Gateway ](/blog/tag/ai-gateway)AI Gateway
  • [AI Security ](/blog/tag/ai-security)AI Security
  • [AIOps ](/blog/tag/aiops)AIOps
  • [API Security ](/blog/tag/api-security)API Security
  • [API Gateway ](/blog/tag/api-gateway)API Gateway
|
    • [API Management ](/blog/tag/api-management)API Management
    • [API Development ](/blog/tag/api-development)API Development
    • [API Design ](/blog/tag/api-design)API Design
    • [Automation ](/blog/tag/automation)Automation
    • [Service Mesh ](/blog/tag/service-mesh)Service Mesh
    • [Insomnia ](/blog/tag/insomnia)Insomnia
    • [Event Gateway ](/blog/tag/event-gateway)Event Gateway
    • [View All Blogs ](/blog/page/1)View All Blogs
We're Entering the Age of AI Connectivity [Read more](/blog/news/the-age-of-ai-connectivity)Read moreProducts & Agents:
    • [Kong AI Gateway](/products/kong-ai-gateway)Kong AI Gateway
    • [Kong API Gateway](/products/kong-gateway)Kong API Gateway
    • [Kong Event Gateway](/products/event-gateway)Kong Event Gateway
    • [Kong Metering & Billing](/products/kong-konnect/features/usage-based-metering-and-billing)Kong Metering & Billing
    • [Kong Insomnia](/products/kong-insomnia)Kong Insomnia
    • [Kong Konnect](/products/kong-konnect)Kong Konnect
  • [Documentation](https://developer.konghq.com)Documentation
  • [Book Demo](/contact-sales)Book Demo
  1. Home
  2. Blog
  3. Engineering
  4. Intelligent Model Routing: NVIDIA Selects the Model, Kong Routes the Traffic
[AI Gateway](/blog/tag/ai-gateway)AI Gateway
August 11, 2026
4 min read

# Intelligent Model Routing: NVIDIA Selects the Model, Kong Routes the Traffic


Harish Madhavan
Head of Partner Solutions, Kong
Grajesh Chandra
Solution Architect, Kong

Every team running production LLMs has had the same idea: not every request needs the frontier model. Intelligent model routing (or LLM routing)— choosing a model per request on criteria such as task complexity, cost, latency, or quality — enables more efficient model usage.

The idea is easy. Shipping it is not. The moment routing logic goes into the request path, it becomes infrastructure — and now it's holding your provider credentials, needs to comply with security requirements (guardrails, auditing, etc.), and is standing between your applications and every model they depend on, which requires reliable, highly available, and redundant infrastructure. Teams can stall at this point because the application or ML team defining the routing strategy should not also have to own the infrastructure and security boundary. Meanwhile, the platform team responsible for that boundary may not have the application context needed to continually tune routing policies and algorithms.

The solution is to separate model selection from traffic management. Intelligent model routing is a combined system: NVIDIA NeMo Switchyard provides the model selection, and Kong AI Gateway provides the connectivity, governance, and traffic management. Switchyard is a customizable selection library with multiple algorithms that teams can configure according to their own requirements. Kong AI Gateway routes the traffic — dispatching each request to the selected model while managing credentials, guardrails, PII masking, rate limits, and auditing through a scalable, reliable gateway.

## Kong AI Gateway: Connectivity and Governance

As AI adoption scales, applications evolve into complex systems of agents, orchestration layers, and context servers. Infrastructure lags behind — struggling with authentication, cost control, and data security across a provider list that changes every quarter.

Kong AI Gateway is the runtime for LLM traffic management those systems already run through. Its Universal API standardizes interfaces across providers, decoupling applications from provider-specific SDKs and centralizing credential management as part of a broader AI governance strategy. On top of that, the gateway enforces what production actually requires:

  • - **Token-based rate limiting and metering** per team, application, and model - the control that caps AI spend, not just optimizes it
  • - **Credentials in a vault, never in application code**, rotated centrally across every provider
  • - **PII sanitization and prompt guardrails** applied before a request ever leaves your network
  • - **Semantic caching** to eliminate redundant inference entirely
  • - **Model and provider routing, with multi-provider failover, retries, and load balancing** — the traffic layer that survives a provider outage
  • - **One control plane for APIs, AI, MCP, and events**, with RBAC, audit, and analytics across all of it
  • - **Highly scalable, performant dataplanes that support hybrid, self-hosted, and air-gapped environments** — providing teams architectural freedom

That last pair matters more than it looks. Your agents don't only call models — they call REST APIs, MCP servers, and event streams. Governing the LLM hop alone leaves most of the attack surface ungoverned. Further, an enterprise network is complex, with workloads running on-prem, across clouds, etc.

## NVIDIA NeMo Switchyard: Configurable Model Selection

Per-request model selection can account for task complexity and other application-specific requirements. NeMo Switchyard is an open-source Apache 2.0 model selection library from NVIDIA. It makes several selection approaches available through an open-source library that teams can configure and extend for their own use cases.

In this integration, teams can deploy it as a model selection service: Kong provides the information needed to evaluate the configured selection policy, Switchyard returns a target, and Kong maintains control of the request path. Its selection algorithms - including complexity-based approaches such as `stage_router` give teams configurable building blocks for evaluating requests and selecting among model targets. Teams choose the algorithm, candidate models, thresholds, and other criteria appropriate to their applications.

Two configurable capabilities can help teams operate this pattern reliably:

  • - **Configurable fallback behavior: **With classifiers that support a strong_target, teams can configure ambiguous requests to use a designated higher-capability model rather than forcing a lower-confidence routing decision.
  • - **Session persistence:** For multi-turn conversations, teams can retain an initial routing decision across a session when appropriate, reducing redundant routing evaluations and their associated latency.

For detailed setup and advanced routing configurations, refer to the NeMo Switchyard documentation.

## Why Intelligent Model Routing Needs Both

Because model selection is one decision, and routing production AI traffic is many more.

A selection library helps teams automate model choice according to their configured policies. It is not intended to replace gateway capabilities such as credential management, per-team token limits, PII masking, cost attribution, or audit logging. And it doesn't govern the other 90% of your traffic — the APIs, MCP servers, and event streams your agents depend on.

Put the router in the data path instead, and you've added a second hop that sees prompts, holds keys, and carries its own CVE surface — with no SLA behind it. Use it as a configurable decision service, and teams can apply their chosen Switchyard selection strategy while Kong keeps routing the traffic and the gateway boundary intact.

## Architecture: Configurable Selection, Centralized Routing and Governance

*Kong AI Gateway routes every request; NVIDIA NeMo Switchyard supplies the model selection*

The data path begins when an application sends a request to Kong AI Gateway. Kong invokes the Switchyard decision service, which evaluates the request using the selection algorithm and criteria configured by the user and returns a model target. Kong then routes the request to that target, applies the organization’s configured policies, and returns the response to the application. If the decision service is unavailable, Kong can use a customer-configured fallback target, helping preserve request-path availability.

## Scale Your AI Strategy

This pairing succeeds because it respects organizational boundaries. Model selection quality is an ML problem. Routing, security, compliance, infrastructure, and even cost control are platform problems. This architecture lets each team excel without overstepping.

  • - **For the ML team:** Choose and tune selection algorithms, candidate models, thresholds, and policies without embedding that logic in the gateway configuration.
  • - **For the platform team:** Support automated model selection while keeping credentials, guardrails, PII masking, auditing, and other traffic controls centralized in the gateway.
  • - **For the business:** Establish an intelligent model routing strategy that balances cost, latency, and quality. Teams can direct different request types to appropriate models and use gateway telemetry to measure usage and cost by request, application, or team.

Routing requirements and algorithms can evolve independently from gateway infrastructure. With a clear contract, teams configure how model targets are selected through Switchyard, while Kong routes, secures, meters, and logs the resulting traffic. This separation allows routing strategies to change without moving the gateway’s security and governance boundary.[](https://konghq.com/contact-sales)

## Performance and Efficiency Benchmarks

To validate this architecture, Kong tested the integration using [_OpenThoughts-TBLite_](https://github.com/open-thoughts/OpenThoughts-TBLite)_OpenThoughts-TBLite_, a benchmark featuring 20 difficulty-calibrated agent tasks across different categories of easy, medium, hard, and expert.

The results demonstrate how precisely tuning a single `stage_router` parameter can optimize the cost-to-performance ratio. By increasing the confidence threshold from 0.3 to 0.5, escalations to the frontier model dropped from 85% to just 17%. This shift yielded a **43.7% reduction in cost per completed task** while maintaining consistent task completion rates.

In this configuration, the efficient tier utilized GLM-5.2, while the frontier tier ran Claude Opus 4.8. While these preliminary results highlight the potential of intelligent routing, Kong is continuing full evaluations against Terminal-Bench 2 for final verification.

*Learn more: *[_*Kong AI Gateway*_](https://developer.konghq.com/ai-gateway/)_*Kong AI Gateway*_*  ·  *[_*NVIDIA NeMo Switchyard*_](https://github.com/NVIDIA-NeMo/Switchyard)_*NVIDIA NeMo Switchyard*_*  ·  *[_*AI Proxy Advanced*_](https://developer.konghq.com/plugins/ai-proxy-advanced/)_*AI Proxy Advanced*_

*Get started with the API & AI platform: *[_*Book a demo*_](https://konghq.com/contact-sales)_*Book a demo*_

**NVIDIA NeMo Switchyard gives teams configurable algorithms for automating model selection according to their own criteria. Kong AI Gateway routes, secures, meters, and logs. Intelligence advances at ML speed; the security boundary never moves.**

- [AI Gateway](/blog/tag/ai-gateway)AI Gateway- [Governance](/blog/tag/governance)Governance- [LLM](/blog/tag/llm)LLM- [Rate Limiting](/blog/tag/rate-limiting)Rate Limiting

Table of Contents

  • Kong AI Gateway: Connectivity and Governance
  • NVIDIA NeMo Switchyard: Configurable Model Selection
  • Why Intelligent Model Routing Needs Both
  • Architecture: Configurable Selection, Centralized Routing and Governance
  • Scale Your AI Strategy
  • Performance and Efficiency Benchmarks
**Topics**
- [AI Gateway](/blog/tag/ai-gateway)AI Gateway- [Governance](/blog/tag/governance)Governance- [LLM](/blog/tag/llm)LLM- [Rate Limiting](/blog/tag/rate-limiting)Rate Limiting
Harish Madhavan
Head of Partner Solutions, Kong
Grajesh Chandra
Solution Architect, Kong

## Get started with the API & AI platform

[Book Demo](/contact-sales)Book Demo

## step-0

    • Company
    • [About Kong ](/company/about-us)About Kong
    • [Customers ](/customer-stories)Customers
    • [Careers ](/company/careers)Careers
    • [Press ](/company/press-room)Press
    • [Events ](/events)Events
    • [Contact ](/company/contact-us)Contact
    • [Pricing ](/pricing)Pricing
      •    * [Terms](/legal/terms-of-use)
      •    * [Privacy](/legal/privacy-policy)
      •    * [Trust and Compliance](https://trust.konghq.com/)
    • Platform
    • [Kong AI Gateway ](/products/kong-ai-gateway)Kong AI Gateway
    • [Kong Konnect ](/products/kong-konnect)Kong Konnect
    • [Kong Gateway ](/products/kong-gateway)Kong Gateway
    • [Kong Event Gateway ](/products/event-gateway)Kong Event Gateway
    • [Kong Insomnia ](/products/kong-insomnia)Kong Insomnia
    • [Documentation ](https://developer.konghq.com)Documentation
    • [Book Demo ](/contact-sales)Book Demo
    • Compare
    • [AI Gateway Alternatives ](/performance-comparison/ai-gateway-alternatives)AI Gateway Alternatives
    • [Kong vs Apigee ](/performance-comparison/kong-vs-apigee)Kong vs Apigee
    • [Kong vs AWS ](/performance-comparison/kong-vsaws)Kong vs AWS
    • [Kong vs IBM ](/performance-comparison/ibm-api-connect-vs-kong)Kong vs IBM
    • [Kong vs Mulesoft ](/performance-comparison/kong-vs-mulesoft)Kong vs Mulesoft
    • [Kong vs Postman ](/performance-comparison/kong-vs-postman)Kong vs Postman
    • Explore More
    • [Kong for Startups ](/solutions/startup-program)Kong for Startups
    • [Open Banking API Solutions ](/solutions/open-banking)Open Banking API Solutions
    • [API Governance Solutions ](/solutions/api-governance)API Governance Solutions
    • [Istio API Gateway Integration ](/solutions/istio-gateway)Istio API Gateway Integration
    • [Kubernetes API Management ](/solutions/build-on-kubernetes)Kubernetes API Management
    • [API Gateway: Build vs Buy ](/campaign/secure-api-scalability)API Gateway: Build vs Buy
    • Open Source
    • [Kong Gateway ](https://developer.konghq.com/gateway/install/)Kong Gateway
    • [Kuma ](https://kuma.io/)Kuma
    • [Insomnia ](https://insomnia.rest/)Insomnia
    • [Kong Community ](/community)Kong Community

Kong enables the connectivity layer for the agentic era – securely connecting, governing, and monetizing APIs and AI tokens across any model or cloud.

  • English
  • Japanese
  • Frenchcoming soon
  • Spanishcoming soon
  • Germancoming soon
[Everything is 200 OK](https://status.konghq.com/)
© Kong Inc. 2026
Interaction mode