AI Glossary · LLM Integration
Model Router
A component that picks the right model per request based on cost, latency, capability or policy — often small model first, large model on fallback.
Definition
What is Model Router?
Model Router is a component that picks the right model per request based on cost, latency, capability or policy — often small model first, large model on fallback.
- Category
- LLM Integration
- Glossary set
- 11 related terms
- Audience
- Enterprise AI leaders
Why does Model Router matter in enterprise AI?
Model Router matters in production LLM integration because it affects reliability, latency, observability, and how AI workflows connect to enterprise systems.
Related terms in LLM Integration
- API Gateway (LLM)
- A managed proxy that routes model calls, enforces quotas, redacts PII, logs prompts, and applies policy across multiple LLM providers.
- Batch Inference
- Running many model calls asynchronously at lower cost — used for backfills, document processing and offline analytics.
- Caching (Prompt/Response)
- Storing repeat prompts or key/value tensors to cut latency and cost; supported natively by most frontier providers.
- Context Assembly
- The pipeline that gathers system prompt, retrieved chunks, tool schemas and history into a single request within the context window.
- Fallback Routing
- Failing over to a secondary model or provider when the primary is slow, rate-limited or degraded — table stakes for production.
- Rate Limits / Quotas
- Provider-imposed ceilings on requests, tokens or concurrency. Design pattern: token-bucket clients, retries with jitter, quota headroom.