FAQ

Frequently asked questions about the AI Router

Basics & Overview

What is the AI Router?

The AI Router automatically selects and routes to the best LLM for each request, helping you save costs and ensure high reliability. It optimizes for quality, cost, and latency while providing automatic fallbacks and full OpenAI API compatibility. Our customers typically save >60% on their LLM costs while maintaining or improving response quality.

What are the different usage modes?

We offer three modes:

1) Model Selection - returns the best model name so you execute the request yourself,
2) Full Privacy - determines the best LLM using embeddings to keep message content private, and
3) Model Routing - automatically calls the selected model on your own provider keys and returns the response.

Which models are used for routing?

AI Router routes across the models of the providers you connect: OpenAI, Anthropic, Google Gemini, Mistral AI, Meta, Qwen, Microsoft. The exact list for your team is on the dashboard Models page. AI Router is built for business customers, and any model a customer needs is added on request. If a model you need is not available yet, contact us. See Providers and Models for more details.

Who is the AI Router best suited for?

The AI Router is ideal for businesses that want to optimize their LLM usage across multiple providers. It's particularly valuable for companies that need to balance cost, performance, and reliability, handle varying workloads, or require privacy-focused solutions. Common use cases include customer support automation, content generation, data analysis, and any applications where consistent LLM performance is crucial, especially RAG applications.

What does BYOK mean for me?

Bring your own key: routed requests execute on your own provider accounts, so model costs stay on your provider bill and you keep your existing provider contracts and data agreements. Connect the keys of the providers you want to use in the dashboard; see the provider keys documentation.

Features & Capabilities

How does the model routing work?

The AI Router has multiple modules with different approaches built in to predict the expected quality, cost, and latency for each individual request for all the supported models. It then selects the best model based on your preferences.

How does the full privacy mode work?

In full privacy mode, the AI Router makes model selections based on embeddings of your messages rather than the raw content. These embeddings can be generated automatically by our SDK or you can generate them yourself and provide them manually. This means your message content stays private and never leaves your infrastructure. Full privacy mode is a plan feature; see the pricing page.

Can I customize the routing preferences?

Yes. Your team's routing policy (default model plus quality, cost, and latency preferences) is set during onboarding and can be changed at any time under Settings → Routing in the dashboard. The enabled models are managed on the Models page. Individual API keys can override the policy, and per-request parameters override everything else. Enterprise customers can also get custom router optimization for their specific use cases.

What happens if an LLM provider like OpenAI is down?

Every routing decision ranks several candidates. If the selected model fails, the AI Router automatically falls back to the next best candidates of the same decision without you even noticing. With your own provider keys connected, fallbacks stay within the providers you can execute on.

Do you support tool calling and structured outputs?

Yes. Tool definitions, response formats and all other OpenAI chat completion parameters are passed through to the selected model. When a request contains tools, only models with tool support are considered for it.

What happens when a model is deprecated?

Deprecated models keep working until their sunset date so your integration does not break, but they are penalized in the ranking and marked in the dashboard. Move to a successor model before the sunset date; removed models are no longer routable. New models of your enabled providers can be adopted automatically, see Providers and Models.

Technical Integration

How do I integrate the AI Router with my existing OpenAI implementation?

Simply change your API endpoint to use api.airouter.io instead of api.openai.com. No other code changes are needed as we maintain full OpenAI API compatibility. See the docs for more details.

Is there an SDK?

You don't need an SDK to use the AI Router for model routing as it is fully compatible with any OpenAI integration, such as the openai lib and langchain. For model selection and full privacy mode, we provide SDKs for Python and NodeJS to make the integration even easier. Check out our SDK documentation for more details.

Is there an HTTP API for model selection without an SDK?

Yes. The Decision API returns a ranked list of candidate models with predicted quality, cost and latency plus a decision id. You execute the completion yourself and report the actual usage through the usage report API.

Do you support streaming responses?

Yes, for model routing we fully support streaming responses just like the OpenAI API.

Do you have a rate limit?

Currently, we don't enforce specific rate limits. If a model's rate limit is reached, our router automatically selects the next best available model to handle your request, ensuring continuous service. This means you can send requests at the rate you need without worrying about hard limits. We plan to introduce configurable rate limits in the future for better predictability and cost control.

Can I restrict what an API key may do?

Yes. When creating an API key you can limit it to a subset of your enabled models and give it its own quality, cost and latency preferences. These overrides are fixed for the lifetime of the key; create a new key to change them.

Performance & Optimization

How do you ensure the quality of model recommendations?

We continuously evaluate and benchmark LLM performance across different types of tasks and contexts. Our router uses multiple sophisticated algorithms to predict model performance for each specific request, considering factors like content, length, and complexity.

I am more interested in low-latency LLMs, is this also possible?

Yes, prioritize latency in your routing policy to favor faster models. Per-request parameters can override that configuration when needed; see the weighting documentation.

What is the latency overhead of using AI Router?

Our best model identification typically adds less than 100ms of latency, this could increase for very large requests. However, this overhead is usually compensated or even outperformed on average by automatically selecting faster models for your requests when speed is a priority.

Pricing & Plans

Do I always save money compared to OpenAI?

Our customers typically save >60% on average compared to using a frontier model for everything. However, the actual savings depend on your priorities - if you choose to prioritize quality or speed over cost, the savings might be lower. Configure the balance between cost, quality, and speed in your routing policy; per-request parameters override that configuration when needed.

Do you offer a free trial?

You can see the AI Router in action in the playground: it compares the routed model with a reference model on sample prompts, and after signing in you can run a limited number of live comparisons on your own prompts. To validate AI Router on your real workloads, book a pilot; the pilot fee is credited towards your subscription.

How do I subscribe?

Plans are set up together with us. Book a pilot call from the pricing page or contact us, and we activate your plan for your team.

What's included in the enterprise plan?

The enterprise plan includes a committed routed-token volume with a custom overage rate and can include custom model routing optimization, support for private models, dedicated support, custom SLAs, and integration assistance. Contact us to discuss your specific needs.

How is my usage calculated?

Your quota is measured in routed tokens. A routed token is any input or output token from a request covered by an AI Router routing decision, whether we proxy the request or you execute it yourself. We notify you as you approach your plan's quota, and overage is billed rather than blocking your traffic. Read the usage and billing documentation, or see our pricing page for plan details.

How do you measure routed-token usage?

We use three sources. For requests we proxy using your provider keys, we measure the usage directly. When you execute a request yourself after a routing decision, you report its usage for that decision through the usage report API. If no usage is reported for a decision, we estimate it.

What happens when usage data is missing?

Billing is based on measured usage. Where a routing decision has no usage report, we estimate its usage conservatively from your measured usage profile, as fixed in your contract. Reporting actual usage avoids estimates altogether.

Why should I report usage for requests I execute?

Reporting selection usage gives you value beyond accurate billing: your savings and opportunity reports are only complete when usage for every routing decision is reported back.

Do I need to pay for the underlying model costs?

Yes. AI Router uses your own provider keys (BYOK), so model usage remains on your provider bill. Your plan price covers the AI Router routing fee.

Security & Privacy

How do you handle API keys and security?

We use industry-standard encryption for all API keys and sensitive data. Provider keys are validated against the provider before they are stored and are never displayed again. In routing mode, we proxy your requests to the selected model providers while maintaining all security best practices.

Do you store any of our prompt data?

We only store usage statistics for your dashboards and for improving our product. Usage reports contain token counters and latency only, never prompt or response content. In full privacy mode, we only see embeddings, not the actual content of your messages.

Can I use my own model infrastructure, e.g. my Azure OpenAI deployment or AWS Bedrock models?

Yes, you can always use the AI Router with your own model infrastructure by using our model selection mode - when you receive the best model you can then call the model yourself. See the privacy modes documentation for examples. In the enterprise plan we can also automatically call private models in model routing mode.

Account & Support

Can I cancel my subscription?

Monthly plans can be cancelled at any time; annual plans run for the agreed contract year. Contact us to cancel or change your plan.

Where can I find my invoices?

Invoices are issued per your plan agreement. If you need a copy, contact us at support@airouter.io.

Can I upgrade or downgrade my plan?

Yes, you can change your plan at any time; contact us and we switch your team to the new plan. Your routed-token quota changes with the plan.

Can I have multiple users on my account?

Yes, you can add team members to your team/organization.

Where can I find documentation?

You can find the full documentation here: Documentation

Do you offer technical support?

We offer email support for all plans and dedicated E-Mail and Slack support for enterprise customers.