Skip to content

Model routing

Model routing decides which model answers a request. Rig supports it at three levels:

  • Between agents: a cheap classifier (or an embedding search) picks one of several specialized agents, and that agent answers.
  • Inside one agent: an agent carries several models under labels, and a hook picks a label before every model call — for example a fast model to plan and call tools, and a stronger one to write the final answer.
  • From configuration: a string like "anthropic/claude-sonnet-5-5" from a config file or a user setting becomes a model at runtime.

All three rely on one fact: an Agent is not generic over its provider. Every model is accepted as impl Into<DynModel<Completion>>, so agents and models from different providers share one type and can sit in the same HashMap, Vec, or struct field.

A router agent classifies the request into one word from a fixed list, and the matching specialist answers. The router’s work is trivial, so it can run on a small model:

flowchart TD
User[User] -->|Query| Router[Router agent]
Router -->|rust| Rust[Rust agent]
Router -->|maths| Maths[Maths agent]
Rust --> Response
Maths --> Response
use std::collections::HashMap;
let openai = OpenAI::from_env()?;
let anthropic = Anthropic::from_env()?;
// Specialists from different providers share the `Agent` type.
let routes: HashMap<&str, Agent> = HashMap::from([
(
"rust",
AgentBuilder::new(anthropic.completion(anthropic::CLAUDE_SONNET_5_5))
.preamble("You are an expert Rust programmer.")
.build(),
),
(
"maths",
AgentBuilder::new(openai.completion(openai::GPT_5_5))
.preamble("You are a mathematician who explains each step.")
.build(),
),
]);
let router = AgentBuilder::new(openai.completion(openai::GPT_5_MINI))
.preamble(
"Classify the user's question. Reply with exactly one word from: rust, maths. \
No other text.",
)
.build();
let question = "How do I use async with Rust?";
let topic = router.prompt(question).await?.output();
let Some(agent) = routes.get(topic.trim()) else {
anyhow::bail!("no route for {topic:?}");
};
println!("{}", agent.prompt(question).await?.output());

Always handle the no-match case: a model can answer with something outside the list.

An LLM classifier is non-deterministic. For steadier matching, embed a description and example questions for each route, then pick the route closest to the incoming question by vector similarity. See Embeddings for how this works.

/// One route. The description and examples are embedded; `name` is the key.
#[derive(Embed, Clone, Debug, Default, Serialize, Deserialize, Eq, PartialEq)]
struct Route {
name: String,
#[embed]
description: String,
#[embed]
examples: Vec<String>,
}
let embedding_model = OpenAI::from_env()?
.embedding(openai::TEXT_EMBEDDING_3_SMALL, None)
.erase();
let routes = vec![
Route {
name: "rust".into(),
description: "Programming and software development in Rust".into(),
examples: vec![
"How do I write an async function in Rust?".into(),
"Why does the borrow checker reject this code?".into(),
],
},
Route {
name: "maths".into(),
description: "Mathematics, calculations and equations".into(),
examples: vec!["Solve this equation".into(), "What is 15% of 200?".into()],
},
];
let embeddings = EmbeddingsBuilder::new(embedding_model.clone())
.documents(routes)?
.build()
.await?;
let index = InMemoryVectorStore::from_documents(embeddings).index(embedding_model);
let request = VectorSearchRequest::builder()
.query("How do I use async with Rust?")
.samples(1)
.build();
let route = index
.top_n::<Route>(request)
.await?
.into_iter()
.next()
.map_or_else(|| "general".to_owned(), |hit| hit.document.name);
println!("route: {route}");

Use the route name to look up an agent as in the previous section. Keep a general fallback for questions that match nothing well; with real data you may also want a minimum hit.score below which you fall back.

Routing between agents picks one agent for the whole request. Sometimes the choice belongs to each model call instead: in a tool-using run, the calls that decide which tool to use can go to a fast model, and the call that writes the final answer to a stronger one. An agent can carry several models for this:

  • AgentBuilder::new(model) registers model under the label default. AgentBuilder::named_model(label, model) does the same under a label you choose.
  • .model_route(label, model) registers another model the run may select.
  • An AgentHook with on_model_select runs before every model call and returns ModelSelectionAction::select(label), ::continue_run() (keep the current candidate), or ::stop(reason) (end the run before the call).

The models can come from different providers; history is kept in Rig’s provider-neutral form, so a tool call made by one model is answered and continued by another.

/// Plan and call tools on the fast model; answer on the strong one.
#[derive(Clone)]
struct FastThenStrong;
impl AgentHook for FastThenStrong {
fn on_model_select(
&self,
context: &HookContext,
_event: ModelSelection<'_>,
) -> ModelSelectionAction {
// `turn()` is the one-based index of the model call in this run.
if context.turn() == 1 {
ModelSelectionAction::select("fast")
} else {
ModelSelectionAction::select("strong")
}
}
}
let agent = AgentBuilder::named_model("fast", OpenAI::from_env()?.completion(openai::GPT_5_MINI))
.model_route(
"strong",
Anthropic::from_env()?.completion(anthropic::CLAUDE_OPUS_5_5),
)
// .tool(search_tool) ...
.add_hook(FastThenStrong)
.build();
let answer = agent
.prompt("Research the question, then write a careful answer.")
.max_turns(4)
.await?;
println!("{answer}");

The hook receives a ModelSelection describing the pending call: the prompt and history it will send, the run’s default_model, the selected_model chosen by earlier hooks, and the previous_model that made the last call. Routing can therefore depend on the conversation itself, not just the turn number. Some details of the contract:

  • Hooks added with AgentBuilder::add_hook apply to every run of the agent; hooks added on a prompt (agent.prompt(p).add_hook(h)) apply to that run only.
  • Several selection hooks run in registration order; each sees the previous one’s choice, the last selection wins, and a stop ends the run.
  • Selection happens again before every model call, including retries and the calls after tool results, so a run can switch models mid-way. A call already in flight never switches.
  • Selecting a label that was never registered fails the run.

The runtime_model_routing example runs this two-model tool round trip end to end with scripted models, without API keys.

Without a routing hook, the agent uses its default model. You can change it for one run or for the agent value itself:

let mut agent = AgentBuilder::named_model("fast", OpenAI::from_env()?.completion(openai::GPT_5_MINI))
.model_route("strong", OpenAI::from_env()?.completion(openai::GPT_5_5))
.build();
// One run: start from a registered label, or from a model value.
let hard = agent.prompt("Prove that sqrt(2) is irrational.").using_model("strong").await?;
let other = agent
.prompt("Say hello.")
.using_model_value(Anthropic::from_env()?.completion(anthropic::CLAUDE_HAIKU_4_5))
.await?;
// This agent value from now on: a registered label, or a new model.
agent.set_model_label("strong");
agent.set_model(Anthropic::from_env()?.completion(anthropic::CLAUDE_SONNET_5_5));

with_model_label and with_model are the by-value forms. Setting the default only changes the starting candidate: routing hooks still run and may choose another label. To force one model, leave routing hooks off, or add a final hook that always selects it. set_model follows value semantics — clones of the agent keep their own default.

Choosing the provider and model at runtime

Section titled “Choosing the provider and model at runtime”

When the provider comes from configuration — a settings file, an environment variable, a user’s choice in your UI — build the model from a string with the provider registry in rig::providers::registry. connect takes a vendor/model reference and an API key and returns a DynModel<Completion>:

use rig::providers::registry;
// e.g. read from a config file
let model_name = "anthropic/claude-sonnet-5-5";
let api_key = std::env::var("ANTHROPIC_API_KEY")?;
let model = registry::connect(model_name, api_key)?;
let agent = AgentBuilder::new(model).preamble("You are a helpful assistant.").build();
println!("{}", agent.prompt("Hello!").await?.output());

The reference may also name a protocol family, which matters for vendors that speak more than one API: vendor[/format]:model, for example "deepseek:deepseek-chat" or "deepseek/openai:deepseek-chat". connect_with takes an HTTP client to send through instead of the shared one.

To read credentials from each provider’s usual environment variables instead of passing a key, parse a ProviderRef. It is a credential-free, serializable recipe, so it is also what you store in configuration:

use rig::providers::registry::ProviderRef;
let reference: ProviderRef = "deepseek:deepseek-chat".parse()?;
let model = reference.completion_model()?; // reads DEEPSEEK_API_KEY
// Serializes as the string "deepseek/openai:deepseek-chat".
let stored = serde_json::to_string(&reference)?;

A model built this way is a normal DynModel, so it can be an agent’s default, a model_route, or a using_model_value for one run.

rig::catalog::Catalog lists known models with their context window, output limit, input modalities, tool and structured-output support, and prices. A router can use it to pick a model that fits a request, and connect accepts a catalog entry directly:

use rig::catalog::Catalog;
use rig::providers::registry;
let spec = Catalog::builtin()
.resolve("openai/gpt-5.5")
.ok_or_else(|| anyhow::anyhow!("model not in catalog"))?;
println!(
"{}: {:?} token context, tools: {}",
spec.display_name, spec.context_window, spec.tools
);
if let Some(pricing) = spec.pricing {
println!("${} / ${} per million input/output tokens", pricing.input, pricing.output);
}
let model = registry::connect(spec, std::env::var("OPENAI_API_KEY")?)?;

Catalog::from_json reads an override file in the same shape, and Catalog::merge lays it over the built-in data.