Local Models
Rig can run entirely on your own hardware. There are two kinds of local setup:
- A local server that Rig talks to over HTTP: Ollama or llama.cpp. Both are built into
rigand need no extra feature. - In-process inference, with no server at all: Candle for completions and fastembed for embeddings. Both are companion crates behind a feature on
rig.
Whichever you pick, the model plugs into AgentBuilder, EmbeddingsBuilder and the vector stores like any hosted model.
Ollama
Section titled “Ollama”Ollama::new() talks to the daemon at http://localhost:11434 with no authentication. Ollama::from_env() reads the optional OLLAMA_API_BASE_URL and OLLAMA_API_KEY for a remote or secured daemon.
Ollama serves two chat APIs, and Rig has a model for each:
client.completion(id)uses Ollama’s OpenAI-compatible/v1/chat/completions.client.native_completion(id)uses Ollama’s native/api/chat. Only this route sends model parameters such asnum_ctx, and it maps thereasoningoption to Ollama’sthink.
use rig::prelude::*;use rig::providers::ollama::{self, Ollama};
let client = Ollama::new();
let agent = AgentBuilder::new(client.completion(ollama::QWEN3)) .preamble("You are a concise assistant.") .build();
println!("{}", agent.prompt("What is Rust's borrow checker?").await?.output());Ollama-only fields are typed in ollama::extension::OllamaOptions. keep_alive goes to both routes; the model parameters (num_ctx, top_k, min_p, repeat_penalty, num_gpu, num_thread, …) only to the native route, and the /v1 route refuses them rather than dropping them:
use rig::prelude::*;use rig::providers::ollama::extension::{KeepAlive, OllamaOptions};use rig::providers::ollama::{self, Ollama};
let options = OllamaOptions::default() .num_ctx(16_384) .keep_alive(KeepAlive::duration("10m"));
let agent = AgentBuilder::new(Ollama::new().native_completion(ollama::QWEN3)) .provider_option(options) .build();Ollama also serves embeddings through /api/embed. Pass the vector width, or None to use the width Rig knows for that model:
use rig::providers::ollama::{self, Ollama};
let embedder = Ollama::new().embedding(ollama::NOMIC_EMBED_TEXT, None);client.list_models().await? lists the models the daemon has pulled. Runnable examples: rag_ollama and vector_search_ollama.
llama.cpp
Section titled “llama.cpp”The llamacpp provider talks to llama-server’s OpenAI-compatible API, by default at http://localhost:8080/v1. A server started without --api-key needs no key, so pass an empty one:
use rig::prelude::*;use rig::providers::llamacpp;
let client = llamacpp::new("");let agent = AgentBuilder::new(client.chat(llamacpp::LLAMA_CPP)) .preamble("You are a concise assistant.") .build();llamacpp::from_env() reads LLAMACPP_API_KEY (required) and LLAMACPP_API_BASE_URL (optional) for a secured or remote server. For an unsecured server elsewhere, set the base URL on a configuration:
use rig::providers::openai::{OpenAIConfig, wire::LLAMACPP};
let client = OpenAIConfig::with_key(&LLAMACPP, "") .with_base_url("http://gpu-box:8080/v1") .client();llamacpp::LLAMA_CPP is the conventional id for a server that hosts one model; for a multi-model router, use an id from client.list_models().await?. The same client serves embedding(id, None) and rerank(id) when the server enables them. llama-server treats a specific-function tool choice as auto, so Rig refuses that tool choice instead of sending it.
Candle
Section titled “Candle”rig::candle (feature candle) runs models in your process on the CPU with Candle. You load the model files yourself and hand Rig the bytes; the crate does no filesystem or network access.
[dependencies]rig = { version = "0.44.0", features = ["candle"] }use rig::candle::{CandleModel, ModelData};use rig::prelude::*;
let model = CandleModel::from_safetensors_async(ModelData { config: std::fs::read("./model/config.json")?, tokenizer: std::fs::read("./model/tokenizer.json")?, weights: std::fs::read("./model/model.safetensors")?,}).await?;
let agent = AgentBuilder::new(model.completion()) .preamble("You are a concise assistant.") .build();println!("{}", agent.prompt("Explain ownership briefly.").await?.output());Only validated checkpoints load; anything else is rejected rather than guessed at:
- Llama 3 instruct: one unsharded safetensors checkpoint.
- SmolLM2-360M-Instruct: Q4_K_M GGUF. Small enough to run in the browser through WASM.
- Qwen3-4B: the official Q4_K_M GGUF, with native tool calling. Native targets only.
Use CandleModel::from_gguf(data)? for GGUF files. GGUF models currently have a 4096-token context. Candle’s local metrics (prompt and generated tokens, prefill time, time to first token, tokens per second) are typed reply extras, read with response.extras::<rig::candle::extension::CandleExt>().
rig::candle::pose adds YOLOv8 pose estimation (CandlePoseModel), which streams the poses found in each frame of a sequence of images.
The Rig repository has three runnable examples: candle_local streams a completion from a local GGUF model and prints Candle’s metrics, candle_pose estimates poses in images, and candle_wasm_chat runs SmolLM2 and a Rig agent inside a browser Web Worker, so chat messages never leave the page.
Local embeddings with fastembed
Section titled “Local embeddings with fastembed”rig::fastembed (feature fastembed) runs fastembed embedding models in your process. The fastembed feature downloads models from Hugging Face and ONNX Runtime binaries on first use; fastembed-hf-hub and fastembed-ort-download-binaries enable those parts separately. It is native-only.
[dependencies]rig = { version = "0.44.0", features = ["fastembed"] }use rig::fastembed::{Fastembed, FastembedModel};use rig::prelude::*;
let model_id = FastembedModel::AllMiniLML6V2Q;let embedder = Fastembed::load(&model_id)?.embedding(&model_id, None)?;
// Use it like any embedding model: EmbeddingsBuilder, vector stores, ...let index = InMemoryVectorStore::<String>::default().index(embedder);Pair it with Ollama, llama.cpp or Candle for a pipeline where neither documents nor prompts leave the machine.
See also
Section titled “See also”- Model Providers: every provider and the shared patterns
- Embeddings and Vector Stores
