Skip to content

Local Models

Rig can run entirely on your own hardware. There are two kinds of local setup:

  • A local server that Rig talks to over HTTP: Ollama or llama.cpp. Both are built into rig and need no extra feature.
  • In-process inference, with no server at all: Candle for completions and fastembed for embeddings. Both are companion crates behind a feature on rig.

Whichever you pick, the model plugs into AgentBuilder, EmbeddingsBuilder and the vector stores like any hosted model.

Ollama::new() talks to the daemon at http://localhost:11434 with no authentication. Ollama::from_env() reads the optional OLLAMA_API_BASE_URL and OLLAMA_API_KEY for a remote or secured daemon.

Ollama serves two chat APIs, and Rig has a model for each:

  • client.completion(id) uses Ollama’s OpenAI-compatible /v1/chat/completions.
  • client.native_completion(id) uses Ollama’s native /api/chat. Only this route sends model parameters such as num_ctx, and it maps the reasoning option to Ollama’s think.
use rig::prelude::*;
use rig::providers::ollama::{self, Ollama};
let client = Ollama::new();
let agent = AgentBuilder::new(client.completion(ollama::QWEN3))
.preamble("You are a concise assistant.")
.build();
println!("{}", agent.prompt("What is Rust's borrow checker?").await?.output());

Ollama-only fields are typed in ollama::extension::OllamaOptions. keep_alive goes to both routes; the model parameters (num_ctx, top_k, min_p, repeat_penalty, num_gpu, num_thread, …) only to the native route, and the /v1 route refuses them rather than dropping them:

use rig::prelude::*;
use rig::providers::ollama::extension::{KeepAlive, OllamaOptions};
use rig::providers::ollama::{self, Ollama};
let options = OllamaOptions::default()
.num_ctx(16_384)
.keep_alive(KeepAlive::duration("10m"));
let agent = AgentBuilder::new(Ollama::new().native_completion(ollama::QWEN3))
.provider_option(options)
.build();

Ollama also serves embeddings through /api/embed. Pass the vector width, or None to use the width Rig knows for that model:

use rig::providers::ollama::{self, Ollama};
let embedder = Ollama::new().embedding(ollama::NOMIC_EMBED_TEXT, None);

client.list_models().await? lists the models the daemon has pulled. Runnable examples: rag_ollama and vector_search_ollama.

The llamacpp provider talks to llama-server’s OpenAI-compatible API, by default at http://localhost:8080/v1. A server started without --api-key needs no key, so pass an empty one:

use rig::prelude::*;
use rig::providers::llamacpp;
let client = llamacpp::new("");
let agent = AgentBuilder::new(client.chat(llamacpp::LLAMA_CPP))
.preamble("You are a concise assistant.")
.build();

llamacpp::from_env() reads LLAMACPP_API_KEY (required) and LLAMACPP_API_BASE_URL (optional) for a secured or remote server. For an unsecured server elsewhere, set the base URL on a configuration:

use rig::providers::openai::{OpenAIConfig, wire::LLAMACPP};
let client = OpenAIConfig::with_key(&LLAMACPP, "")
.with_base_url("http://gpu-box:8080/v1")
.client();

llamacpp::LLAMA_CPP is the conventional id for a server that hosts one model; for a multi-model router, use an id from client.list_models().await?. The same client serves embedding(id, None) and rerank(id) when the server enables them. llama-server treats a specific-function tool choice as auto, so Rig refuses that tool choice instead of sending it.

rig::candle (feature candle) runs models in your process on the CPU with Candle. You load the model files yourself and hand Rig the bytes; the crate does no filesystem or network access.

[dependencies]
rig = { version = "0.44.0", features = ["candle"] }
use rig::candle::{CandleModel, ModelData};
use rig::prelude::*;
let model = CandleModel::from_safetensors_async(ModelData {
config: std::fs::read("./model/config.json")?,
tokenizer: std::fs::read("./model/tokenizer.json")?,
weights: std::fs::read("./model/model.safetensors")?,
})
.await?;
let agent = AgentBuilder::new(model.completion())
.preamble("You are a concise assistant.")
.build();
println!("{}", agent.prompt("Explain ownership briefly.").await?.output());

Only validated checkpoints load; anything else is rejected rather than guessed at:

  • Llama 3 instruct: one unsharded safetensors checkpoint.
  • SmolLM2-360M-Instruct: Q4_K_M GGUF. Small enough to run in the browser through WASM.
  • Qwen3-4B: the official Q4_K_M GGUF, with native tool calling. Native targets only.

Use CandleModel::from_gguf(data)? for GGUF files. GGUF models currently have a 4096-token context. Candle’s local metrics (prompt and generated tokens, prefill time, time to first token, tokens per second) are typed reply extras, read with response.extras::<rig::candle::extension::CandleExt>().

rig::candle::pose adds YOLOv8 pose estimation (CandlePoseModel), which streams the poses found in each frame of a sequence of images.

The Rig repository has three runnable examples: candle_local streams a completion from a local GGUF model and prints Candle’s metrics, candle_pose estimates poses in images, and candle_wasm_chat runs SmolLM2 and a Rig agent inside a browser Web Worker, so chat messages never leave the page.

rig::fastembed (feature fastembed) runs fastembed embedding models in your process. The fastembed feature downloads models from Hugging Face and ONNX Runtime binaries on first use; fastembed-hf-hub and fastembed-ort-download-binaries enable those parts separately. It is native-only.

[dependencies]
rig = { version = "0.44.0", features = ["fastembed"] }
use rig::fastembed::{Fastembed, FastembedModel};
use rig::prelude::*;
let model_id = FastembedModel::AllMiniLML6V2Q;
let embedder = Fastembed::load(&model_id)?.embedding(&model_id, None)?;
// Use it like any embedding model: EmbeddingsBuilder, vector stores, ...
let index = InMemoryVectorStore::<String>::default().index(embedder);

Pair it with Ollama, llama.cpp or Candle for a pipeline where neither documents nor prompts leave the machine.