LM struct is the provider client: a thin wrapper over OpenAI-compatible APIs with built-in retries, optional response caching, and history tracking. You rarely call it directly; a predictor uses an adapter to format a signature and send the result through the configured LM, which keeps business logic separate from transport. Configure one globally with configure(lm), or attach one per predictor with PredictBuilder::lm(...).
Responsibilities
LM handles three core responsibilities:
- Configuration - Stores provider credentials, model selection, and inference parameters (eg: temperature)
-
API Execution - Takes pre-formatted
Chatmessages and executes HTTP calls to the LLM provider - Response Caching - Optionally stores input/output pairs to avoid duplicate API calls
Structure
LM is built using the builder pattern. The builder collects an LMConfig, the serializable data half, and build() initializes the live client. The config holds:
model- Model identifier (e.g., “gpt-4o-mini” or “openai:gpt-4o-mini”)api_key- Provider API credentials (optional for local servers)base_url- API endpoint URL (optional, inferred from model provider)temperature- Sampling temperature (default: 0.7)max_tokens- Maximum completion tokens (default: 512)max_tool_iterations- Upper bound on tool-loop round trips (default: 10)max_retries- Additional attempts after a transient failure (default: 2)retry_base_delay_ms- Base delay for exponential retry backoff (default: 250)cache- Enable response caching (default: false)
LM adds:
client- Internal provider client (initialized during build)cache_handler- Optional response cache (initialized during build if enabled)
LM is cheap - clones share the same HTTP client and cache via Arc, making them ideal for concurrent use.
Construction and configuration
TheLM::builder() must be awaited with .build().await because client initialization is async.
Local server usage
For local OpenAI-compatible servers (vLLM, Ollama, etc.), providebase_url without an api_key:
Custom OpenAI-compatible endpoints
For custom endpoints requiring authentication, provide bothbase_url and api_key:
- Clone semantics:
LMimplementsClone; clones share the underlying client and cache viaArc, so they see the same history while carrying their own config copy.
API Reference
You can browse the fullLM module reference on docs.rs.
Global vs explicit usage
- Global:
configure(lm)sets the process-wide default LM used by predictors. - Per-instance override: Attach an LM to a specific predictor with
PredictBuilder::lm(...), which bypasses the global; or build a secondLMand callconfigure(lm)before the specific call.
Async execution and sync entry
- Async: LM building and calls are
async; prefer using an async runtime (Tokio). - Sync-style: If you need a plain
fn main, create a runtime andblock_onthe async work.
- Async (Tokio)
- Sync
Inspecting history
inspect_historyrequires caching to be enabled (.cache(true)); it panics on an LM built without caching. Entries areCacheEntryvalues served byResponseCache; see Utils. Only tool-free calls are cached, since tool loops execute side-effectful user code.
Configuration options
AllLM builder parameters have sensible defaults, so you only need to override what you need.
Example with custom settings
Provider Support
DSRs supports multiple LLM providers through Rig. Use theprovider:model format to specify which provider to use. Bare model names default to OpenAI.
Supported providers:
openai- OpenAI models (requiresOPENAI_API_KEY)anthropic- Anthropic models (requiresANTHROPIC_API_KEY)gemini- Google Gemini models (requiresGEMINI_API_KEY)groq- Groq models (requiresGROQ_API_KEY)openrouter- OpenRouter (requiresOPENROUTER_API_KEY)ollama- Local Ollama models (no API key required)
.api_key() if you want to override the default environment variable.
You can also use base_url to connect to any OpenAI-compatible server (vLLM, LiteLLM, etc.).
Usage examples
Tool sets and Code Mode
LM::call accepts tools directly, but repeated calls with a fixed set of tools should build a ToolSet once and reuse it via LM::call_with_toolset. A ToolSet pre-fetches every tool definition and indexes the executors by name.
With
ToolSet::code_mode, instead of emitting one JSON tool call per step, the model writes JavaScript against the tools as a JS API and composes their results in one execution. The returned set drops into any tool loop (LM::call_with_toolset, Predict) exactly like a normal ToolSet. It errors if two tool names mangle to the same JS identifier. See Code Mode for the full sandbox surface.
