Skip to main content
The LM struct is the provider client: a thin wrapper over OpenAI-compatible APIs with built-in retries, optional response caching, and history tracking. You rarely call it directly; a predictor uses an adapter to format a signature and send the result through the configured LM, which keeps business logic separate from transport. Configure one globally with configure(lm), or attach one per predictor with PredictBuilder::lm(...).

Responsibilities

LM handles three core responsibilities:
  1. Configuration - Stores provider credentials, model selection, and inference parameters (eg: temperature)
  2. API Execution - Takes pre-formatted Chat messages and executes HTTP calls to the LLM provider
  3. Response Caching - Optionally stores input/output pairs to avoid duplicate API calls

Structure

LM is built using the builder pattern. The builder collects an LMConfig, the serializable data half, and build() initializes the live client. The config holds:
  • model - Model identifier (e.g., “gpt-4o-mini” or “openai:gpt-4o-mini”)
  • api_key - Provider API credentials (optional for local servers)
  • base_url - API endpoint URL (optional, inferred from model provider)
  • temperature - Sampling temperature (default: 0.7)
  • max_tokens - Maximum completion tokens (default: 512)
  • max_tool_iterations - Upper bound on tool-loop round trips (default: 10)
  • max_retries - Additional attempts after a transient failure (default: 2)
  • retry_base_delay_ms - Base delay for exponential retry backoff (default: 250)
  • cache - Enable response caching (default: false)
The live LM adds:
  • client - Internal provider client (initialized during build)
  • cache_handler - Optional response cache (initialized during build if enabled)
Cloning an LM is cheap - clones share the same HTTP client and cache via Arc, making them ideal for concurrent use.

Construction and configuration

The LM::builder() must be awaited with .build().await because client initialization is async.

Local server usage

For local OpenAI-compatible servers (vLLM, Ollama, etc.), provide base_url without an api_key:

Custom OpenAI-compatible endpoints

For custom endpoints requiring authentication, provide both base_url and api_key:
  • Clone semantics: LM implements Clone; clones share the underlying client and cache via Arc, so they see the same history while carrying their own config copy.

API Reference

You can browse the full LM module reference on docs.rs.

Global vs explicit usage

  • Global: configure(lm) sets the process-wide default LM used by predictors.
  • Per-instance override: Attach an LM to a specific predictor with PredictBuilder::lm(...), which bypasses the global; or build a second LM and call configure(lm) before the specific call.

Async execution and sync entry

  • Async: LM building and calls are async; prefer using an async runtime (Tokio).
  • Sync-style: If you need a plain fn main, create a runtime and block_on the async work.

Inspecting history

inspect_history requires caching to be enabled (.cache(true)); it panics on an LM built without caching. Entries are CacheEntry values served by ResponseCache; see Utils. Only tool-free calls are cached, since tool loops execute side-effectful user code.

Configuration options

All LM builder parameters have sensible defaults, so you only need to override what you need.

Example with custom settings

Provider Support

DSRs supports multiple LLM providers through Rig. Use the provider:model format to specify which provider to use. Bare model names default to OpenAI. Supported providers:
  • openai - OpenAI models (requires OPENAI_API_KEY)
  • anthropic - Anthropic models (requires ANTHROPIC_API_KEY)
  • gemini - Google Gemini models (requires GEMINI_API_KEY)
  • groq - Groq models (requires GROQ_API_KEY)
  • openrouter - OpenRouter (requires OPENROUTER_API_KEY)
  • ollama - Local Ollama models (no API key required)
API keys are automatically read from environment variables. You only need to provide .api_key() if you want to override the default environment variable. You can also use base_url to connect to any OpenAI-compatible server (vLLM, LiteLLM, etc.).

Usage examples

All provider integrations are powered by Rig, which handles the provider-specific API details.

Tool sets and Code Mode

LM::call accepts tools directly, but repeated calls with a fixed set of tools should build a ToolSet once and reuse it via LM::call_with_toolset. A ToolSet pre-fetches every tool definition and indexes the executors by name. With ToolSet::code_mode, instead of emitting one JSON tool call per step, the model writes JavaScript against the tools as a JS API and composes their results in one execution. The returned set drops into any tool loop (LM::call_with_toolset, Predict) exactly like a normal ToolSet. It errors if two tool names mangle to the same JS identifier. See Code Mode for the full sandbox surface.

See also