Configuration
max_attempts includes the initial call. This example makes one retry after a
transient failure. Configuration is validated when the shared LLM backend is
created, before an experiment starts.
Model roles
Roles are resolved only by
LLMBackend. Coding-agent models—such as the Claude
Code model in coding_agent.model—remain explicit because their provider and
authentication mode are part of that agent’s configuration.
Existing explicit LLM model strings also remain valid:
utility route. Historical web-search inputs such as
gpt-4.1-mini and gpt-5-mini remain compatibility aliases and resolve to the
configured web_search model when used with a web-search method.
Kapso.research() uses the active mode’s web_search route and retry policy.
depth="light" requests medium search context; depth="deep" requests high
search context. Constructing Researcher(model="vendor/model") still selects an
explicit provider model, while Researcher() uses the default route.
An evolution run constructs one configured backend and shares it with search
strategies, RepoMemory, commit-message generation, and experiment insight
extraction. Knowledge-search reranking and navigation construct equivalent
backends from the same mode settings. This keeps route overrides consistent
throughout the run and cost accounting centralized for its shared callers.
Partial route overrides are supported; unspecified roles retain their defaults.
Unknown role keys and empty model values are configuration errors.
Retry semantics
Kapso retries only failures that are usually safe to repeat:- connection and timeout failures;
- HTTP
408,409,425, and429responses; - HTTP
500,502,503, and504responses; - provider exceptions classified as rate limits, timeouts, connection errors, service unavailability, or internal server errors.
n is capped exponential backoff:
jitter: true, Kapso uses full jitter—a random delay between zero and the
calculated cap—to avoid synchronized retry storms. With jitter: false, the
calculated delay is used exactly.
When transient failures exhaust all attempts, LLMRetryError reports the
operation, resolved model, and attempt count while preserving the provider
exception as its cause.
Parallel calls
Each model in a parallel request owns its retry loop. If one of five calls is temporarily rate-limited, the four successful calls are not submitted again. Successful call costs are recorded once, and returned response order continues to match input model order. Synchronous web search, parallel web search, standard synchronous completion, and standard parallel completion all share the same classifier andRetryPolicy.
The Researcher API retains its existing best-effort contract: after the
shared backend has exhausted a transient failure—or immediately rejects a
non-transient one—the individual research mode logs the failure and returns an
empty typed result.
Standalone use
ModelRouter, RetryPolicy, LLMRetryError, and
is_transient_llm_error are public from kapso.core for components that need
to inspect or compose the shared behavior.