September 27, 2026
3 min read
LiteLLM — One API for 100+ LLM Providers
Python SDK and AI gateway that calls OpenAI, Anthropic, Gemini, Bedrock, Azure, Ollama and many more through one OpenAI-compatible interface — with retries, fallbacks, load balancing, caching and cost tracking.
Installation
uv add litellm # Python SDK
uv tool install 'litellm[proxy]' # AI gateway (proxy server) as a CLI tool
Section Map
File
Topics
01 Completion & Providers
completion, model prefixes, params, streaming, async, embeddings, exceptions
02 Tools & Structured Output
Function calling, tool loop, JSON schema, Pydantic responses, vision
03 Router & Reliability
Retries, fallbacks, Router, load balancing, caching, cost and budgets
04 Proxy (AI Gateway)
config.yaml, virtual keys, budgets, Docker, OpenAI SDK as the client
05 Observability & Testing
Callbacks, OpenTelemetry, custom loggers, mock_response, pytest patterns
SDK or Proxy?
Python SDK
Proxy (AI Gateway)
Runs
Inside your Python process
Separate service (:4000)
Clients
Python only
Any language, any OpenAI-compatible SDK
Keys
Provider keys in every app
Provider keys only in the gateway; apps get virtual keys
Budgets, rate limits per team
Manual
Built-in (needs Postgres)
Good for
Scripts, one service, test suites, evals
Platform teams, many services, central spend control
Minimal Example
from litellm import completion
# ANTHROPIC_API_KEY / OPENAI_API_KEY are read from the environment
messages = [{ "role" : "user" , "content" : "Give me three edge cases for a login form." }]
for model in [ "anthropic/claude-sonnet-5" , "openai/gpt-4o-mini" , "ollama/llama3.2" ]:
response = completion ( model = model , messages = messages , max_tokens = 300 )
print ( model , response . choices [ 0 ] . message . content )
The response has the OpenAI shape for every provider: choices[0].message, usage, model, id.
Quick Commands
Command
Use
litellm --model anthropic/claude-sonnet-5
Start a proxy for one model on :4000
litellm --config config.yaml
Start a proxy from a config
litellm --config config.yaml --detailed_debug
Proxy with verbose request logs
LITELLM_LOG=DEBUG python app.py
SDK debug logging
curl localhost:4000/health/liveliness
Proxy liveness probe
Quick Rules
Always use the provider prefix (anthropic/, openai/, bedrock/) — it removes guessing and routing surprises.
Set timeout and num_retries on every call — defaults are generous and LLM APIs do hang.
Configure fallbacks to another provider , not only another model of the same provider.
Enable drop_params when one code path serves many providers — unsupported params are dropped instead of raising.
Track cost per call (response._hidden_params["response_cost"]) and send it to your metrics.
Use mock_response in unit tests — no network, no keys, no spend.
Pin the exact version in the lockfile — the gateway holds all provider keys, which makes it a high-value supply-chain target.
Centralize keys in the proxy once more than one service calls LLMs.
See also
libraries
litellm
llm
python