Skip to content
  • 3 min read

LiteLLM — Tools & Structured Output

Function Calling

Tools are always described in the OpenAI format. LiteLLM converts them to Anthropic tool_use, Gemini function declarations, Bedrock Converse tools and so on.

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_order_status",
            "description": "Return the current status of an order by its ID.",
            "parameters": {
                "type": "object",
                "properties": {
                    "order_id": {"type": "string", "description": "Order ID, e.g. ord-42"},
                },
                "required": ["order_id"],
            },
        },
    }
]

response = completion(
    model="anthropic/claude-sonnet-5",
    messages=[{"role": "user", "content": "Where is my order ord-42?"}],
    tools=tools,
    tool_choice="auto",          # "auto" | "none" | "required" | {"type": "function", "function": {"name": ...}}
)

message = response.choices[0].message
for call in message.tool_calls or []:
    print(call.id, call.function.name, call.function.arguments)   # arguments is a JSON string

The Tool Loop

import json

from litellm import completion

HANDLERS = {"get_order_status": lambda order_id: {"order_id": order_id, "status": "shipped"}}


def run_agent(question: str, model: str = "anthropic/claude-sonnet-5", max_steps: int = 5) -> str:
    messages = [{"role": "user", "content": question}]
    for _ in range(max_steps):
        response = completion(model=model, messages=messages, tools=tools, timeout=60)
        message = response.choices[0].message
        messages.append(message.model_dump())           # keep the assistant turn with tool_calls

        if not message.tool_calls:
            return message.content

        for call in message.tool_calls:
            args = json.loads(call.function.arguments)
            result = HANDLERS[call.function.name](**args)
            messages.append({
                "role": "tool",
                "tool_call_id": call.id,
                "content": json.dumps(result),
            })
    raise RuntimeError("agent did not finish in max_steps")

Keep a step limit: models can loop on tool calls. Validate arguments before executing — they are model output, not trusted input.

Structured Output

Pydantic model as response_format

from pydantic import BaseModel, Field

from litellm import completion


class BugReport(BaseModel):
    title: str = Field(max_length=80)
    severity: str = Field(pattern="^(blocker|critical|major|minor)$")
    steps: list[str]
    expected: str
    actual: str


response = completion(
    model="openai/gpt-4o-mini",
    messages=[{"role": "user", "content": "Turn this into a bug report: checkout button does nothing on Safari"}],
    response_format=BugReport,
)
report = BugReport.model_validate_json(response.choices[0].message.content)

Raw JSON schema

response_format = {
    "type": "json_schema",
    "json_schema": {
        "name": "labels",
        "schema": {
            "type": "object",
            "properties": {"label": {"type": "string", "enum": ["bug", "feature", "question"]}},
            "required": ["label"],
            "additionalProperties": False,
        },
        "strict": True,
    },
}

Provider support

from litellm import supports_response_schema

supports_response_schema(model="anthropic/claude-sonnet-5")
Mode What you get
Native JSON schema (OpenAI, Gemini, recent Anthropic models) Output constrained to the schema
Fallback via tool calling LiteLLM wraps the schema in a forced tool call and returns its arguments
{"type": "json_object"} Valid JSON, but no schema guarantee

Always validate

Even with native schema support, validate with Pydantic. Refusals, truncation (finish_reason="length") and older models can still return invalid JSON. litellm.enable_json_schema_validation = True makes LiteLLM validate the response against the schema and raise on mismatch.

Vision (Images)

response = completion(
    model="anthropic/claude-sonnet-5",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "List UI defects visible on this screenshot."},
            {"type": "image_url", "image_url": {"url": "data:image/png;base64," + b64_png}},
        ],
    }],
)

Handy in UI test pipelines: send a failed Playwright screenshot and get a first-pass description of what went wrong.

Prompt Caching

Anthropic and Bedrock cache prompts marked with cache_control; OpenAI and Gemini cache long prefixes automatically. LiteLLM passes the marker through:

messages = [
    {
        "role": "system",
        "content": [{
            "type": "text",
            "text": LONG_SPEC,                                # e.g. 20k tokens of API docs
            "cache_control": {"type": "ephemeral"},
        }],
    },
    {"role": "user", "content": "Generate negative test cases for POST /orders"},
]
response = completion(model="anthropic/claude-sonnet-5", messages=messages)
print(response.usage.prompt_tokens_details)                   # cached_tokens on cache hits

Put stable content (system prompt, docs, few-shot examples) first and the changing question last — caching works on prefixes.

Reasoning Models

response = completion(
    model="anthropic/claude-sonnet-5",
    messages=messages,
    reasoning_effort="medium",          # mapped to the provider's thinking/reasoning settings
)
print(response.choices[0].message.reasoning_content)          # when the provider returns it

See also