LangGraph — Testing LangGraph Apps¶
A LangGraph app is ordinary Python around a non-deterministic model. Test the deterministic parts (nodes, routing, state, interrupts, graph shape) without any LLM, and keep a small, separate layer of tests and evaluations against real models.
Test Layers¶
| Layer | What is checked | Model | Speed | When |
|---|---|---|---|---|
| Tools and node functions | Pure logic, prompt building, parsing | None or fake | ms | Every commit |
| Routing functions | Every branch of every conditional edge | None | ms | Every commit |
| Graph flow | Path through nodes, state after each step, tool calls | Scripted fake | ms | Every commit |
| Persistence and HITL | Threads, interrupts, resume / reject, partial runs | Scripted fake + InMemorySaver |
ms | Every commit |
| Graph structure | Nodes and edges did not change by accident | None | ms | Every commit |
| Trajectory and quality evals | Right tools, right order, good answers | Real model | seconds, costs money | Nightly / before release |
Make the Graph Testable¶
The single most useful design decision: build the graph in a factory that receives the model (and checkpointer). Keep routing and review logic in module-level functions.
# app/support_graph.py
from typing import Literal
from langchain_core.language_models import BaseChatModel
from langchain_core.messages import AIMessage
from langchain_core.tools import tool
from langgraph.graph import END, START, MessagesState, StateGraph
from langgraph.prebuilt import ToolNode
from langgraph.types import Command, interrupt
SYSTEM_PROMPT = "You are a support agent. Use tools for order data."
RISKY_TOOLS = {"issue_refund"}
@tool
def get_order_status(order_id: str) -> str:
"""Return the delivery status of an order."""
return f"{order_id}: shipped"
@tool
def issue_refund(order_id: str, amount: float) -> str:
"""Refund an order."""
return f"refunded {amount} for {order_id}"
TOOLS = [get_order_status, issue_refund]
def route_after_agent(state: MessagesState) -> Literal["tools", "human_review", "__end__"]:
last = state["messages"][-1]
if not isinstance(last, AIMessage) or not last.tool_calls:
return END
if any(call["name"] in RISKY_TOOLS for call in last.tool_calls):
return "human_review"
return "tools"
def human_review(state: MessagesState) -> Command[Literal["tools", "__end__"]]:
calls = state["messages"][-1].tool_calls
decision = interrupt({"tool_calls": calls})
if decision == "approve":
return Command(goto="tools")
rejection = AIMessage("Refund was not approved by a human reviewer.")
return Command(goto=END, update={"messages": [rejection]})
def build_graph(model: BaseChatModel, checkpointer=None):
llm = model.bind_tools(TOOLS)
def agent(state: MessagesState) -> dict:
return {"messages": [llm.invoke([("system", SYSTEM_PROMPT), *state["messages"]])]}
builder = StateGraph(MessagesState)
builder.add_node("agent", agent)
builder.add_node("tools", ToolNode(TOOLS))
builder.add_node("human_review", human_review)
builder.add_edge(START, "agent")
builder.add_conditional_edges("agent", route_after_agent)
builder.add_edge("tools", "agent")
return builder.compile(checkpointer=checkpointer)
Layout used on this page:
app/__init__.py
app/support_graph.py # build_graph(model, checkpointer)
tests/__init__.py
tests/fakes.py # ScriptedModel, tool_call()
tests/conftest.py # checkpointer, config fixtures
tests/test_*.py
tests/snapshots/support_graph.mmd
pyproject.toml
Test dependencies: uv add --dev pytest pytest-asyncio agentevals.
Project test settings:
# pyproject.toml
[tool.pytest.ini_options]
pythonpath = ["."]
asyncio_mode = "strict" # pytest-asyncio; async tests use @pytest.mark.asyncio
markers = ["llm: tests that call a real LLM (need API keys)"]
Fake Chat Models¶
langchain_core.language_models.fake_chat_models ships several fakes:
| Fake | Returns | Notes |
|---|---|---|
GenericFakeChatModel(messages=iter([...])) |
Next item of the iterator (AIMessage or str) |
Supports tool_calls on AIMessage and token streaming; raises when the iterator is exhausted |
FakeListChatModel(responses=[...]) |
Next string, cycles back to the start | Text only |
FakeMessagesListChatModel(responses=[...]) |
Next BaseMessage, cycles |
Full messages without streaming by chunks |
ParrotFakeChatModel() |
Echoes the last input message | Checks what reaches the model |
None of them implements bind_tools(), which every tool-calling graph and create_agent call. A tiny subclass fixes that and also records what the model received:
# tests/fakes.py
from langchain_core.language_models.fake_chat_models import GenericFakeChatModel
from langchain_core.messages import AIMessage, BaseMessage
from pydantic import Field
class ScriptedModel(GenericFakeChatModel):
"""Fake chat model: returns scripted messages in order, records its inputs, accepts bind_tools()."""
received: list[list[BaseMessage]] = Field(default_factory=list)
def bind_tools(self, tools, **kwargs):
return self
def _generate(self, messages, stop=None, run_manager=None, **kwargs):
self.received.append(list(messages))
return super()._generate(messages, stop=stop, run_manager=run_manager, **kwargs)
def scripted(*messages: AIMessage | str) -> ScriptedModel:
return ScriptedModel(messages=iter(messages))
def tool_call(name: str, args: dict, call_id: str = "call_1") -> AIMessage:
return AIMessage(content="", tool_calls=[{"name": name, "args": args, "id": call_id}])
# tests/conftest.py
import uuid
import pytest
from langgraph.checkpoint.memory import InMemorySaver
@pytest.fixture
def checkpointer():
return InMemorySaver() # fresh, isolated memory per test
@pytest.fixture
def config():
return {"configurable": {"thread_id": f"test-{uuid.uuid4()}"}, "recursion_limit": 20}
Script exactly the calls you expect
An exhausted GenericFakeChatModel fails the run (RuntimeError: generator raised StopIteration). That is a feature: an unexpected extra model call — an extra loop iteration — breaks the test instead of passing silently.
Unit Tests: Nodes, Tools, Routing¶
Routing functions are plain functions — parametrize every branch:
# tests/test_routing.py
import pytest
from langchain_core.messages import AIMessage, HumanMessage
from app.support_graph import route_after_agent
from tests.fakes import tool_call
@pytest.mark.parametrize(
("last_message", "expected"),
[
(AIMessage("Your order has shipped."), "__end__"),
(tool_call("get_order_status", {"order_id": "ord-1"}), "tools"),
(tool_call("issue_refund", {"order_id": "ord-1", "amount": 10.0}), "human_review"),
(HumanMessage("hello"), "__end__"),
],
ids=["final-answer", "safe-tool", "risky-tool", "not-ai-message"],
)
def test_route_after_agent(last_message, expected):
assert route_after_agent({"messages": [last_message]}) == expected
A single node can run in isolation through graph.nodes["name"].invoke(state) — it returns the node's update, not the new state. Tools decorated with @tool are tested with .invoke({...}). ToolNode itself needs the graph runtime, so exercise it through a graph run.
# tests/test_nodes.py
from langchain_core.messages import AIMessage, HumanMessage, SystemMessage
from app.support_graph import SYSTEM_PROMPT, build_graph, get_order_status, issue_refund
from tests.fakes import scripted
def test_agent_node_sends_system_prompt_first():
model = scripted(AIMessage("Hello!"))
graph = build_graph(model)
update = graph.nodes["agent"].invoke({"messages": [HumanMessage("hi")]}) # one node, no graph run
assert update["messages"][-1].content == "Hello!"
[sent] = model.received
assert isinstance(sent[0], SystemMessage) and sent[0].content == SYSTEM_PROMPT
assert sent[-1].content == "hi"
def test_tools_are_plain_functions_too():
assert get_order_status.invoke({"order_id": "ord-1"}) == "ord-1: shipped"
assert issue_refund.invoke({"order_id": "ord-1", "amount": 5}) == "refunded 5.0 for ord-1"
Graph Tests with a Scripted Model¶
Assert the path (from stream_mode="updates"), the tool calls and results, and the final state — not the model's wording.
# tests/test_graph.py
import pytest
from langchain_core.messages import AIMessage, HumanMessage, ToolMessage
from langgraph.errors import GraphRecursionError
from langgraph.types import Command
from app.support_graph import build_graph
from tests.fakes import scripted, tool_call
def test_tool_loop_follows_expected_path(config):
model = scripted(
tool_call("get_order_status", {"order_id": "ord-42"}),
AIMessage("Order ord-42 has shipped."),
)
graph = build_graph(model)
path = []
for update in graph.stream({"messages": [HumanMessage("Where is ord-42?")]}, config,
stream_mode="updates"):
path.extend(update)
assert path == ["agent", "tools", "agent"]
def test_tool_message_contains_tool_output(config):
model = scripted(tool_call("get_order_status", {"order_id": "ord-42"}), AIMessage("Shipped."))
result = build_graph(model).invoke({"messages": [HumanMessage("Where is ord-42?")]}, config)
tool_messages = [m for m in result["messages"] if isinstance(m, ToolMessage)]
assert [m.content for m in tool_messages] == ["ord-42: shipped"]
assert result["messages"][-1].content == "Shipped."
def test_refund_pauses_for_human_review(checkpointer, config):
model = scripted(tool_call("issue_refund", {"order_id": "ord-7", "amount": 25.0}))
graph = build_graph(model, checkpointer=checkpointer)
result = graph.invoke({"messages": [HumanMessage("Refund ord-7")]}, config)
[pending] = result["__interrupt__"]
assert pending.value["tool_calls"][0]["name"] == "issue_refund"
assert graph.get_state(config).next == ("human_review",)
assert not any(isinstance(m, ToolMessage) for m in result["messages"]) # nothing executed yet
def test_approved_refund_is_executed(checkpointer, config):
model = scripted(
tool_call("issue_refund", {"order_id": "ord-7", "amount": 25.0}),
AIMessage("Refund done."),
)
graph = build_graph(model, checkpointer=checkpointer)
graph.invoke({"messages": [HumanMessage("Refund ord-7")]}, config)
result = graph.invoke(Command(resume="approve"), config)
assert any(m.content == "refunded 25.0 for ord-7" for m in result["messages"])
assert result["messages"][-1].content == "Refund done."
def test_rejected_refund_is_not_executed(checkpointer, config):
model = scripted(tool_call("issue_refund", {"order_id": "ord-7", "amount": 25.0}))
graph = build_graph(model, checkpointer=checkpointer)
graph.invoke({"messages": [HumanMessage("Refund ord-7")]}, config)
result = graph.invoke(Command(resume="reject"), config)
assert not any(isinstance(m, ToolMessage) for m in result["messages"])
assert "not approved" in result["messages"][-1].content
assert graph.get_state(config).next == ()
def test_thread_keeps_history_between_turns(checkpointer, config):
graph = build_graph(scripted(AIMessage("Hi!"), AIMessage("You asked about ord-42.")),
checkpointer=checkpointer)
graph.invoke({"messages": [HumanMessage("Hello, I have a question about ord-42")]}, config)
result = graph.invoke({"messages": [HumanMessage("What did I ask?")]}, config)
assert [m.type for m in result["messages"]] == ["human", "ai", "human", "ai"]
def test_endless_tool_loop_is_stopped(config):
always_tools = scripted(*[tool_call("get_order_status", {"order_id": "ord-1"}, f"c{i}") for i in range(50)])
graph = build_graph(always_tools)
with pytest.raises(GraphRecursionError):
graph.invoke({"messages": [HumanMessage("loop")]}, {**config, "recursion_limit": 6})
- A fresh
InMemorySaverand a uniquethread_idper test keep tests independent and parallel-safe (pytest-xdist). graph.get_state(config).nextis the cleanest assertion for "where did the graph stop?".- For interrupts assert both sides: the payload shown to the human, and the effect of each possible resume value.
Partial Execution¶
update_state(..., as_node=...) injects a node's output; interrupt_after=[...] (or interrupt_before) stops the run at a chosen point. Together they let you test one segment of a long graph without scripting everything before it.
# tests/test_partial.py
from langchain_core.messages import AIMessage, HumanMessage
from app.support_graph import build_graph
from tests.fakes import scripted, tool_call
def test_resume_from_tools_node_without_calling_the_model_first(checkpointer, config):
model = scripted(AIMessage("Order ord-9 has shipped.")) # only the final answer is scripted
graph = build_graph(model, checkpointer=checkpointer)
# Pretend "agent" already produced a tool call, then run only what follows
graph.update_state(
config,
{"messages": [HumanMessage("Where is ord-9?"), tool_call("get_order_status", {"order_id": "ord-9"})]},
as_node="agent",
)
assert graph.get_state(config).next == ("tools",)
result = graph.invoke(None, config, interrupt_after=["tools"]) # stop right after the tool
assert result["messages"][-1].content == "ord-9: shipped"
assert graph.get_state(config).next == ("agent",)
result = graph.invoke(None, config) # continue to the end
assert result["messages"][-1].content == "Order ord-9 has shipped."
Graph Structure Snapshot¶
Refactors can silently drop an edge. Pin the structure with a set of edges and a Mermaid snapshot reviewed in the PR diff:
# tests/test_structure.py
from pathlib import Path
from app.support_graph import build_graph
from tests.fakes import scripted
SNAPSHOT = Path(__file__).parent / "snapshots" / "support_graph.mmd"
def test_graph_edges():
graph = build_graph(scripted()).get_graph()
edges = {(e.source, e.target, e.conditional) for e in graph.edges}
assert set(graph.nodes) == {"__start__", "agent", "tools", "human_review", "__end__"}
assert ("tools", "agent", False) in edges
assert ("agent", "human_review", True) in edges
assert ("human_review", "tools", True) in edges
def test_graph_matches_mermaid_snapshot():
mermaid = build_graph(scripted()).get_graph().draw_mermaid()
if not SNAPSHOT.exists(): # first run: record, then review in the PR diff
SNAPSHOT.parent.mkdir(exist_ok=True)
SNAPSHOT.write_text(mermaid)
assert mermaid == SNAPSHOT.read_text()
draw_mermaid() is deterministic and needs no extra packages (draw_mermaid_png() calls a remote renderer by default — avoid it in CI).
Streaming and Async¶
With pytest-asyncio, test async graphs and the token stream the UI depends on:
# tests/test_async.py
import pytest
from langchain_core.messages import AIMessage, HumanMessage
from app.support_graph import build_graph
from tests.fakes import scripted
@pytest.mark.asyncio
async def test_streams_tokens_from_agent_node(config):
graph = build_graph(scripted(AIMessage("All good here")))
tokens = []
async for token, metadata in graph.astream({"messages": [HumanMessage("status?")]}, config,
stream_mode="messages"):
if metadata["langgraph_node"] == "agent":
tokens.append(token.content)
assert "".join(tokens) == "All good here"
Trajectory Evaluation¶
A trajectory is the sequence of messages — especially tool calls — the agent produced. Assert it directly, or use agentevals for configurable matching.
# tests/test_trajectory.py
from agentevals.trajectory.match import create_trajectory_match_evaluator
from langchain_core.messages import AIMessage, HumanMessage, ToolMessage
from app.support_graph import build_graph
from tests.fakes import scripted, tool_call
def tool_names(messages) -> list[str]:
return [call["name"] for m in messages if isinstance(m, AIMessage) for call in m.tool_calls]
def run_agent(question: str, *script):
return build_graph(scripted(*script)).invoke({"messages": [HumanMessage(question)]})
def test_tool_sequence():
result = run_agent(
"Where is ord-42?",
tool_call("get_order_status", {"order_id": "ord-42"}),
AIMessage("Shipped."),
)
assert tool_names(result["messages"]) == ["get_order_status"]
def test_trajectory_matches_reference():
result = run_agent(
"Where is ord-42?",
tool_call("get_order_status", {"order_id": "ord-42"}),
AIMessage("Shipped."),
)
reference = [
HumanMessage("Where is ord-42?"),
tool_call("get_order_status", {"order_id": "ord-42"}),
ToolMessage("ord-42: shipped", tool_call_id="call_1"),
AIMessage("Your order has shipped."), # final wording is not compared
]
evaluator = create_trajectory_match_evaluator(trajectory_match_mode="strict")
evaluation = evaluator(outputs=result["messages"], reference_outputs=reference)
assert evaluation["score"] is True
trajectory_match_mode |
Passes when the actual tool calls… |
|---|---|
strict |
Match the reference messages and tool calls in the same order |
unordered |
Contain the same tool calls in any order |
subset |
Are all contained in the reference (no extra tools used) |
superset |
Contain at least the reference tool calls (extras allowed) |
tool_args_match_mode (exact, ignore, subset, superset) controls how arguments are compared. For open-ended quality ("was this a sensible plan?") use an LLM-as-judge — agentevals.trajectory.llm.create_trajectory_llm_as_judge, DeepEval, or the evaluators in Phoenix, Langfuse and MLflow (06).
Tests Against a Real Model¶
Keep them few, marked, and focused on behaviour: which tool was chosen, with which arguments, whether the graph finished.
# tests/test_real_llm.py
import os
import pytest
from langchain_core.messages import AIMessage, HumanMessage
from app.support_graph import build_graph
pytestmark = [
pytest.mark.llm, # run with: pytest -m llm
pytest.mark.skipif(not os.getenv("ANTHROPIC_API_KEY"), reason="no API key"),
]
def test_real_model_calls_order_tool(config):
from langchain_anthropic import ChatAnthropic
graph = build_graph(ChatAnthropic(model="claude-sonnet-5", temperature=0))
result = graph.invoke({"messages": [HumanMessage("Where is order ord-42?")]}, config)
calls = [c for m in result["messages"] if isinstance(m, AIMessage) for c in m.tool_calls]
assert [c["name"] for c in calls] == ["get_order_status"] # behaviour, not wording
assert calls[0]["args"]["order_id"] == "ord-42"
- Run with
pytest -m llmnightly or before release, not on every commit;pytest -m "not llm"in the PR pipeline. - Use
temperature=0, an explicitrecursion_limitand a small model where possible; repeat flaky-prone cases 3–5 times and assert a pass rate. - Trace these runs with
tags/metadata(suite, case ID, git SHA) to debug failures from the trace, not from logs.
What to Test in Every LangGraph App¶
| Risk | Test |
|---|---|
| Endless loop | Scripted model that always calls a tool + small recursion_limit → expect GraphRecursionError or a graceful stop |
| Risky tool without approval | Scripted risky tool call → assert interrupt, assert nothing executed |
| Wrong branch | Parametrized routing tests for every Literal value |
| Lost memory | Two turns on one thread_id → history present; new thread_id → empty |
| Tool failure | Tool raises → the agent returns an error ToolMessage (e.g. ToolErrorMiddleware) or the run fails loudly — whichever is specified |
| Broken tool-call history | After trimming / handoffs, every ToolMessage still follows its AIMessage |
| Cross-user leakage | Two users in one Store → each sees only their namespace |
| Graph drift | Structure snapshot |
Testing Checklist¶
- Graph is built by a factory that receives the model and checkpointer
- Every routing function has a parametrized test per branch
- Fake model scripts exactly the expected calls; no network in unit tests
- Each test uses its own checkpointer and
thread_id - Interrupts are tested for payload, approve and reject paths
- Graph structure is pinned with an edges test or a Mermaid snapshot
- Real-model tests are marked, cheap, behaviour-based and traced
- Trajectory evals run on a golden set before each release
See also¶
- LangGraph — Stateful Agent Orchestration
- LangGraph — Observability & Deployment
- Pytest
- DeepEval — LLM Testing Guide
- Agentic AI — Testing, Evaluation & Observability
- LiteLLM — Observability & Testing