Test and evaluate
In shortWrite a unit test with a mocked model, then run a small evaluation set against your weather agent with the 10xgraph eval command.
- 6 min read
- 8 sections
- Updated
- v0.10.0
- Markdown
Test your agent in two layers. A unit test swaps the real model for TestAgent, which returns scripted replies, so you can check wiring offline in milliseconds. An evaluation runs the real graph against a small dataset and scores the results. This step adds one of each.
What you build in this step
You write a pytest file that checks the graph without any model call, then an eval file with three cases that you run with 10xgraph eval. Unit tests answer “is my graph wired correctly?”. Evaluations answer “does my agent still behave well after I changed the prompt?”.
| Item | Unit test | Evaluation |
|---|---|---|
| Model | Mocked (TestAgent) |
Real, from your graph |
| Needs API key | No | Yes |
| Speed | Milliseconds | Seconds per case |
| Checks | Wiring, routing, tool calls | Response quality and tool choice |
| Run with | pytest |
10xgraph eval |
Prerequisites
- The weather agent from Build a graph by hand, saved as
agent_with_tools.pywith the compiled graph in a variable namedapp. - Test tooling:
pip install pytest pytest-asyncio - The CLI from
10xgraph-api(pip install 10xgraph-api) for the10xgraph evalcommand.
Unit test the graph with a mocked model
TestAgent is a drop-in replacement for Agent. It returns your responses list in order, cycling when the list runs out, and records every call so you can assert on it.
Write the first test
Create test_agent.py. The test builds a one-node graph around a TestAgent, runs it and checks the reply.
# test_agent.py
import pytest
from tenxgraph.core.graph import StateGraph
from tenxgraph.core.state import Message
from tenxgraph.qa.testing import TestAgent
from tenxgraph.utils import END
@pytest.mark.asyncio
async def test_agent_responds():
# The mocked model: no API call is made
test_agent = TestAgent(
model="test-model",
responses=["Hello! I can help you with the weather."],
)
# A minimal graph: MAIN -> END
graph = StateGraph()
graph.add_node("MAIN", test_agent)
graph.set_entry_point("MAIN")
graph.add_edge("MAIN", END)
compiled = graph.compile()
result = await compiled.ainvoke(
{"messages": [Message.text_message("Say hello")]},
config={},
)
# The last message is the assistant reply
last_message = result["messages"][-1]
assert "weather" in last_message.text()
test_agent.assert_called_times(1)Run it:
pytest test_agent.py -vYou should see the test pass:
test_agent.py::test_agent_responds PASSEDUse QuickTest for common patterns
QuickTest builds the graph for you. It has class methods for a single turn, a multi-turn conversation, a tool call and a custom agent, and each returns a TestResult with chainable assertions. Add these tests to test_agent.py:
# test_agent.py (additional tests)
from tenxgraph.qa.testing import QuickTest
@pytest.mark.asyncio
async def test_single_turn():
result = await QuickTest.single_turn(
user_message="What's the weather?",
agent_response="I'll check the weather for you.",
)
result.assert_contains("weather").assert_no_errors()
@pytest.mark.asyncio
async def test_multi_turn():
# Each tuple is (user message, scripted agent reply)
result = await QuickTest.multi_turn(
[
("Hi", "Hello there!"),
("How are you?", "I'm doing great!"),
]
)
result.assert_contains("great")Note that single_turn takes agent_response first and user_message second (default "Hello"). Use keyword arguments, as above, to avoid mixing them up.
Test that a tool is called
QuickTest.with_tools wires a TestAgent and a ToolNode together. For every name in tools it creates a mock function that records the call and returns the text from tool_responses (or "Mock result from <name>"). The scripted agent calls each tool on its first turn, then returns response.
# test_agent.py (additional test)
@pytest.mark.asyncio
async def test_tool_call():
result = await QuickTest.with_tools(
query="What's the weather in NYC?",
response="It's sunny in New York today.",
tools=["get_weather"],
tool_responses={"get_weather": "Sunny, 22C"},
)
result.assert_tool_called("get_weather")
result.assert_contains("sunny")The TestResult assertions are assert_contains, assert_not_contains, assert_equals, assert_tool_called, assert_tool_not_called, assert_message_count and assert_no_errors. Each returns the result so you can chain them.
Isolate a node with TestContext
TestContext is an optional helper that gives each test its own dependency container and in-memory store. Use ctx.create_graph() and ctx.create_test_agent() to keep tests from sharing state.
# test_agent.py (additional test)
from tenxgraph.qa.testing import TestContext
@pytest.mark.asyncio
async def test_with_context():
with TestContext() as ctx:
test_agent = ctx.create_test_agent(responses=["Test response from agent"])
graph = ctx.create_graph()
graph.add_node("MAIN", test_agent)
graph.set_entry_point("MAIN")
graph.add_edge("MAIN", END)
compiled = graph.compile()
await compiled.ainvoke(
{"messages": [Message.text_message("Test")]},
config={},
)
# TestAgent records how often it ran
test_agent.assert_called()
assert test_agent.call_count == 1Run the whole file with pytest test_agent.py -v. All tests should pass without any API key set. For MockToolRegistry, MockMCPClient and InMemoryStore, see Unit tests.
Evaluate the agent on a dataset
An evaluation runs your real graph on each case in an EvalSet and scores the result against what you expected. Because it uses the real model, results can vary between runs, so you set thresholds instead of exact matches. Use it to catch regressions when you change a prompt, a tool or a model.
Create the eval file
10xgraph eval discovers files named *_eval.py or eval_*.py under evals/ (or the directory set in 10xgraph.json). Each file must expose get_eval_set(), and may expose get_eval_config(). The CLI uses a module-level app as the graph; if there is none, it loads the graph from the agent entry in 10xgraph.json. Create evals/weather_eval.py:
# evals/weather_eval.py
from agent_with_tools import app # the compiled graph from the previous step
from tenxgraph.qa.evaluation import (
CriteriaConfig,
CriterionConfig,
EvalCase,
EvalConfig,
EvalSet,
ToolCall,
)
def get_eval_set() -> EvalSet:
"""Three cases: a weather question, a tool call and a greeting."""
return EvalSet(
eval_set_id="weather-agent-basic",
name="Basic weather agent",
description="Simple checks for the weather agent",
eval_cases=[
EvalCase.single_turn(
eval_id="weather-1",
user_query="What is the weather in London?",
expected_response="The weather in London is sunny and 22C.",
expected_tools=[ToolCall(name="get_weather")],
description="Calls the weather tool and reports the result",
),
EvalCase.single_turn(
eval_id="weather-2",
user_query="Is it nice in Paris today?",
expected_response="The weather in Paris is sunny and 22C.",
expected_tools=[ToolCall(name="get_weather")],
description="Calls the weather tool for a different city",
),
EvalCase.single_turn(
eval_id="weather-3",
user_query="Hello",
expected_response="Hello! How can I help you?",
description="Answers a greeting",
),
],
)
def get_eval_config() -> EvalConfig:
"""Score tool choice exactly and responses by word overlap (no judge model)."""
return EvalConfig(
criteria=CriteriaConfig(
tool_name_match=CriterionConfig.tool_name_match(threshold=1.0),
rouge_match=CriterionConfig.rouge_match(threshold=0.4),
)
)tool_name_match checks that the tools called match expected_tools. rouge_match scores token overlap with expected_response and needs no extra model call. For paraphrase-aware scoring, CriterionConfig.response_match uses an LLM judge instead (it calls a judge model, so it costs tokens). If you leave out get_eval_config(), the CLI uses built-in defaults (tool_name_match at 1.0, rouge_match at 0.5, node_order at 0.8) and a confeval.py next to your evals overrides them globally.
Run the evaluation
Point 10xgraph eval at the folder. It runs every case, prints a summary and writes HTML and JSON reports to eval_reports/ unless you pass --no-report.
10xgraph eval evals/The output varies with your scores, but looks like this (example):
Criteria weather_eval.py (source: per-file)
tool_name_match threshold=1.0
rouge_match threshold=0.4
...
HTML report: eval_reports/<report file>.htmlEach case passes only when every criterion meets its threshold. Open the HTML report to see the conversation, tool calls and per-criterion scores of each failed case.
| Flag | Short | Default | Purpose |
|---|---|---|---|
--output |
-o |
eval_reports |
Directory for reports |
--no-report |
off | Print the summary only | |
--threshold |
-t |
none | Exit with an error if the overall pass rate is below this value (0.0 to 1.0) |
--open |
off | Open the HTML report when done | |
--parallel |
-p |
off | Run cases concurrently |
--max-concurrency |
-c |
4 | Concurrent cases with --parallel |
Add --threshold 0.8 in CI so a drop in quality fails the build. To run evaluations from inside pytest instead, see Evals in pytest.
When to use each approach
Use both. They catch different failures, and the unit tests keep your feedback loop fast.
- Unit tests cover graph structure, routing, tool wiring and error handling. They are free and deterministic, so run them on every commit.
- Evaluations cover answer quality, tool choice and regressions after prompt or model changes. They cost API calls and vary slightly, so run them before a release or on a schedule.
Do not use evaluations to test wiring bugs: a failing case does not tell you whether the cause is the graph or the model. Do not use TestAgent to judge answer quality: it only returns what you scripted.
What you learned
TestAgentreplaces the model with scripted replies, andQuickTestbuilds common test graphs in one call.TestResultassertions are chainable, andTestContextisolates each test.- An
EvalSetofEvalCaseobjects plus anEvalConfigdescribes an evaluation, and10xgraph evalruns it against your real graph and writes reports.
Next steps
You have finished the tutorial track. Continue with:
- Testing for unit test patterns, criteria, presets and user simulation.
- Run the server to serve the agent over HTTP with
10xgraph apiand10xgraph play. - Production checklist before you deploy.
- Prebuilt agents and Concepts for the next level of detail.
Frequently asked questions
- Do unit tests with TestAgent call a real model?
- No. TestAgent returns the responses you give it, so the test runs offline, instantly and with the same result every time.
- Does 10xgraph eval call a real model?
- Yes. It runs your real graph against each case, so it needs your provider API key. Cases that use only rouge_match and tool_name_match need no extra judge model.
- Where should eval files live?
- By default 10xgraph eval looks in an evals/ folder (or the directory set in 10xgraph.json) for files named *_eval.py or eval_*.py. You can also pass a file or folder path.