Running Evaluations in pytest
In shortIntegrate agent evaluations into pytest to measure quality alongside unit tests. Use decorators and assertion helpers to verify agents meet quality thresholds.
- 6 min read
- 11 sections
- Updated
- v0.10.0
- Markdown
Embed agent evaluations into your pytest test suite to measure quality alongside your unit tests. Unlike unit tests, evaluations run real LLM calls and score your agent’s responses and tool usage against expected outcomes. This guide covers how to integrate the evaluation framework into pytest using decorators, standalone functions, and custom assertions.
Prerequisites
You have already installed 10xGraph and pytest:
pip install pytest pytest-asyncio "10xgraph[google-genai]"tenxgraph.qa.evaluation is built into 10xgraph, no extra package needed.
You have eval cases defined in an evalset file (.evalset.json) or as Python code. See Building eval sets for how to structure cases and Criteria reference for scoring rules. You also have a compiled graph ready for evaluation (see Quick start for setup examples).
Quick start
The simplest way is to use the @eval_test decorator on an async test function that returns a (graph, collector) tuple:
import pytest
from tenxgraph.qa.evaluation.testing import eval_test, create_eval_app
# Fixture that builds your graph once
@pytest.fixture(scope="session")
def eval_graph():
from my_agents import build_weather_agent
graph = build_weather_agent()
app, collector = create_eval_app(graph)
return app, collector
# Evaluation test
@eval_test("tests/fixtures/weather_agent.evalset.json")
async def test_weather_agent_quality(eval_graph):
return eval_graphRun it with pytest:
pytest test_evals.py::test_weather_agent_quality -vWith no config, @eval_test uses EvalConfig.default(): an exact tool trajectory match (threshold 1.0) plus an LLM response match (threshold 0.8), so a judge model API key must be available. Pass config=EvalPresets.tool_usage() or EvalPresets.quick_check() to change that. The decorator’s default threshold is 1.0, so every case must pass. It fails with details of which cases failed and why.
Set up evaluation fixtures
Create a shared fixture that compiles your graph with evaluation instrumentation. This fixture wires in a TrajectoryCollector to capture the agent’s execution trace, which the evaluator then scores.
# conftest.py
import pytest
from tenxgraph.qa.evaluation.testing import create_eval_app
@pytest.fixture(scope="session")
def weather_agent_app():
"""Compile the weather agent for evaluation.
The create_eval_app helper does the plumbing: it creates a
TrajectoryCollector, wires in the callback manager, and compiles
the graph. You only provide the uncompiled StateGraph.
"""
from my_agents import build_weather_agent
graph = build_weather_agent() # Uncompiled StateGraph
app, collector = create_eval_app(
graph,
capture_all_events=True, # Capture all runtime events for detailed scoring
)
return app, collectorThe fixture returns (compiled_app, collector). The app is the compiled graph; the collector records every node execution and tool call. Pass both to your evaluation tests.
Run evals with the @eval_test decorator
The @eval_test decorator wraps a test function, runs an entire eval set, and asserts that the overall pass rate meets a threshold. Point it at an evalset file:
from tenxgraph.qa.evaluation.testing import eval_test
@eval_test(
eval_file="tests/fixtures/weather_agent.evalset.json",
threshold=0.9, # Require 90% of cases to pass
)
async def test_weather_agent_regression(weather_agent_app):
"""Evaluate the weather agent against a regression suite."""
return weather_agent_appThe test function receives the fixture and returns (graph, collector). The decorator:
- Runs every case in the eval file against your graph.
- Scores each case using the criteria from
config(orEvalConfig.default()). - Asserts that the pass rate is >=
threshold. - On failure, prints which cases failed and why.
If the test fails:
AssertionError: Evaluation failed: 80.0% pass rate (threshold: 90.0%)
Failed cases:
- weather_london: tool_trajectory_avg_score
- booking_flight: response_match_scoreAuto-detect eval files
If you omit eval_file, the decorator searches for a file named after your test:
@eval_test(threshold=0.95)
async def test_booking_agent(booking_agent_app):
# Searches for: tests/fixtures/booking_agent.evalset.json
# tests/eval/booking_agent.evalset.json
# eval/booking_agent.evalset.json
return booking_agent_appRun evals with the standalone run_eval function
For tests that need more control, use run_eval to run an evaluation and capture the full report:
from tenxgraph.qa.evaluation.testing import run_eval
from tenxgraph.qa.evaluation.config.presets import EvalPresets
async def test_agent_with_custom_checks(weather_agent_app):
"""Run evals and make custom assertions on the report."""
app, collector = weather_agent_app
# Run evaluation with a specific config
config = EvalPresets.tool_usage(threshold=0.8)
report = await run_eval(
graph=app,
collector=collector,
eval_set_path="tests/fixtures/weather_agent.evalset.json",
config=config,
verbose=True,
)
# Make custom assertions
assert report.summary.pass_rate >= 0.8
assert len(report.failed_cases) <= 2
print(f"Passed {len(report.passed_cases)} cases")This approach gives you the full EvalReport object, so you can inspect and assert on:
report.summary.pass_rate: overall pass rate (0.0 to 1.0)report.summary.total_cases: how many cases ranreport.passed_cases: list of cases that passedreport.failed_cases: list of cases that failed with error detailsreport.summary.criterion_stats: per-criterion score breakdowns, keyed by criterion name
Assert on specific criteria
Use assert_criterion_passed to verify that a single criterion met a minimum average score. The criterion argument is the criterion’s reported name (for example tool_name_match_score, rouge_match, tool_trajectory_avg_score, response_match_score), and it must be enabled in the config you ran with, otherwise the helper raises “not found in report”:
from tenxgraph.qa.evaluation.config.presets import EvalPresets
from tenxgraph.qa.evaluation.testing import (
run_eval,
assert_criterion_passed,
assert_eval_passed,
)
async def test_tool_selection_quality(weather_agent_app):
"""Ensure tool selection is accurate."""
app, collector = weather_agent_app
report = await run_eval(
graph=app,
collector=collector,
eval_set_path="tests/fixtures/weather_agent.evalset.json",
config=EvalPresets.comprehensive(use_llm_judge=False),
)
# At least 95% of cases must have passed overall
assert_eval_passed(report, min_pass_rate=0.95)
# Tool name accuracy must be high
assert_criterion_passed(
report,
criterion="tool_name_match_score",
min_score=0.98, # Average score across all cases
)
# Response overlap must be good
assert_criterion_passed(
report,
criterion="rouge_match",
min_score=0.85,
)This pattern is useful when you care about specific aspects of agent quality and want to fail the test if a particular criterion degrades.
Parametrize tests with individual eval cases
Use parametrize_eval_cases to run each eval case as a separate pytest test. This produces granular pass/fail reporting. Test ids are the case eval_id values. To run them in parallel you need a plugin such as pytest-xdist:
from tenxgraph.qa.evaluation.testing import parametrize_eval_cases
@parametrize_eval_cases("tests/fixtures/weather_agent.evalset.json")
async def test_weather_agent_case(weather_agent_app, eval_case):
"""Test a single eval case."""
from tenxgraph.qa.evaluation import AgentEvaluator
from tenxgraph.qa.evaluation.config.presets import EvalPresets
app, collector = weather_agent_app
config = EvalPresets.tool_usage()
evaluator = AgentEvaluator(app, collector, config=config)
result = await evaluator.evaluate_case(eval_case)
assert result.passed, f"Case failed: {', '.join(c.criterion for c in result.failed_criteria)}"Running this test:
pytest test_evals.py::test_weather_agent_case -vproduces:
test_evals.py::test_weather_agent_case[weather_london] PASSED
test_evals.py::test_weather_agent_case[weather_tokyo] PASSED
test_evals.py::test_weather_agent_case[weather_paris] FAILEDEach case is a separate test item, so you can run a subset:
pytest test_evals.py::test_weather_agent_case[weather_london] -vBuild eval sets programmatically
Instead of loading from a .evalset.json file, create eval sets in code using create_simple_eval_set:
from tenxgraph.qa.evaluation.config.presets import EvalPresets
from tenxgraph.qa.evaluation.testing import create_simple_eval_set, run_eval
async def test_agent_with_inline_cases(weather_agent_app):
"""Run evals with cases defined inline."""
app, collector = weather_agent_app
# Create a simple eval set
eval_set = create_simple_eval_set(
eval_set_id="weather-quick-check",
cases=[
("What is the weather in London?", "London", "weather_london"),
("Weather in Tokyo?", "Tokyo", "weather_tokyo"),
("Tell me about Paris weather.", "Paris", "weather_paris"),
],
)
# run_eval passes the EvalSet through to AgentEvaluator.evaluate, which
# accepts either an EvalSet object or a path to a JSON file.
report = await run_eval(
graph=app,
collector=collector,
eval_set_path=eval_set,
config=EvalPresets.quick_check(),
)
assert report.summary.pass_rate == 1.0Each tuple is (user_query, expected_response, case_name); case ids are generated as case_0, case_1, and so on. quick_check() scores response text with ROUGE overlap only (no LLM). Without a config, EvalConfig.default() applies (exact tool trajectory plus LLM response match), and these cases have no expected tools. For more control, build an EvalSet directly using the eval-sets API (see Building eval sets).
Common errors and fixes
“Eval file not found: auto-detected”
Cause: You used @eval_test without eval_file, and the decorator could not find a matching file.
Fix: Either provide the full path:
@eval_test("tests/fixtures/my_agent.evalset.json")
async def test_my_agent(app_fixture):
...Or place the eval file in one of the auto-detected locations:
tests/fixtures/{test_name}.evalset.json
tests/eval/{test_name}.evalset.json
eval/{test_name}.evalset.json“eval_test decorated function must return (graph, collector) tuple”
Cause: Your test function returned something other than a tuple of two items.
Fix: Ensure you return exactly (compiled_graph, collector):
@eval_test("tests/fixtures/my_eval.evalset.json")
async def test_my_agent(my_fixture):
# my_fixture should be (graph, collector)
return my_fixture # Correct
# Wrong:
async def test_my_agent(my_fixture):
return my_fixture.graph # Missing collector“Evaluation failed: … pass rate (threshold: …)”
Cause: One or more eval cases failed to meet the criteria thresholds.
Fix: Inspect the failure details in the error message. It lists which cases failed and which criteria they missed. Then either:
- Improve the agent to pass those cases.
- Adjust the eval set expectations if they are wrong.
- Lower the
thresholdparameter temporarily for debugging.
Use verbose=True in run_eval to see detailed scoring for each case:
report = await run_eval(
...,
verbose=True,
)Next steps
- Building eval sets: structure eval cases, multi-turn scenarios, and custom expectations
- Criteria reference: scoring rules and thresholds for each criterion
- How to run evaluations: CLI commands, parallel runs, and CI integration
- Reports: HTML dashboards and JSON export for evals
Frequently asked questions
- How are evaluation tests different from unit tests?
- Unit tests mock the LLM and tools, running fast and deterministically. Evaluation tests run your agent against real LLM calls and measure quality (correct tool usage, accurate responses, safety). They run at a different speed and cost.
- Can I run evals without pytest?
- Yes. Use AgentEvaluator, QuickEval, or the CLI directly. pytest integration is optional, for teams that want evals to run as part of their test suite.
- Do evaluations run in parallel in pytest?
- No. pytest runs each test sequentially. For parallel evaluation runs, use the --parallel flag from the CLI instead.