Evaluation example

In shortWalk through the weather-agent evaluation example: eval cases, QuickEval, trajectory criteria, user simulation and CI reports.

  • 6 min read
  • 11 sections
  • Updated
  • v0.10.0
  • Markdown

This example evaluates a small weather agent end to end: it defines eval cases, runs quick checks, scores tool trajectories, simulates multi-turn users and writes JSON, HTML and JUnit reports. It shows how to measure whether a real agent behaves correctly, which unit tests with a mocked model cannot do.

What the example contains

The example lives in agentflow/examples/evaluation. One shared weather agent graph backs five test folders, each focused on a different part of the evaluation API.

Path What it shows
test_graph/__init__.py The weather agent graph and create_app_and_collector()
samples.py Shared EvalCase and EvalSet definitions, including trajectory cases
test1/ Criteria on single cases (the module is entirely commented out, kept as a reference)
test2/ Every criterion type in AgentEvaluator
test3/ UserSimulator, BatchSimulator and SimulationGoalsCriterion
test4/ QuickEval, EvalSetBuilder, presets and reporters
test5/ evaluate_case, evaluate, rubrics, presets and EvaluationRunner

Run the example

The tests call a live Gemini model, so install the Google extra and set an API key first. Each test module calls pytest.skip at import time so a root pytest run stays offline; delete that call in the module you want to run.

Terminal
# From the agentflow/ directory of the repo
pip install "10xgraph[google-genai]"
export GEMINI_API_KEY="your-key"   # GOOGLE_API_KEY also works

# After removing the module-level pytest.skip in test4
pytest examples/evaluation/test4/ -v -s

The tests use relative imports (from ..test_graph import ...), so run pytest from the agentflow/ directory rather than from inside the example folder.

The weather agent under test

The shared graph has one model node and one tool node, wired with the usual tool loop. Two plain Python functions are the tools, and create_app_and_collector() compiles the graph with a trajectory collector so evaluators can see which tools ran.

Python
# examples/evaluation/test_graph/__init__.py (condensed)
from tenxgraph.core.graph import Agent, StateGraph, ToolNode
from tenxgraph.core.state import AgentState
from tenxgraph.qa.evaluation.testing import create_eval_app
from tenxgraph.utils.constants import END


def get_weather(location: str) -> str:
    """Get the current weather for a location."""
    return f"The weather in {location} is sunny"


def get_forecast(location: str, days: int = 3) -> str:
    """Get a multi-day weather forecast for a location."""
    return f"{days}-day forecast for {location}: sunny, cloudy, sunny"


agent = Agent(
    model="gemini-2.5-flash",
    provider="google",
    system_prompt=[{"role": "system", "content": "You are a helpful weather assistant."}],
    tool_node="TOOL",
)


def should_use_tools(state: AgentState) -> str:
    last = state.context[-1] if state.context else None
    if last is None:
        return "TOOL"
    if getattr(last, "tools_calls", None) and last.role == "assistant":
        return "TOOL"
    if last.role == "tool":
        return "MAIN"
    return END


graph = StateGraph()
graph.add_node("MAIN", agent)
graph.add_node("TOOL", ToolNode([get_weather, get_forecast]))
graph.add_conditional_edges("MAIN", should_use_tools, {"TOOL": "TOOL", END: END})
graph.add_edge("TOOL", "MAIN")
graph.set_entry_point("MAIN")

# Returns (compiled_graph, collector)
compiled_graph, collector = create_eval_app(graph)

The example file uses a fuller system prompt and builds the graph inside a function; the structure is the same. Each test folder exposes compiled_graph and collector as session-scoped pytest fixtures in its conftest.py.

Define eval cases and an eval set

Eval cases describe what the agent should do for one input. samples.py is the single source of truth, so every test folder reuses the same cases.

Python
# examples/evaluation/samples.py (excerpt)
from tenxgraph.qa.evaluation.dataset import EvalCase, ToolCall
from tenxgraph.qa.evaluation.dataset.eval_set import EvalSet

NYC = EvalCase.single_turn(
    eval_id="nyc_happy",
    name="NYC happy path",
    user_query="Please call the get_weather function for New York City",
    expected_response="The weather in New York City is sunny",
    expected_tools=[ToolCall(name="get_weather")],
)

CAPITAL_QUESTION = EvalCase.single_turn(
    eval_id="capital_no_tool",
    name="General knowledge, no tool",
    user_query="What is the capital of France?",
    expected_response="Paris",
    expected_tools=[],  # empty list asserts zero tool calls
)

EVAL_SET = EvalSet(
    eval_set_id="weather_agent_eval",
    name="Weather Agent Evaluation Suite",
    eval_cases=[NYC, CAPITAL_QUESTION],
)

samples.py also defines trajectory cases such as WEATHER_THEN_FORECAST, whose expected_tools list get_weather then get_forecast. See Eval sets for the full dataset format.

Get fast feedback with QuickEval

QuickEval scores a query without building an eval set or config. Test4 uses check() for one query, batch() for query and response pairs, tool_usage() for tool expectations and conversation_flow() for a short multi-turn run.

Python
# examples/evaluation/test4/test_single_turn.py (excerpt)
from tenxgraph.qa.evaluation import QuickEval, assert_eval_passed

report = await QuickEval.check(
    graph=compiled_graph,
    collector=collector,
    query="What is the weather in London?",
    expected_response_contains="sunny",
    expected_tools=["get_weather"],
    threshold=0.5,
    print_results=True,
)
assert_eval_passed(report)  # raises if the pass rate is below 1.0

# Several query and expected-response pairs at once
report = await QuickEval.batch(
    graph=compiled_graph,
    collector=collector,
    test_pairs=[
        ("Weather in NYC?", "The weather in NYC is sunny"),
        ("Capital of Spain?", "The capital of Spain is Madrid"),
    ],
    threshold=0.3,
)

Test4 also covers EvalSetBuilder with QuickEval.from_builder() and QuickEval.preset() with EvalPresets.quick_check().

Score tool usage and trajectories

Trajectory criteria compare the tools the agent actually called against expected_tools. Test2 combines the criteria that need no judge model in one config, so a single agent run is scored several ways.

Python
# examples/evaluation/test2/test_weather_agent.py (excerpt)
from tenxgraph.qa.evaluation import AgentEvaluator, CriterionConfig, EvalConfig, MatchType

config = EvalConfig(
    criteria={
        "tool_name_match_score": CriterionConfig.tool_name_match(threshold=1.0),
        "tool_trajectory_avg_score": CriterionConfig.trajectory(
            threshold=1.0,
            match_type=MatchType.EXACT,
        ),
        "rouge_match": CriterionConfig.rouge_match(threshold=0.4),
        "contains_keywords": CriterionConfig.contains_keywords(
            keywords=["New York", "sunny"], threshold=0.5
        ),
    },
    reporter={"enabled": True},
)
evaluator = AgentEvaluator(compiled_graph, collector, config=config)
result = await evaluator.evaluate_case(WEATHER_NYC)
assert result.passed, [c.criterion for c in result.failed_criteria]

MatchType controls how strict the comparison is.

Match type Passes when
EXACT The same tools are called with the same arguments and order
IN_ORDER Expected tools appear in order, extra calls are allowed
ANY_ORDER Expected tools all appear in any order, extra calls are allowed

The result object carries the pass state, a score per criterion, the actual response, the actual tool calls and the duration. Test5 asserts on each of these fields. See Criteria for every criterion.

Simulate a multi-turn user

A simulated user plays a scenario against your agent and a judge scores whether the goals were met. Use it when success depends on the whole conversation rather than one reply. Test3 runs a scenario, attaches SimulationGoalsCriterion, and runs several scenarios with BatchSimulator.

Python
# examples/evaluation/test3/test_weather_simulator.py (excerpt)
from tenxgraph.qa.evaluation import (
    BatchSimulator,
    ConversationScenario,
    CriterionConfig,
    SimulationGoalsCriterion,
    UserSimulator,
)

scenario = ConversationScenario(
    scenario_id="weather_single_city",
    description="User wants to know the current weather in Tokyo for trip planning",
    starting_prompt="I'm thinking of visiting Tokyo soon. Can you help me?",
    conversation_plan=(
        "1. User hints at travel interest\n"
        "2. User explicitly asks for Tokyo weather\n"
        "3. User confirms they got the information they needed"
    ),
    goals=["Get weather information for Tokyo"],
    max_turns=4,
)

judge = SimulationGoalsCriterion(config=CriterionConfig(enabled=True, threshold=0.5))
simulator = UserSimulator(model="gemini/gemini-2.5-flash", criteria=[judge])

result = await simulator.run(compiled_graph, scenario)
print(result.turns, result.criterion_scores["simulation_goals"])

# Run many scenarios, then summarize
batch = BatchSimulator(simulator=simulator)
results = await batch.run_batch(compiled_graph, [scenario])
print(batch.summary(results))

UserSimulator makes its own model calls to play the user, so a simulation costs more than a single-turn check. See User simulation for the options.

Write reports for CI

Reporters turn an evaluation run into files a CI system can store. Test5 and test4 configure ReporterConfig and then check which files appear.

Python
# examples/evaluation/test5/test_multi_turn.py (excerpt)
from tenxgraph.qa.evaluation import AgentEvaluator, CriterionConfig, EvalConfig, ReporterConfig

config = EvalConfig(
    criteria={"response_match_score": CriterionConfig.response_match(threshold=0.3)},
    reporter=ReporterConfig(
        enabled=True,
        output_dir="./eval_reports",
        console=False,
        json_report=True,
        html=True,
        junit_xml=True,        # off by default
        timestamp_files=False, # stable filenames for CI artifacts
    ),
)
evaluator = AgentEvaluator(compiled_graph, collector, config=config)
report = await evaluator.evaluate(EVAL_SET)

json_report, html and console are on by default; junit_xml is off, and timestamp_files defaults to True. See Reports for the report formats.

Layer evaluation in CI

Run cheap deterministic checks first and the model-backed ones later, so a broken graph fails fast before you spend on judge or simulator calls.

flowchart TD
    A["Unit tests: mocked model, graph logic"]
    B["QuickEval smoke checks"]
    C["AgentEvaluator suites: full criteria"]
    D["Simulation scenarios: multi-turn goals"]
    E["Upload JSON, HTML and JUnit reports"]
    A --> B --> C --> D --> E

Use evaluation when you need to know whether the agent called the right tools in the right order, whether the response matches the expected answer, or whether a conversation reached its goals. Do not use it to test deterministic graph structure; unit tests are faster and free of model cost.

What to try next

  • Run the same cases with EvalPresets.comprehensive(threshold=0.3, use_llm_judge=False) and compare the criteria count against EvalPresets.quick_check(), as test5 does.
  • Aggregate several eval sets with EvaluationRunner.
  • Add a case with ToolCall arguments and CriterionConfig.trajectory(check_args=True).
  • Run evals from the command line with Run evals, or inside pytest with Evals in pytest.

For the concepts behind this example, read Evaluation. For the next production example, see Graceful shutdown.

Frequently asked questions

Why are the example tests skipped?
Most test modules call pytest.skip at import time so a root-level pytest run does not call a live model. Remove that call in the module you want to run.
Which API key does the example need?
The shared weather agent uses the Google provider with gemini-2.5-flash, so set GEMINI_API_KEY or GOOGLE_API_KEY in your environment. The user simulator also defaults to a Gemini model.
When should I use evaluation instead of unit tests?
Use unit tests for deterministic graph logic with a mocked model. Use evaluation when you need to score a real agent's tool calls, trajectory, response quality or multi-turn goals.
Last updated for v0.10.0Edit this page on GitHubReport an issue