Evaluation

In shortRun evaluations to measure agent quality: correct tool usage, accurate responses, and safety.

  • 4 min read
  • 5 sections
  • Updated
  • v0.10.0
  • Markdown

Evaluation runs your agent against test cases and scores the results. Unlike unit tests, evaluations measure quality: whether the agent called the right tools, gave a semantically correct response, avoided hallucinations, and stayed safe.

The evaluation stack has four layers:

plaintext
EvalSet / Scenarios  →  AgentEvaluator / UserSimulator  →  Criteria  →  EvalReport
       ↑                           ↑                            ↑              ↓
  Test cases               Runs the graph               Scores results   HTML / JSON

Quick start

1. Define test cases

Create an EvalSet with test cases using EvalSetBuilder:

Python
from tenxgraph.qa.evaluation import EvalSetBuilder

eval_set = (
    EvalSetBuilder("weather-agent")
    .add_tool_test(
        query="What is the weather in London?",
        tool_name="get_weather",
        tool_args={"location": "London"},
        expected_response="London",
        case_id="weather_london",
    )
    .add_tool_test(
        query="Weather in Tokyo?",
        tool_name="get_weather",
        tool_args={"location": "Tokyo"},
        expected_response="Tokyo",
        case_id="weather_tokyo",
    )
    .build()
)

2. Run programmatically

The graph must be compiled with the collector’s callback manager so tool calls and node visits are captured. Build your uncompiled StateGraph (graph below), wire the collector in at compile time, then pass the compiled graph to AgentEvaluator and await the result:

Python
from tenxgraph.qa.evaluation import (
    AgentEvaluator,
    TrajectoryCollector,
    make_trajectory_callback,
)
from tenxgraph.qa.evaluation.config.presets import EvalPresets

collector, callback_manager = make_trajectory_callback(
    TrajectoryCollector(capture_all_events=True)
)
your_graph = graph.compile(callback_manager=callback_manager)
config = EvalPresets.tool_usage(threshold=0.6)

evaluator = AgentEvaluator(your_graph, collector, config=config)
report = await evaluator.evaluate(eval_set)

print(f"Pass rate: {report.summary.pass_rate:.0%}")

3. Quick one-liner with QuickEval

For a single test without building an EvalSet, use QuickEval.check():

Python
from tenxgraph.qa.evaluation import QuickEval

report = await QuickEval.check(
    graph=your_graph,  # compiled as in step 2
    collector=collector,
    query="Weather in London?",
    expected_response_contains="sunny",
    expected_tools=["get_weather"],
)

4. Or use the CLI

Put eval files in an evals/ directory and run the command:

Terminal
10xgraph eval

The CLI discovers *_eval.py and eval_*.py files, loads your agent from 10xgraph.json, runs all cases, and writes HTML and JSON reports to eval_reports/ automatically. For full CLI options, see How to run evaluations.


Core concepts

EvalSet and EvalCase

An EvalSet is a named collection of test cases (EvalCase objects). Each case defines a user query (or multi-turn conversation), the expected response, and optionally the expected tool calls and node visit order.

See Building eval sets for the full API: single-turn, multi-turn, and trajectory-based cases.

Criteria

A criterion scores one evaluation case by comparing the agent’s trajectory and response against the expected outcome. Each criterion returns a score between 0 and 1. A case passes if all criteria meet their thresholds.

10xGraph provides 12 criteria:

Criterion Type What it checks
tool_name_match No-LLM Tool names called match expected
trajectory No-LLM Tool sequence matches (EXACT / IN_ORDER / ANY_ORDER)
node_order No-LLM Graph nodes visited in expected order
rouge_match No-LLM ROUGE-1 token overlap between actual and expected response
contains_keywords No-LLM Required keywords appear in the response
response_match LLM judge Semantic equivalence of actual and expected response
llm_judge LLM judge Same semantic check, reported separately
rubric_based LLM judge Your own written grading rules
factual_accuracy LLM judge Factual correctness of stated facts
hallucination LLM judge Is the response grounded in tool results?
safety LLM judge Safety across harmful content, hate speech, privacy, misinformation
simulation_goals LLM judge Goal achievement in a multi-turn user simulation

See Criteria reference for details and thresholds.

EvalConfig and EvalPresets

EvalConfig specifies which criteria to run and their thresholds. EvalPresets offers ready-made configs:

Python
from tenxgraph.qa.evaluation.config.presets import EvalPresets

config = EvalPresets.tool_usage(threshold=0.6)       # No LLM: tool names + sequence
config = EvalPresets.response_quality(threshold=0.7) # LLM judge on response accuracy
config = EvalPresets.quick_check()                   # ROUGE-only, no LLM cost
config = EvalPresets.comprehensive(threshold=0.8)    # Tool, trajectory, ROUGE and LLM-judged criteria

See Presets and configuration for how to build custom configs.

User simulation

UserSimulator uses an LLM to play a user role and drive dynamic multi-turn conversations with your agent. You define a ConversationScenario with goals; the simulator generates messages turn by turn and scores goal achievement.

Python
from tenxgraph.qa.evaluation import ConversationScenario, UserSimulatorConfig

SIMULATOR_CONFIG = UserSimulatorConfig(
    model="gemini/gemini-2.5-flash",
    max_invocations=8,
)

def get_scenarios() -> list[ConversationScenario]:
    return [
        ConversationScenario(
            scenario_id="travel_planning",
            description="User planning a trip wants weather and packing advice",
            starting_prompt="Hi! I'm planning a trip to Paris this weekend.",
            goals=[
                "User receives weather information for Paris",
                "User gets clothing or packing advice",
            ],
            max_turns=8,
        ),
    ]

See User simulation for the full API and the get_scenarios() protocol.

Reports

By default, an evaluation run writes an HTML visual dashboard (pass rates, criterion scores, failure details) and a JSON report to eval_reports/. JUnit XML for CI integrations is opt-in through ReporterConfig(junit_xml=True).

See Reports for output formats and CI setup.


Running evaluations

For quick local evaluation, use AgentEvaluator or QuickEval as shown above. For repeatable, production evaluations with CLI commands, parallel execution, configuration, and CI integration, see How to run evaluations.


Next steps

Frequently asked questions

How do evaluations differ from unit tests?
Evaluations measure agent quality: whether it called the right tools, gave correct responses, and avoided hallucinations. Unit tests verify code behavior.
Can I run evaluations without LLM costs?
Yes. Tool names, trajectory, node order, ROUGE and keywords run without an LLM (for example `EvalPresets.tool_usage()` or `quick_check()`). The default `EvalConfig` includes the LLM-judged `response_match`.
Can I run evaluations from code or just the CLI?
Both. Use AgentEvaluator or QuickEval for code; use `10xgraph eval` for CLI discovery and parallel runs.
Last updated for v0.10.0Edit this page on GitHubReport an issue