Evaluation
In shortRun evaluations to measure agent quality: correct tool usage, accurate responses, and safety.
- 4 min read
- 5 sections
- Updated
- v0.10.0
- Markdown
Evaluation runs your agent against test cases and scores the results. Unlike unit tests, evaluations measure quality: whether the agent called the right tools, gave a semantically correct response, avoided hallucinations, and stayed safe.
The evaluation stack has four layers:
EvalSet / Scenarios → AgentEvaluator / UserSimulator → Criteria → EvalReport
↑ ↑ ↑ ↓
Test cases Runs the graph Scores results HTML / JSONQuick start
1. Define test cases
Create an EvalSet with test cases using EvalSetBuilder:
from tenxgraph.qa.evaluation import EvalSetBuilder
eval_set = (
EvalSetBuilder("weather-agent")
.add_tool_test(
query="What is the weather in London?",
tool_name="get_weather",
tool_args={"location": "London"},
expected_response="London",
case_id="weather_london",
)
.add_tool_test(
query="Weather in Tokyo?",
tool_name="get_weather",
tool_args={"location": "Tokyo"},
expected_response="Tokyo",
case_id="weather_tokyo",
)
.build()
)2. Run programmatically
The graph must be compiled with the collector’s callback manager so tool calls and node visits are captured. Build your uncompiled StateGraph (graph below), wire the collector in at compile time, then pass the compiled graph to AgentEvaluator and await the result:
from tenxgraph.qa.evaluation import (
AgentEvaluator,
TrajectoryCollector,
make_trajectory_callback,
)
from tenxgraph.qa.evaluation.config.presets import EvalPresets
collector, callback_manager = make_trajectory_callback(
TrajectoryCollector(capture_all_events=True)
)
your_graph = graph.compile(callback_manager=callback_manager)
config = EvalPresets.tool_usage(threshold=0.6)
evaluator = AgentEvaluator(your_graph, collector, config=config)
report = await evaluator.evaluate(eval_set)
print(f"Pass rate: {report.summary.pass_rate:.0%}")3. Quick one-liner with QuickEval
For a single test without building an EvalSet, use QuickEval.check():
from tenxgraph.qa.evaluation import QuickEval
report = await QuickEval.check(
graph=your_graph, # compiled as in step 2
collector=collector,
query="Weather in London?",
expected_response_contains="sunny",
expected_tools=["get_weather"],
)4. Or use the CLI
Put eval files in an evals/ directory and run the command:
10xgraph evalThe CLI discovers *_eval.py and eval_*.py files, loads your agent from 10xgraph.json, runs all cases, and writes HTML and JSON reports to eval_reports/ automatically. For full CLI options, see How to run evaluations.
Core concepts
EvalSet and EvalCase
An EvalSet is a named collection of test cases (EvalCase objects). Each case defines a user query (or multi-turn conversation), the expected response, and optionally the expected tool calls and node visit order.
See Building eval sets for the full API: single-turn, multi-turn, and trajectory-based cases.
Criteria
A criterion scores one evaluation case by comparing the agent’s trajectory and response against the expected outcome. Each criterion returns a score between 0 and 1. A case passes if all criteria meet their thresholds.
10xGraph provides 12 criteria:
| Criterion | Type | What it checks |
|---|---|---|
tool_name_match |
No-LLM | Tool names called match expected |
trajectory |
No-LLM | Tool sequence matches (EXACT / IN_ORDER / ANY_ORDER) |
node_order |
No-LLM | Graph nodes visited in expected order |
rouge_match |
No-LLM | ROUGE-1 token overlap between actual and expected response |
contains_keywords |
No-LLM | Required keywords appear in the response |
response_match |
LLM judge | Semantic equivalence of actual and expected response |
llm_judge |
LLM judge | Same semantic check, reported separately |
rubric_based |
LLM judge | Your own written grading rules |
factual_accuracy |
LLM judge | Factual correctness of stated facts |
hallucination |
LLM judge | Is the response grounded in tool results? |
safety |
LLM judge | Safety across harmful content, hate speech, privacy, misinformation |
simulation_goals |
LLM judge | Goal achievement in a multi-turn user simulation |
See Criteria reference for details and thresholds.
EvalConfig and EvalPresets
EvalConfig specifies which criteria to run and their thresholds. EvalPresets offers ready-made configs:
from tenxgraph.qa.evaluation.config.presets import EvalPresets
config = EvalPresets.tool_usage(threshold=0.6) # No LLM: tool names + sequence
config = EvalPresets.response_quality(threshold=0.7) # LLM judge on response accuracy
config = EvalPresets.quick_check() # ROUGE-only, no LLM cost
config = EvalPresets.comprehensive(threshold=0.8) # Tool, trajectory, ROUGE and LLM-judged criteriaSee Presets and configuration for how to build custom configs.
User simulation
UserSimulator uses an LLM to play a user role and drive dynamic multi-turn conversations with your agent. You define a ConversationScenario with goals; the simulator generates messages turn by turn and scores goal achievement.
from tenxgraph.qa.evaluation import ConversationScenario, UserSimulatorConfig
SIMULATOR_CONFIG = UserSimulatorConfig(
model="gemini/gemini-2.5-flash",
max_invocations=8,
)
def get_scenarios() -> list[ConversationScenario]:
return [
ConversationScenario(
scenario_id="travel_planning",
description="User planning a trip wants weather and packing advice",
starting_prompt="Hi! I'm planning a trip to Paris this weekend.",
goals=[
"User receives weather information for Paris",
"User gets clothing or packing advice",
],
max_turns=8,
),
]See User simulation for the full API and the get_scenarios() protocol.
Reports
By default, an evaluation run writes an HTML visual dashboard (pass rates, criterion scores, failure details) and a JSON report to eval_reports/. JUnit XML for CI integrations is opt-in through ReporterConfig(junit_xml=True).
See Reports for output formats and CI setup.
Running evaluations
For quick local evaluation, use AgentEvaluator or QuickEval as shown above. For repeatable, production evaluations with CLI commands, parallel execution, configuration, and CI integration, see How to run evaluations.
Next steps
- Building eval sets: define test cases and multi-turn scenarios
- Criteria reference: all 12 criteria explained
- Presets and configuration: ready-made configs and custom thresholds
- User simulation: LLM-driven multi-turn testing
- Reports: HTML, JSON, JUnit XML output formats
- How to run evaluations: CLI commands, parallel runs, CI integration
Frequently asked questions
- How do evaluations differ from unit tests?
- Evaluations measure agent quality: whether it called the right tools, gave correct responses, and avoided hallucinations. Unit tests verify code behavior.
- Can I run evaluations without LLM costs?
- Yes. Tool names, trajectory, node order, ROUGE and keywords run without an LLM (for example `EvalPresets.tool_usage()` or `quick_check()`). The default `EvalConfig` includes the LLM-judged `response_match`.
- Can I run evaluations from code or just the CLI?
- Both. Use AgentEvaluator or QuickEval for code; use `10xgraph eval` for CLI discovery and parallel runs.