Building Eval Sets
In shortHow to build evaluation datasets with EvalSetBuilder, single-turn cases, multi-turn conversations, tool call assertions, and loading from files.
- 7 min read
- 7 sections
- Updated
- v0.10.0
- Markdown
An EvalSet is a named collection of test cases, each defining a scenario that an agent should handle in a specific way. Each case specifies user input, expected behavior (response text, tool calls, node visits), and metadata. You build eval sets using the fluent EvalSetBuilder API, load them from JSON files, or construct them directly.
Eval sets are the input to evaluations: you pass an EvalSet to the evaluator, which runs your agent against each case and compares the actual behavior to the expected behavior using criteria. The result is a report showing how many cases passed, which criteria failed, and why.
Creating eval sets with EvalSetBuilder
The EvalSetBuilder class provides a fluent, chainable API for building eval sets incrementally. Every call to add_case, add_multi_turn, or add_tool_test returns the builder itself, so you can chain many calls before calling build().
Single-turn case
The simplest eval case is a single user message and an expected response. Use add_case:
from tenxgraph.qa.evaluation import EvalSetBuilder
eval_set = (
EvalSetBuilder("customer-support")
.add_case(
query="How do I reset my password?",
expected="visit the account settings page",
case_id="reset_password",
)
.build()
)The query is the user’s input. The expected string is the response the agent should produce (used by response-matching criteria like ExactMatchCriterion and RougeMatchCriterion). The case_id is an optional identifier; if omitted, it is auto-generated. Other optional fields: name and description for human-readable labels in reports.
Adding expected tool calls
Most agents use tools. Specify which tools the agent should call with expected_tools. Pass either tool names as strings or ToolCall objects for finer control:
from tenxgraph.qa.evaluation import EvalSetBuilder, ToolCall
eval_set = (
EvalSetBuilder("weather-agent")
.add_case(
query="Weather in London?",
expected="The weather in London is sunny.",
expected_tools=["get_weather"], # tool name only, args not checked
)
.add_case(
query="Weather in Tokyo?",
expected="Raining in Tokyo.",
expected_tools=[
ToolCall(name="get_weather", args={"location": "Tokyo"}),
], # with expected args
)
.build()
)When ToolCall objects include args, the trajectory matching criterion can optionally verify that the agent passed the correct arguments, controlled by setting check_args=True in the criterion config. By default, argument checking is off.
Verifying node execution order
Some tests require that the agent visit nodes in a specific sequence. Use expected_node_order to define the exact path through the graph:
.add_case(
query="Search and summarise the latest AI news",
expected="Here is a summary of recent developments...",
expected_tools=["search_web"],
expected_node_order=["MAIN", "TOOL", "MAIN"],
)The evaluator will check that nodes were visited in this exact sequence. This is useful for ensuring graph logic is working as expected (e.g., agent -> tool -> agent -> END).
Shortcut for tool-focused tests
When the primary concern is that a specific tool is called with specific arguments, use add_tool_test. It is a convenience method that automatically sets the expected response to a placeholder:
.add_tool_test(
query="What is the weather in Berlin?",
tool_name="get_weather",
tool_args={"location": "Berlin"},
expected_response="Berlin", # optional; defaults to "Result from get_weather"
case_id="berlin_weather",
)This is equivalent to calling add_case with expected_tools=[ToolCall(name="get_weather", args={"location": "Berlin"})], but shorter when you do not care about the final response text.
Multi-turn conversations
Eval cases can also be multi-turn conversations, where the agent receives multiple user messages and must maintain state across turns:
.add_multi_turn(
conversation=[
("Hello", "Hi! How can I help?"),
("What can you do?", "I can check weather, search the web, and more."),
("Check weather in Paris", "It is 18°C in Paris."),
],
expected_tools=["get_weather"],
case_id="multi_turn_weather",
)The conversation parameter is a list of (user_query, expected_response) tuples. The evaluator will send each user message in sequence and accumulate tool calls, node visits, and messages across all turns. Criteria then score that entire conversation once, not once per turn. The expected_tools is attached to the first turn; trajectory criteria will check for these tools to be called at any point in the conversation.
Quick builders from pairs
For rapid prototyping, use EvalSetBuilder.quick to create eval sets from a series of (query, expected) pairs:
from tenxgraph.qa.evaluation import EvalSetBuilder
eval_set = EvalSetBuilder.quick(
("Hello", "Hi!"),
("What is 2+2?", "4"),
("Capital of France?", "Paris"),
)Each pair becomes a single-turn case with auto-generated IDs. This method returns a built EvalSet directly (not a builder).
Creating from conversation logs
If you have conversation logs in a standard format (list of dicts with "user" and "assistant" keys), use from_conversations:
conversations = [
{"user": "Hello", "assistant": "Hi!"},
{"user": "Bye", "assistant": "Goodbye!"},
]
eval_set = EvalSetBuilder.from_conversations(conversations, name="smoke-tests")This method also returns a built EvalSet directly.
Persisting eval sets to files
Eval sets are Pydantic models that serialize to JSON. You can save and load them to share across teams or version-control them.
Saving to JSON
Build and save in one operation:
eval_set = (
EvalSetBuilder("weather-tests")
.add_case(query="London weather", expected="sunny")
.save("evals/weather.json")
)Or save an existing eval set:
eval_set.to_file("evals/weather.json")Both methods write a JSON file with all cases and metadata. The JSON is human-readable and can be edited by hand.
Loading from JSON
Load an eval set from disk:
from tenxgraph.qa.evaluation import EvalSet
eval_set = EvalSet.from_file("evals/weather.json")Loading and modifying
You can also load an eval set into a builder, add more cases, and rebuild:
builder = EvalSetBuilder.from_file("evals/weather.json")
builder.add_case(query="Weather in Seoul?", expected="Seoul weather")
eval_set = builder.build()This pattern is useful for extending existing eval sets without manually editing JSON.
Constructing eval cases directly
For advanced use cases, construct EvalCase objects without the builder using the class factory methods single_turn and multi_turn. This gives you full control over all fields:
Single-turn case
from tenxgraph.qa.evaluation import EvalCase, ToolCall
case = EvalCase.single_turn(
eval_id="london_weather",
user_query="Weather in London?",
expected_response="It is sunny in London.",
expected_tools=[ToolCall(name="get_weather", args={"location": "London"})],
expected_node_order=["MAIN", "TOOL", "MAIN"],
name="London weather check",
description="Verifies the weather tool is called for London queries",
)Multi-turn case
case = EvalCase.multi_turn(
eval_id="multi_weather",
conversation=[
("Weather in London?", "Sunny."),
("And in Tokyo?", "Rainy."),
],
expected_tools=[ToolCall(name="get_weather")],
)You can then add these cases to an EvalSet manually:
eval_set = EvalSet(name="my-cases", eval_cases=[case])Understanding the data model
EvalSet structure
An EvalSet is a container of evaluation cases with metadata:
| Attribute | Type | Description |
|---|---|---|
eval_set_id |
str |
Unique ID for this eval set (UUID, auto-generated) |
name |
str |
Human-readable name shown in reports |
description |
str |
Optional description of what this set tests |
eval_cases |
list[EvalCase] |
The test cases |
metadata |
dict[str, Any] |
Free-form metadata for your use |
EvalCase structure
Each case in an eval set represents one test scenario:
| Attribute | Type | Description |
|---|---|---|
eval_id |
str |
Unique ID for this case |
name |
str |
Optional human-readable name |
description |
str |
Optional details about what the case tests |
conversation |
list[Invocation] |
One invocation per turn (single-turn case has one) |
session_input |
SessionInput |
Initial session config: app_name, user_id (default "test_user"), state, config |
tags |
list[str] |
Optional tags used by EvalSet.filter_by_tags() |
metadata |
dict[str, Any] |
Free-form metadata; read by some criteria (e.g., hallucinations_v1 expects "context") |
Invocation structure
An Invocation is a single turn in a conversation. The builder’s add_case and add_multi_turn methods create these automatically, but you can inspect them:
| Attribute | Type | Description |
|---|---|---|
user_content |
MessageContent |
The user’s message for this turn |
expected_tool_trajectory |
list[ToolCall] |
Tools expected to be called during this turn |
expected_node_order |
list[str] |
Nodes expected to be visited during this turn |
expected_final_response |
MessageContent | None |
Expected assistant response for this turn (or None) |
To read the expected response text of a single-turn case:
case.conversation[0].expected_final_response.get_text()ToolCall structure
ToolCall represents a tool invocation, either expected or actual:
from tenxgraph.qa.evaluation import ToolCall
# Tool name only (args not checked in evaluation)
tool_call = ToolCall(name="get_weather")
# With arguments (args can be checked when check_args=True)
tool_call = ToolCall(name="get_weather", args={"location": "London"})The args field defaults to {}. Argument checking is disabled by default; enable it per criterion in the evaluation config with CriterionConfig.trajectory(check_args=True).
Common patterns
Case with tags
Use tags to group related cases for filtering:
builder = EvalSetBuilder("customer-support")
for topic in ["password", "billing", "refund"]:
builder.add_case(
query=f"Help with {topic}",
expected="Here is how to...",
case_id=f"case_{topic}",
)
# Add tags to the case after building (requires manual construction)
# Or build cases with tags by constructing EvalCase directly
case = EvalCase.single_turn(
eval_id="support_billing",
user_query="How do I get a refund?",
expected_response="You can request a refund...",
)
case.tags = ["billing", "support"]Case with session settings
session_input holds per-case session settings. The evaluator passes session_input.user_id and the entries of session_input.config into the run config for each case. For example, to run a case as a specific user:
case = EvalCase.single_turn(
eval_id="user_context_test",
user_query="What can you tell me?",
expected_response="Based on your account...",
)
case.session_input.user_id = "premium_user"
case.session_input.config = {"tier": "premium"}The session_input.state field exists on the model, but the evaluator does not currently apply it to the agent’s state.
Next steps
- Criteria reference: the scoring criteria available and how to use them
- Presets: ready-made evaluation configurations for common patterns
- Running evaluations: how to execute eval sets with
10xgraph evaland the API - Evaluation overview: conceptual background on evaluation as a testing methodology
Frequently asked questions
- How do I specify what tool calls an agent should make?
- Use `expected_tools` with a list of tool names or `ToolCall` objects. The evaluator will check that the agent called these tools during execution.
- Can I check that tool arguments match exactly?
- Yes. Include `args` in `ToolCall` objects, then enable `check_args=True` in the criterion config when running evaluations.
- What's the difference between a single-turn and multi-turn case?
- Single-turn cases have one user message and one expected response. Multi-turn cases contain a conversation with multiple exchanges, and criteria score the accumulated tool calls and messages across all turns.