Testing and evaluation

In this sectionBuild confidence in agent behavior: unit test graph logic with TestAgent and mocks, then evaluate response quality and tool use with criteria and LLM-as-judge scoring.

  • 10 pages
  • About 74 min to read all

Testing in 10xGraph comes in two independent layers, both in the tenxgraph.qa package: unit tests that verify graph logic with mocked models, and evaluations that measure real agent behavior against criteria. When you build an agent, you start with unit tests to catch routing mistakes in milliseconds with zero cost. When you ship or iterate, you add evaluations to prove the agent still meets your quality bar and to catch regressions.

This section is for teams that want confidence their agents work correctly, and who want that confidence expressed in code that gates a CI pipeline.

Quick start

Begin with unit tests: they show you TestAgent, QuickTest, and MockToolRegistry, which let you assert the right tool was called with the right arguments without any LLM. Write these first. They are fast (milliseconds per case), cost nothing, and catch most routing mistakes.

Then move to evaluation for scored runs against real or staged models. Eval sets define the test cases and expected tool sequences. Criteria describes what can be scored (tool accuracy, response quality, hallucination, safety, and more). Presets give ready-made criterion configurations to speed up setup. Reports explains the HTML, JSON, and JUnit output formats.

For conversational agents, user simulation lets a model play the user and drives realistic multi-turn interactions. Integration with pytest is in evals in pytest, which runs evals as part of your test suite via helpers like eval_test and run_eval.

Wiring both layers into CI pipelines is covered in run tests and run evaluations.

Why two layers

Unit testing and evaluation serve different purposes and run at different speeds. The comparison shows where each fits:

Item Unit testing Evaluation
Goal Verify graph logic and tool routing Measure response quality and agent behavior
LLM calls None; fully mocked Yes (required by default)
Speed Milliseconds per case Seconds per case
Entry point TestAgent in pytest, or 10xgraph test CLI AgentEvaluator in code, or 10xgraph eval CLI
Output pytest pass/fail with coverage reports HTML dashboard, JSON, JUnit with criterion scores
Cost Free Charged by LLM provider per case
Best for Catch regressions early; gate commits Gate releases; measure production behavior

Use both together in CI: quick unit tests first to fail fast, then evaluations as a later stage. Or use one without the other if your needs fit only one layer.

Unit testing layer

Unit tests verify graph structure and tool routing without making any LLM API calls. The main classes are:

  • TestAgent: A drop-in replacement for the real Agent. Returns predefined responses you give it, records every LLM call, tracks tool names.
  • QuickTest: One-liner helpers for single-turn agent calls, multi-turn conversations, and tool-routing scenarios.
  • MockToolRegistry: Registers mock tool functions and tracks all invocations by name and arguments.
  • TestContext: Helper that sets up an isolated dependency container, in-memory store, and test graph factory.
  • MockMCPClient: Mock MCP client for testing MCP tool integrations without a real server.
  • TestResult: Fluent assertion helpers for the final response, tool calls, and message count.

See unit tests for detailed examples and run tests for CLI and CI integration.

Evaluation layer

Evaluation runs the real agent against test cases and scores results across one or more criteria. Two modes are supported:

Fixed test cases. You define the input and expected behavior:

  • EvalSet and EvalSetBuilder: Fluent API for defining test cases with expected responses and expected tool sequences (trajectory).
  • EvalConfig and EvalPresets: Configuration for which criteria to use and what score thresholds pass a case. EvalPresets provides one-line ready-made configurations for common patterns.
  • Criteria: Built-in criteria cover tool accuracy (tool name and arguments match), response quality (keyword presence, exact match, ROUGE similarity), trajectory (node order and tool sequence), safety, hallucination, factual accuracy, and custom rubrics. LLM-as-judge criteria score anything with a language model.
  • AgentEvaluator: Orchestrates running cases and scoring. Supports sequential and parallel execution.
  • Reports: HTML dashboard with per-criterion scores, JSON, and JUnit XML for CI systems.

User simulation. An LLM plays the user and drives real conversations:

  • ConversationScenario: Define goals and topics; the simulator generates realistic user messages turn by turn.
  • UserSimulator: The LLM-powered user. Stops when goals are achieved or max turns is reached.
  • SimulationGoalsCriterion: Scores the full conversation transcript against your stated goals.

See evaluation for a quick start, criteria for the full criterion reference, and run evaluations for CLI and CI integration.

Reading order

If you are new to testing agents, start here:

  1. Unit tests: write your first test with TestAgent
  2. Run tests: wire unit tests into CI
  3. Evaluation: run your agent against real cases
  4. Eval sets: define test cases
  5. Criteria: understand scoring
  6. Run evaluations: automate evaluations in CI

Then explore deeper topics like presets, user simulation, and evals in pytest as your needs grow.

Pair testing with Replay-safe tools: mock registries are a good way to verify that tools with side effects (like refunds or database deletes) are called only when intended and not replayed unexpectedly. See also Stream a graph for capturing execution in depth.

All pages in Testing and evaluation

Unit tests

  1. Unit TestingWrite fast, deterministic tests for agents using TestAgent, QuickTest, MockToolRegistry, and mock storage without making real LLM API calls.7 min
  2. Run TestsRun your 10xGraph project's test suite with 10xgraph test. Configure coverage thresholds and CI gates via 10xgraph.json.6 min

Evaluation

  1. EvaluationRun evaluations to measure agent quality: correct tool usage, accurate responses, and safety.4 min
  2. Building Eval SetsHow to build evaluation datasets with EvalSetBuilder, single-turn cases, multi-turn conversations, tool call assertions, and loading from files.7 min
  3. Evaluation CriteriaScore agent outputs with deterministic criteria or LLM judges: tools, nodes, ROUGE, accuracy, safety.8 min
  4. Evaluation Presets and ConfigurationReady-made evaluation presets for common scenarios and how to build custom EvalConfig for specific needs.9 min
  5. User SimulationTest agents with the 10xGraph user simulator and goal-driven conversations using get_scenarios(), UserSimulator, BatchSimulator, and SimulationGoalsCriterion.10 min
  6. Evaluation ReportsHow 10xGraph evaluation reports work, HTML dashboard, JSON output, JUnit XML for CI, ReporterConfig, and how to interpret results.5 min
  7. Run EvaluationsRun agent evaluations with 10xgraph eval: covers parallel execution, eval file protocols, EvalPresets, and CI integration.12 min
  8. Running Evaluations in pytestIntegrate agent evaluations into pytest to measure quality alongside unit tests. Use decorators and assertion helpers to verify agents meet quality thresholds.6 min