Testing and QA
In this sectionHow to test and evaluate 10xGraph agents: mocked unit tests with TestAgent and QuickTest, and scored evaluation sets and simulated users with the eval command.
- 7 pages
- About 42 min to read all
Testing and QA in 10xGraph means two separate layers in one package, tenxgraph.qa. Unit tests check graph logic and tool routing with mocked models, so they run in milliseconds and cost nothing. Evaluations run the real agent against test cases or a simulated user and score the results. This section is for engineers who need to know an agent change did not break behavior, and who want that check to gate a CI pipeline.
Start here
Start with the unit-testing guide. It shows TestAgent, QuickTest and MockToolRegistry, which let you assert that the right tool was called with the right arguments without any LLM call. Write these first: they are fast and catch most routing mistakes.
Then read the evaluation guide for scored runs. Eval sets define the cases and expected tool sequences, criteria lists what can be scored, presets gives ready-made configurations, and reports covers the HTML, JSON and JUnit output. For multi-turn behavior, user simulation lets a model play the user.
Testing pairs with the runtime guarantees in Replay-safe tools: mocked tool registries are a good way to check that a tool with side effects, such as a refund, is only invoked when intended. For wiring into CI, see How to run tests and How to run evaluations.
The two layers compared
| Item | Unit testing | Evaluation |
|---|---|---|
| Goal | Verify graph logic and tool routing | Measure response quality and agent behaviour |
| LLM calls | None, fully mocked | Optional (LLM-as-judge criteria) |
| Speed | Fast (milliseconds per case) | Slower (real inference per case) |
| Entry point | agentflow test CLI or pytest directly |
agentflow eval CLI or AgentEvaluator |
| Output | pytest pass/fail + coverage | HTML + JSON report with per-criterion scores |
Both layers are independent, you can use one without the other, or run them together in CI.
Unit testing
The unit-testing layer lets you test the graph logic of your agent without making any LLM API calls:
TestAgent, drops into any node, returns predefined responses, records every call.QuickTest, one-liner helpers for single-turn, multi-turn, and tool-call scenarios.MockToolRegistry, registers mock tool functions and tracks all invocations.TestResult, fluent assertion helpers on top of the raw graph output.agentflow test, CLI wrapper around pytest that reads defaults from10xgraph.json.
Evaluation
The evaluation layer runs your real (or staging) agent against test cases and scores results across multiple criteria. It supports two modes of testing:
Fixed test cases, you define the query and expected output:
EvalSetBuilder, fluent API for defining test cases with expected responses and tool sequences.EvalConfig/EvalPresets, configure which criteria to use and at what thresholds.EvalPresetsprovides one-line ready-made configs.- Criteria, ten built-in criteria covering tool accuracy, response quality, hallucination, factual accuracy, safety, and custom rubrics.
AgentEvaluator, orchestrates execution and scoring. Supports sequential and parallel case runs.- Reports, HTML dashboard, JSON, and JUnit XML output.
User simulation, an LLM plays the user and drives real conversations:
ConversationScenario, define goals and a conversation plan. The simulator generates realistic user messages turn by turn.UserSimulator, LLM-powered user agent. Stops when all goals are achieved ormax_turnsis reached.BatchSimulator, runs multiple scenarios concurrently.SimulationGoalsCriterion, scores the full conversation transcript against stated goals.
agentflow eval, CLI that auto-discovers eval files, runs all cases from all files in a flat parallel pool, and always writes reports.
CLI commands at a glance
# Run the test suite
agentflow test
# Run with coverage
agentflow test --coverage --html
# Run evaluations (sequential)
agentflow eval
# Run evaluations in parallel across all cases and files
agentflow eval --parallel --max-concurrency 8
# Evaluate a specific file and open the report
agentflow eval evals/my_agent_eval.py --open
# Enforce a pass-rate threshold (useful in CI)
agentflow eval --threshold 0.8See also:
All pages in Testing and QA
Unit testing
Evaluation
- EvaluationEvaluate 10xGraph agents with eval sets, criteria presets, user simulation and parallel runs, then run them in CI with the 10xgraph eval command.7 min
- Building Eval SetsHow to build evaluation datasets with EvalSetBuilder — single-turn cases, multi-turn conversations, tool call assertions, and loading from files.4 min
- Evaluation CriteriaThe 10xGraph evaluation criteria: tool and node matching, ROUGE, semantic response, LLM-as-judge, rubrics, factual accuracy, hallucination, and safety.8 min
- Eval Presets & ConfigurationReady-made EvalPresets (tool_usage, response_quality, conversation_flow, comprehensive, safety_check, quick_check) and how to build a custom EvalConfig.4 min
- Evaluation ReportsHow 10xGraph evaluation reports work — HTML dashboard, JSON output, JUnit XML for CI, ReporterConfig, and how to interpret results.5 min
- User SimulationTest agents with the 10xGraph user simulator and goal-driven conversations using get_scenarios(), UserSimulator, BatchSimulator, and SimulationGoalsCriterion.9 min