Run Evaluations

In shortRun agent evaluations with 10xgraph eval: covers parallel execution, eval file protocols, EvalPresets, and CI integration.

  • 12 min read
  • 10 sections
  • Updated
  • v0.10.0
  • Markdown

What is evaluation and why it matters

Evaluation is the measurement of your agent’s behavior against criteria that matter to your application. The 10xgraph eval command discovers evaluation files in your project, runs all test cases in a single async event loop, and produces machine-readable reports so you can track agent quality over time, gate merges on minimum pass rates, and identify regressions before they reach production.

Unlike unit tests (which verify isolated functions), evaluations score your entire agent graph end-to-end against realistic scenarios, criteria like correct tool usage or response accuracy, and can include LLM-based judges to assess semantic quality. Cases run sequentially unless you pass --parallel. The command generates both HTML dashboards and JSON results for CI tooling, and supports everything from quick local checks to comprehensive test suites with custom scoring rules.

Prerequisites

Your project must have been initialized with 10xgraph init. The standard project layout includes an evals/ directory where evaluation files live. When you choose the Production setup during initialization, the directory structure and a sample evaluation file are generated for you.

If you created a project without the Production setup, create an evals/ directory manually at the project root, next to 10xgraph.json.

Quick start

From the folder that contains 10xgraph.json:

Terminal
10xgraph eval

The CLI scans evals/ for files matching *_eval.py or eval_*.py, collects every case from every file into a single pool, runs them all, and writes timestamped reports to eval_reports/:

plaintext
eval_reports/
  <eval-set-id>_20260513_142301.html
  <eval-set-id>_20260513_142301.json

EvalSetBuilder assigns a random UUID as the eval set id, so that is what appears in the file name for a single set. When a run covers more than one eval set, the results are merged into one report named combined_eval_<timestamp>.

The HTML report shows a visual dashboard: summary pass rate, criterion scores, and per-case results. The JSON report is machine-readable for CI systems or custom analysis. A progress display updates as cases complete.

Running evaluations

Run a specific file or directory

To evaluate only certain cases, pass a path:

Terminal
# One file
10xgraph eval evals/weather_eval.py

# All files in a subdirectory
10xgraph eval evals/regression/

When a file is given, only that file runs. When a directory is given, all matching *_eval.py and eval_*.py files are discovered recursively. Results from several eval sets are merged into one combined report.

Run in parallel

By default all cases run sequentially. For faster feedback on large suites, use --parallel:

Terminal
10xgraph eval --parallel
10xgraph eval --parallel --max-concurrency 8

How it works: All cases from all files are collected into a flat pool before execution starts. A single asyncio event loop runs the entire pool under a concurrency semaphore capped at --max-concurrency (default 4). Cases complete out of order as they finish; this is expected and intentional.

The run ends with a summary line such as:

plaintext
Results: 47/50 passed (94.0%)

You can enable parallel by default in 10xgraph.json (see Configure evaluation defaults).

Reports and output

Every run produces two files in eval_reports/ (or a custom directory via --output):

File Format Purpose
<eval-id>_<timestamp>.html HTML + interactive Visual dashboard with summary, criterion bars, and per-case details for humans to review
<eval-id>_<timestamp>.json JSON Machine-readable results for CI tooling, dashboards, or programmatic analysis

The progress display shows each case as <file>::<case> with its status and duration. Both report files are written after all cases finish. To skip writing files and see only console output:

Terminal
10xgraph eval --no-report

To open the HTML report automatically in your browser:

Terminal
10xgraph eval --open

Exit codes and thresholds

The command exits with code 0 only when every case passes (100% pass rate) and 1 otherwise. For CI this already gates merges on any regression.

Terminal
10xgraph eval --threshold 0.8

--threshold makes the command print an error (“Pass rate … is below threshold …”) when the overall pass rate is under 80%. It does not relax the exit code: a run at 90% still exits 1. The threshold can also be set in 10xgraph.json (see below).

Custom output directory

Terminal
10xgraph eval --output ci/reports

Reports are written to ci/reports/ instead of the default eval_reports/.

Configure evaluation defaults in 10xgraph.json

Add an evaluation block to 10xgraph.json to set project-level defaults. CLI flags always take precedence over these settings.

JSON
{
  "agent": "graph.agent:app",
  "evaluation": {
    "directory": "evals",
    "output_dir": "eval_reports",
    "threshold": 0.75,
    "parallel": false,
    "max_concurrency": 4
  }
}
Field Default Description
directory "evals" Directory scanned when no TARGET argument is given
output_dir "eval_reports" Directory where HTML and JSON reports are written
threshold (none) Pass rate below which an error message is printed; the exit code is 1 for any pass rate under 100%
parallel false Run all cases concurrently instead of sequentially
max_concurrency 4 Maximum cases running simultaneously when parallel is true

Report filenames always include a timestamp; there is no configuration setting for it.

Enforce minimum quality in CI

YAML
# .github/workflows/ci.yml
- name: Run evaluations
  run: 10xgraph eval --parallel

The step fails (exit code 1) whenever any case fails, which gates the merge.

How criteria work and what the defaults are

Every evaluation case is scored against one or more criteria. A criterion is a rule that produces a pass/fail or a numeric score. The CLI applies default criteria when none are specified; you can override them per-file, globally, or use preset configurations.

Default criteria (applied when no config is specified)

When an eval file has no get_eval_config() or EVAL_CONFIG, and no confeval.py is found, these built-in defaults apply:

Criterion Threshold Purpose
tool_name_match 1.0 Agent must call the correct tool (exact match required)
rouge_match 0.5 Agent response must have at least 50% token overlap with expected response (fast, no LLM)
node_order 0.8 Agent must visit graph nodes in the expected order (score of at least 0.8)

These defaults are intentionally strict on tool correctness (1.0) but lenient on response content (0.5 ROUGE), because the order of tool calls often matters more than exact text matching. For applications where response wording must match precisely, you would override with stricter criteria.

Eval file protocols

An eval file is any module matching *_eval.py or eval_*.py. The CLI auto-detects which protocol you are using. Each protocol suits different evaluation needs.

get_eval_set(), standard fixed cases

The simplest protocol: define your test cases once, the CLI runs them repeatedly. You only provide the cases; the CLI loads the agent from 10xgraph.json, applies criteria, and produces results.

Python
# evals/weather_eval.py
from tenxgraph.qa.evaluation import EvalSet, EvalSetBuilder

def get_eval_set() -> EvalSet:
    return (
        EvalSetBuilder(name="weather-regression")
        .add_tool_test(
            query="What is the weather in London?",
            tool_name="get_weather",
            tool_args={"location": "London"},
            expected_response="London",
            case_id="london_forecast",
        )
        .add_tool_test(
            query="What is the weather in Tokyo?",
            tool_name="get_weather",
            tool_args={"location": "Tokyo"},
            expected_response="Tokyo",
            case_id="tokyo_forecast",
        )
        .build()
    )

When to use: Regression suites with fixed inputs and known good outputs. Best for deterministic behavior like tool selection, API behavior, or structured data extraction.

get_eval_config() or EVAL_CONFIG, per-file criteria

Override the criteria and thresholds for a specific file. The recommended approach is to use one of the preset configurations, which cover common patterns:

Python
from tenxgraph.qa.evaluation import EvalConfig, EvalSet, EvalSetBuilder
from tenxgraph.qa.evaluation.config.presets import EvalPresets

def get_eval_config() -> EvalConfig:
    return EvalPresets.tool_usage(threshold=0.6)

def get_eval_set() -> EvalSet:
    return (
        EvalSetBuilder(name="weather-regression")
        .add_tool_test(...)
        .build()
    )

Or as a constant instead of a function:

Python
from tenxgraph.qa.evaluation.config.presets import EvalPresets

EVAL_CONFIG = EvalPresets.tool_usage(threshold=0.6)

Available EvalPresets

Preset Default threshold What it checks
response_quality(threshold, use_llm_judge=True) 0.7 LLM response match plus LLM judge on response quality
tool_usage(threshold, strict=True, check_args=True) 1.0 Tool name match plus trajectory match (exact order by default, arguments checked)
conversation_flow(threshold) 0.8 LLM response match plus in-order tool trajectory
quick_check() 0.5 (fixed) Fast ROUGE token overlap, no LLM cost, instant feedback
comprehensive(threshold, use_llm_judge=True) 0.8 Tool name match and trajectory (fixed 1.0), ROUGE, plus LLM judge, factual accuracy, hallucination and safety
safety_check(threshold) 0.8 Hallucination and safety criteria via LLM judge

Presets that use an LLM judge accept an optional judge_model parameter (defaults to "gemini-2.5-flash"; tool_usage and quick_check use no judge):

Python
config = EvalPresets.response_quality(threshold=0.7, judge_model="gpt-4o")

You can combine multiple presets:

Python
from tenxgraph.qa.evaluation.config.presets import EvalPresets

def get_eval_config():
    return EvalPresets.combine(
        EvalPresets.tool_usage(threshold=0.7),
        EvalPresets.response_quality(threshold=0.6),
    )

confeval.py, global evaluation config

Place a file named exactly confeval.py in your evals/ directory (or the project root) to set a global default EvalConfig that applies to every eval file that does not define its own get_eval_config() or EVAL_CONFIG. If a file does provide its own config, that takes precedence.

The file must expose either a module-level EVAL_CONFIG variable or a callable get_eval_config() that returns an EvalConfig:

Python
# evals/confeval.py
from tenxgraph.qa.evaluation import CriteriaConfig, CriterionConfig, EvalConfig

EVAL_CONFIG = EvalConfig(
    criteria=CriteriaConfig(
        tool_name_match=CriterionConfig.tool_name_match(threshold=1.0),
        rouge_match=CriterionConfig.rouge_match(threshold=0.5),
        node_order=CriterionConfig.node_order(threshold=0.8),
    )
)

Or as a function:

Python
# confeval.py
from tenxgraph.qa.evaluation import CriteriaConfig, CriterionConfig, EvalConfig

def get_eval_config() -> EvalConfig:
    return EvalConfig(
        criteria=CriteriaConfig(
            tool_name_match=CriterionConfig.tool_name_match(threshold=1.0),
            rouge_match=CriterionConfig.rouge_match(threshold=0.5),
        )
    )

This is useful for enforcing consistent evaluation standards across many eval files without repeating the config in each one.

Annotated functions -> EvalSet, multiple eval sets per file

Any module-level function with return type annotation -> EvalSet is auto-discovered as an eval set. This lets you organize multiple related eval sets in one file:

Python
from tenxgraph.qa.evaluation import EvalSet, EvalSetBuilder
from tenxgraph.qa.evaluation.config.presets import EvalPresets

def get_eval_config():
    return EvalPresets.tool_usage(threshold=0.6)

def weather_cases() -> EvalSet:
    return EvalSetBuilder(name="weather").add_tool_test(...).build()

def booking_cases() -> EvalSet:
    return EvalSetBuilder(name="booking").add_tool_test(...).build()

Both weather_cases and booking_cases are discovered and run, and their cases are merged into the report. A file that has get_eval_set() uses that instead of annotation discovery. A function annotated -> EvalConfig is also picked up as the file’s config.

get_scenarios() or SCENARIOS, user simulator (dynamic conversations)

Use this protocol when you want the LLM to drive a multi-turn conversation against your agent, rather than using fixed prompt/response pairs. The simulator generates contextual follow-up messages after each agent response, letting you test realistic conversation flows.

You define the scenarios; the CLI handles running the simulator, scoring goal achievement, and writing results:

Python
# evals/simulator_eval.py
from tenxgraph.qa.evaluation import ConversationScenario, UserSimulatorConfig

# Optional: configure simulator model and settings for this file.
# If omitted, defaults are used (gemini-2.5-flash, max 10 invocations, temperature 0.7).
SIMULATOR_CONFIG = UserSimulatorConfig(
    model="gemini/gemini-2.5-flash",
    max_invocations=8,
    temperature=0.7,
)

def get_scenarios() -> list[ConversationScenario]:
    return [
        ConversationScenario(
            scenario_id="weather_travel",
            description="User planning a trip wants weather and packing advice",
            starting_prompt="Hi! I'm planning a trip to Paris this weekend.",
            conversation_plan=(
                "1. Ask about current weather in Paris\n"
                "2. Ask whether to bring a jacket\n"
                "3. Ask about outdoor activity timing"
            ),
            goals=[
                "User receives weather information for Paris",
                "User gets clothing or packing advice",
                "User learns about outdoor activity timing",
            ],
            max_turns=8,
        ),
        ConversationScenario(
            scenario_id="flight_booking",
            description="User wants help finding a flight",
            starting_prompt="I need to fly from London to New York next Friday.",
            goals=[
                "User receives flight options",
                "User gets pricing information",
            ],
            max_turns=10,
        ),
    ]

How it works:

  1. The CLI detects get_scenarios() or a SCENARIOS constant and switches to simulator mode.
  2. Each ConversationScenario becomes one eval case.
  3. The simulator drives up to max_turns turns, generating contextual user messages after each agent response.
  4. An LLM judge scores how many of the stated goals were achieved across the conversation.
  5. Pass/fail and results appear in the report like regular eval cases.

ConversationScenario fields:

Field Required Description
scenario_id Yes Unique ID for the scenario (appears in reports)
description No Human-readable name shown in the report
starting_prompt Yes First user message to start the conversation
conversation_plan No Hints to the simulator about how to progress
goals Yes List of outcomes the user wants to achieve (scored by LLM judge)
max_turns No Maximum conversation turns (default: 10)

SIMULATOR_CONFIG fields:

Field Default Description
model "gemini-2.5-flash" LLM used to generate user messages
max_invocations 10 Maximum turns per scenario
temperature 0.7 Temperature for user message generation

When to use: Multi-turn conversations where the exact dialogue matters. Useful for customer support agents, question-answering over sessions, or any scenario where realistic back-and-forth is important.

Config priority and precedence

When the same setting is configured in multiple places, this precedence applies (highest first):

plaintext
Run settings (parallel, max concurrency, threshold, output directory):
1. CLI flags
2. 10xgraph.json "evaluation" section
3. Built-in defaults

Criteria (the 10xgraph.json "evaluation" section holds no criteria):
1. Per-file config   get_eval_config() / EVAL_CONFIG inside the eval file
2. confeval.py       get_eval_config() / EVAL_CONFIG (global fallback)
3. Built-in defaults (tool_name_match 1.0, rouge_match 0.5, node_order 0.8)

A CLI flag overrides 10xgraph.json for run settings. For criteria, the eval file wins over confeval.py, which wins over the built-in defaults.

Common scenarios and recipes

Fast local check on a single file, open the report:

Terminal
10xgraph eval evals/weather_eval.py --open

Run one file sequentially, write the report, and open it in your browser for inspection.

Parallel run with 8 concurrent cases:

Terminal
10xgraph eval --parallel --max-concurrency 8

All cases run concurrently (up to 8 at a time) for faster feedback on large suites.

Strict CI gate at 80% pass rate:

In 10xgraph.json:

JSON
{
  "evaluation": {
    "threshold": 0.8,
    "parallel": true,
    "max_concurrency": 8
  }
}

Then in your CI:

Terminal
10xgraph eval

The command prints an error when the pass rate is below 80% and exits with code 1 (it also exits 1 for any pass rate under 100%).

Run only a regression suite in a subdirectory:

Terminal
10xgraph eval evals/regression/ --output reports/regression

All eval files under evals/regression/ run, reports go to reports/regression/.

Mix regular eval cases and user simulator in one run:

plaintext
evals/
  weather_eval.py        ← get_eval_set() protocol
  user_simulator_eval.py ← get_scenarios() protocol
Terminal
10xgraph eval --parallel

Both files are discovered. Cases and scenarios are collected into the same flat pool and run concurrently. Results appear in a single merged report.

Troubleshooting

“Eval directory ‘evals/’ not found”

The CLI cannot find your evaluation directory. Either:

  • Create an evals/ directory at the project root
  • Pass the correct path: 10xgraph eval path/to/evals
  • Run 10xgraph init and choose the Production setup to scaffold the standard layout

“No eval cases found”

Eval files were discovered but contained no test cases. Ensure:

  • At least one file exposes get_eval_set(), get_scenarios(), SCENARIOS, or a function annotated -> EvalSet
  • Files match *_eval.py or eval_*.py (rename if needed, or pass the file explicitly)

“File skipped with warning”

A file was found but does not expose any recognized entry point. Add get_eval_set() or get_scenarios() to the file.

Exit code 1 although most cases pass

The exit code is 1 whenever the pass rate is below 100%, including cases that errored. Check the summary line and the HTML report for cases with status error.

Simulator scenarios always fail

Ensure:

  • The agent is reachable (either app is exported in the eval file, or "agent" is set in 10xgraph.json)
  • Goals are specific enough for the LLM judge to verify. Vague goals like “have a conversation” will not score well
  • max_turns is large enough for the agent to satisfy all goals (increase if needed)

Different results on different runs

Simulator scenarios and LLM-based criteria include randomness (temperature > 0). Exact reproducibility is not guaranteed. For deterministic evaluation, use quick_check() or fixed get_eval_set() cases only.

Last updated for v0.10.0Edit this page on GitHubReport an issue