# Evaluation Reports

> How 10xGraph evaluation reports work — HTML dashboard, JSON output, JUnit XML for CI, ReporterConfig, and how to interpret results.

Source: https://10xgraph.com/docs/qa/evaluation/reports
Last updated: 2026-07-21

After every evaluation run, 10xGraph generates structured reports so you can understand what passed, what failed, and why. Reports are always generated by default — no flags required.

---

## Report types

| Format | File | Best for |
|---|---|---|
| HTML | `eval_reports/<id>_<timestamp>.html` | Human review — visual dashboard |
| JSON | `eval_reports/<id>_<timestamp>.json` | Programmatic analysis, dashboards |
| JUnit XML | `eval_reports/<id>_<timestamp>_junit.xml` | CI systems (GitHub Actions test summary) |
| Console | stdout | Quick local feedback during development |

---

## HTML report

The HTML report is a self-contained dashboard that opens in any browser. It contains:

- **Summary header** — eval set name, total cases, passed/failed count, overall pass rate
- **Criterion bar chart** — a breakdown of scores for each active criterion across all cases
- **Per-case detail** — expandable sections showing the query, actual response, expected response, and score for each criterion per case

Open it manually:

```bash
open eval_reports/weather-agent-regression_20260519_142301.html
```

Or open automatically after the run:

```bash
agentflow eval --open
```

---

## JSON report

The JSON report contains every piece of data from the run in a machine-readable format. You can consume it in post-processing scripts, ingest it into a metrics dashboard, or archive it for trend analysis.

Structure (simplified):

```json
{
  "eval_set_id": "weather-agent",
  "eval_set_name": "Weather Agent Regression",
  "summary": {
    "total_cases": 5,
    "passed_cases": 4,
    "failed_cases": 1,
    "error_cases": 0,
    "pass_rate": 0.8,
    "criterion_stats": {
      "tool_trajectory_avg_score": {
        "total": 5, "passed": 5, "failed": 0,
        "avg_score": 1.0, "pass_rate": 1.0
      },
      "rouge_match": {
        "total": 5, "passed": 4, "failed": 1,
        "avg_score": 0.72, "pass_rate": 0.8
      }
    }
  },
  "results": [
    {
      "eval_id": "london_weather",
      "name": "London weather",
      "passed": true,
      "criterion_results": [
        {"criterion": "tool_trajectory_avg_score", "score": 1.0, "passed": true, "threshold": 1.0},
        {"criterion": "rouge_match", "score": 0.65, "passed": true, "threshold": 0.5}
      ],
      "actual_response": "The weather in London is sunny.",
      "actual_tool_calls": [{"name": "get_weather", "args": {"location": "London"}}]
    }
  ]
}
```

Note that `criterion_results` is a **list**, and per-criterion aggregates live under `summary.criterion_stats` keyed by criterion name.

---

## JUnit XML report

JUnit XML output lets CI systems (GitHub Actions, Jenkins, GitLab CI) display eval results as structured test results — with pass/fail summaries, per-case details, and failure messages in the test tab.

Enable it via `ReporterConfig`:

```python
from tenxgraph.qa.evaluation import ReporterConfig, ReporterManager

manager = ReporterManager(
    ReporterConfig(
        output_dir="eval_reports",
        junit_xml=True,
        html=True,
        json_report=True,
    )
)
manager.run_all(report)
```

Or via the `agentflow eval` command — JUnit XML is off by default. You can enable it programmatically in your `run()` function.

GitHub Actions example using the JUnit XML output:

```yaml
- name: Run evaluations
  run: agentflow eval

- name: Publish eval results
  uses: EnricoMi/publish-unit-test-result-action@v2
  if: always()
  with:
    files: eval_reports/*.xml
```

---

## EvalReport object

When running programmatically, `AgentEvaluator.evaluate()` returns an `EvalReport`:

```python
from tenxgraph.qa.evaluation import AgentEvaluator, EvalPresets

evaluator = AgentEvaluator(app, collector, config=EvalPresets.tool_usage())
report = await evaluator.evaluate(eval_set)
```

### Summary

```python
summary = report.summary

summary.total_cases            # int — total number of cases
summary.passed_cases           # int — cases where all criteria passed
summary.failed_cases           # int — cases that failed a criterion without erroring
summary.error_cases            # int — cases that raised during the run
summary.pass_rate              # float 0.0–1.0
summary.criterion_stats        # dict[str, dict] — per-criterion aggregates
summary.total_duration_seconds # float
summary.total_token_usage      # TokenUsage — agent + judge tokens
```

Each entry in `criterion_stats` is keyed by the criterion's reported name and holds `total`, `passed`, `failed`, `avg_score` and `pass_rate`:

```python
for name, stats in report.summary.criterion_stats.items():
    print(f"{name}: avg {stats['avg_score']:.2f}, {stats['passed']}/{stats['total']} passed")
```

The counting invariant is `total_cases == passed_cases + failed_cases + error_cases` — an errored case is counted in `error_cases`, never also in `failed_cases`.

### Per-case results

`criterion_results` is a list of `CriterionResult` objects, not a dict:

```python
for result in report.results:
    print(result.eval_id, "PASS" if result.passed else "FAIL")

    for cr in result.criterion_results:
        print(f"  {cr.criterion}: {cr.score:.2f} ({'PASS' if cr.passed else 'FAIL'})")

    # Look one up by name
    traj = result.get_criterion_result("tool_trajectory_avg_score")
```

`result.failed_criteria` and `result.passed_criteria` are ready-made filters over the same list.

### Print to console

```python
from tenxgraph.qa.evaluation import print_report

print_report(report, verbose=False, use_color=True)
```

`report.format_summary()` returns the same kind of overview as a plain string if you would rather log it than print it.

---

## ReporterConfig reference

`ReporterConfig` controls all automatic reporting. Pass it to `ReporterManager` when generating reports manually.

```python
from tenxgraph.qa.evaluation import ReporterConfig

config = ReporterConfig(
    enabled=True,
    output_dir="eval_reports",
    console=True,
    json_report=True,
    html=True,
    junit_xml=False,
    verbose=True,
    include_details=True,
    include_trajectory=True,
    include_node_responses=True,
    include_actual_response=True,
    include_tool_call_details=True,
    timestamp_files=True,
)
```

| Field | Default | Description |
|---|---|---|
| `enabled` | `True` | Master switch — when `False`, no reporters run |
| `output_dir` | `eval_reports` | Directory for generated files |
| `console` | `True` | Print summary to stdout |
| `json_report` | `True` | Write JSON file |
| `html` | `True` | Write HTML file |
| `junit_xml` | `False` | Write JUnit XML file |
| `verbose` | `True` | Show all cases; when `False`, show only failures |
| `include_details` | `True` | Include per-criterion details in file reports |
| `include_trajectory` | `True` | Include tool call trajectory in JSON |
| `include_node_responses` | `True` | Include per-node intermediate data |
| `include_actual_response` | `True` | Include the agent's final response |
| `include_tool_call_details` | `True` | Include tool arguments and results |
| `timestamp_files` | `True` | Append timestamp so runs do not overwrite each other |

---

## Running reporters manually

```python
from tenxgraph.qa.evaluation import ReporterManager, ReporterConfig

manager = ReporterManager(
    ReporterConfig(
        output_dir="my_reports",
        html=True,
        json_report=True,
        console=False,
        timestamp_files=True,
    )
)

output = manager.run_all(report)

if output.html_path:
    print(f"HTML: {output.html_path}")
if output.json_path:
    print(f"JSON: {output.json_path}")
if output.has_errors:
    for name, err in output.errors:
        print(f"Reporter error [{name}]: {err}")
```

---

## Combined reports from multiple eval files

When `agentflow eval` runs more than one eval file, results from all files are merged into a single combined report. The merged report has `eval_set_id="combined_eval"` and contains all cases from all files.

If you run multiple files manually, merge them the same way:

```python
from tenxgraph.qa.evaluation import EvalReport

merged = EvalReport.create(
    eval_set_id="combined",
    results=[*report_a.results, *report_b.results],
    eval_set_name="Combined Evaluation",
)
```

`EvalReport.create()` recomputes `summary` from the results you pass in.

---

## Interpreting results

### Pass rate

A case passes only when **all enabled criteria** meet their thresholds. The pass rate is `passed_cases / total_cases`.

- `1.0` (100%) — every criterion passed for every case
- `< 1.0` — at least one case has at least one criterion below threshold

The `agentflow eval` CLI exits with code `1` whenever pass rate is below `1.0` or below the configured threshold. Code `0` means a perfect 100% pass rate (or threshold was met and no errors occurred).

### Diagnosing failures

1. Open the HTML report and look at the per-case criterion bars.
2. Find which criterion scored lowest.
3. If it is a trajectory criterion, the trajectory section shows which tool calls were expected vs actual.
4. If it is an LLM criterion, the detail section shows the judge's reasoning (when `include_details=True`).
5. Adjust the eval set (add missing tool expectations), tighten the prompt, or revise the threshold.

---

## Next steps

- [Eval sets](/docs/qa/evaluation/eval-set) — defining test cases
- [Criteria reference](/docs/qa/evaluation/criteria) — understanding each score
- [Presets](/docs/qa/evaluation/presets) — choosing the right config
- [How to run evaluations](/docs/how-to/api-cli/run-evals) — CLI reference
