# Errors and limits

> Exception taxonomy, execution limits, retry behavior, and error recovery in graph execution.

Source: https://10xgraph.com/docs/concepts/errors-and-limits
Last updated: 2026-10-08

## The error model

Graph execution raises structured exceptions when nodes fail, storage operations conflict, or limits are exceeded. Rather than crash the system, 10xGraph distinguishes retryable errors (which the runtime can recover from automatically) from fatal errors (which are surfaced immediately for the caller to handle).

This lets you build resilient agents: brief network glitches are invisible, but genuine problems surface quickly with enough context to diagnose and fix them.

## Exception taxonomy

Graph-level exceptions (`GraphError`, `NodeError`, `NodeTimeoutError`, `GraphRecursionError`) live in `tenxgraph.core.exceptions`. Storage exceptions derive from `StorageError`, a separate base. Both bases carry an error code, a human-readable message, and a structured context dictionary.

### Graph and node errors

**GraphError** is the base exception for graph-related failures. It includes:
- `message`: Human-readable description
- `error_code`: Unique code (e.g., `GRAPH_000`)
- `context`: Dictionary with extra details (node name, input size, etc.)

**NodeError** is raised when a node fails to execute. This includes tool errors, user code exceptions inside a custom node, and LLM call failures that cannot be retried. The exception includes the node name in context so you can see which step failed.

**GraphRecursionError** occurs when the graph exceeds its recursion limit (default 25 steps). This is a safety mechanism to prevent infinite loops. Each node execution counts as one step, including a `ToolNode` run.

### Timeout errors

**NodeTimeoutError** (a subclass of `NodeError`) is raised when a node or tool execution exceeds its timeout. Without a bound, a hanging tool (a half-open socket in an MCP server, a custom tool that never returns, a database query with no read timeout) blocks the entire graph forever. The recursion limit never trips because no step completes, and stop signals are only checked between nodes, never within them.

Timing out the node converts that indefinite hang into a reportable error. Timeouts are on by default; see [Timeouts](#timeouts).

### Storage and concurrency errors

**StorageError** is the base for persistence layer failures: database errors, serialization problems, schema version mismatches, and conflicts. It is not a subclass of `GraphError`.

**TransientStorageError** is a retryable storage error: a connection drop, timeout, or transient database unavailability. `PgCheckpointer` retries connection-level failures itself (up to 3 attempts with exponential back-off) and raises this when it still cannot store state.

**SerializationError** indicates that state or messages could not be serialized to or deserialized from the checkpointer. This is usually fatal (your state structure is corrupted or incompatible), so it does not retry.

**SchemaVersionError** occurs when the checkpointer detects a schema version mismatch, usually after a code upgrade changes the state structure. May require migration; the exact recovery depends on your checkpointer.

**StaleStateError** is raised during an optimistic-concurrency check: two graph runs tried to update the same thread simultaneously, and one write succeeded before the other. The losing write is rejected to prevent a lost update. The HTTP API returns 409 Conflict. Reload the latest state and retry. `PgCheckpointer` raises it.

**ResourceNotFoundError** indicates a resource (thread, file, stored message) does not exist. Non-retryable.

### Other errors

**GraphStopRequested** is raised when a stop signal (from the API or client) is received while a node is running. This is control flow, not a failure: the execution loop catches it, marks the run stopped, and returns cleanly. It lets the stop check reach the caller even if a node is mid-execution.

**MetricsError** is the exception type for metrics emission failures (for example, OTEL export unavailable).

**UnsupportedMediaInputError** is raised before the provider call when a model cannot accept the given media type. For example, trying to send an image to a text-only model, or a document to a model that does not support documents. Includes a suggestion to switch models or change input format.

## Limits

### Recursion limit

The default recursion limit is 25 steps. Each node execution counts as one step. An Agent node and a `ToolNode` each count, so one model call plus one round of tool calls is two steps.

This prevents accidental infinite loops. Legitimate multi-turn flows (a ReAct loop with several tool rounds, for example) can exceed it. Increase it per call:

```python
result = graph.invoke(
    {"messages": [...]},
    config={"thread_id": "t1", "recursion_limit": 50},
)
```

When the limit is hit, the run raises `GraphRecursionError`.

### Timeouts

Both timeouts are on by default:

| Config key | Default | Applies to |
|---|---|---|
| `node_timeout` | 900 seconds | One node execution |
| `tool_timeout` | 300 seconds | One tool call |

Override them per run, or disable one with `None` or `0`. A timed-out node raises `NodeTimeoutError`:

```python
result = graph.invoke(
    {"messages": [...]},
    config={"thread_id": "t1", "node_timeout": 120, "tool_timeout": 60},
)
```

Defaults are in `tenxgraph/utils/constants.py`. Individual tools can also set their own timeouts with libraries that support them (for example, HTTP clients with read timeouts).

### State and message size

There is no enforced maximum state size, but each state snapshot is persisted to the checkpointer. Storing very large state (gigabytes of uncompressed text) will stress your database. Trim context regularly with `MessageContextManager` or `SummaryContextManager` to keep state size reasonable. See [Context management](/docs/concepts/context-management).

## Retry and fallback behavior

The `Agent` class accepts `retry_config` and `fallback_models` parameters to control LLM call retry and fallback logic.

### RetryConfig

The `RetryConfig` dataclass controls retry behavior:

```python
from tenxgraph.core.graph import Agent, RetryConfig

retry = RetryConfig(
    max_retries=3,               # 3 retry attempts for the primary model
    initial_delay=1.0,           # Start with 1 second delay
    max_delay=30.0,              # Cap delay at 30 seconds
    backoff_factor=2.0,          # Double delay after each retry (1s, 2s, 4s, 8s...)
    retryable_status_codes=frozenset({429, 500, 502, 503, 529}),  # Retryable HTTP codes
    circuit_breaker_enabled=False,   # Optional: track failures per model
    circuit_breaker_threshold=5,     # Open circuit after 5 consecutive failures
    circuit_breaker_reset_timeout=30.0,  # Try again after 30 seconds
)

agent = Agent(
    model="openai/gpt-4o",
    retry_config=retry
)
```

Retries are on by default with `RetryConfig()` (3 retries, 1s initial delay, 2x backoff). Pass `retry_config=False` to disable them.

### What gets retried

LLM calls retry on:
- HTTP status codes 429 (rate limit), 500 (server error), 502/503 (gateway/service unavailable), 529 (provider overloaded)
- Connection-level errors (`ConnectionError`, `TimeoutError`, `OSError`) and exceptions whose class name contains `timeout`, `connection` or `unavailable`

These are temporary; the same request often succeeds on the next attempt.

### What does not retry

Anything outside the retryable set surfaces immediately:
- 401/403 (authentication/authorization failure): Fix credentials or permissions
- 404 (model not found): Use a valid model name
- Serialization/schema errors in storage: The state structure is broken
- User code exceptions in a custom node: Fix the code
- GraphRecursionError: Design the graph to use fewer steps

### Fallback models

If retries are exhausted on the primary model, try fallbacks:

```python
agent = Agent(
    model="openai/gpt-4o",
    fallback_models=["gpt-4o-mini", ("gemini-2.0-flash", "google")],
    retry_config=RetryConfig(max_retries=2),
)
```

Fallback occurs after `max_retries` on the current model. Each entry is a model name or a `(model, provider)` tuple. Each fallback gets its own retry budget. This lets you degrade gracefully: use your fast/cheap primary model, but fall back to a more capable model if rate-limited or if the primary is unavailable.

### Circuit breaker

The circuit breaker is off by default. When enabled, it tracks consecutive failures per (provider, model) pair. After `circuit_breaker_threshold` failures (default 5), the circuit opens and that model is skipped for `circuit_breaker_reset_timeout` (default 30s). This avoids repeatedly hammering a known-dead provider:

```python
retry = RetryConfig(
    circuit_breaker_enabled=True,
    circuit_breaker_threshold=3,
    circuit_breaker_reset_timeout=60.0,
)
```

Enable this if you have flaky providers or your quota is exhausted and retries fail fast anyway.

## Error recovery in streams

When a run fails while you stream it with `astream()`, the stream first yields a `StreamChunk` with `event=StreamEvent.ERROR` and the failure text in `chunk.data["reason"]`, and then the original exception is raised. Wrap the loop in `try/except` to handle the exception type:

```python
from tenxgraph.core.exceptions import GraphRecursionError, NodeError, StaleStateError
from tenxgraph.core.state import Message, StreamEvent

try:
    async for chunk in app.astream(
        {"messages": [Message.text_message("Hello")]},
        config={"thread_id": "t1"},
    ):
        if chunk.event == StreamEvent.ERROR:
            print("run failed:", chunk.data["reason"])
except StaleStateError:
    # Another run updated this thread. Reload the thread and retry.
    ...
except GraphRecursionError:
    # Raise recursion_limit or simplify the graph.
    ...
except NodeError as e:
    print(e.error_code, e.message)
    raise
```

On failure, the error is recorded in the state and the state is persisted before the exception propagates.

## HTTP API error responses

The REST API maps exceptions to status codes. Notable ones:

- **401 Unauthorized** and **403 Forbidden**: missing or invalid authentication, or insufficient permissions
- **404 Not Found**: thread or resource does not exist
- **409 Conflict**: `StaleStateError`; reload state and retry
- **422 Unprocessable Entity**: request or validation errors, and `SchemaVersionError`
- **429 Too Many Requests**: rate limited
- **500 Internal Server Error**: `GraphError`, `NodeError`, `GraphRecursionError`, `StorageError` and other server-side failures; messages are sanitized when `MODE=production`
- **503 Service Unavailable**: `TransientStorageError`

Most errors use this body shape:

```json
{
  "error": {
    "code": "NODE_000",
    "message": "Node failed",
    "details": []
  },
  "metadata": {}
}
```

The 409 conflict response has its own shape:

```json
{
  "error": "state_conflict",
  "detail": "This thread was updated by another run while yours was in flight. Reload the thread and retry.",
  "thread_id": "abc123"
}
```

## Best practices

**Design for transience.** Assume any network-dependent operation (LLM call, database query, tool invocation) may fail once and succeed the next time. Set `retry_config` appropriately for your use case.

**Handle StaleStateError.** If your application has concurrent runs on the same thread, catch `StaleStateError` (or handle the 409) and reload state before retrying. Or serialize access to threads to avoid the conflict in the first place.

**Log and monitor.** Errors are logged with structured context; forward logs to your observability stack (Logfire, LangSmith, Datadog) to spot patterns.

**Validate input.** Catch errors early: validate user input before invoking the graph. Invalid input will fail all retries and delay error reporting.

**Test error paths.** Unit test your error handling: mock providers to return 429 or 503, raise `TransientStorageError` from a fake checkpointer, and check that fallback models are tried when the primary fails.

## Related pages

- [Checkpointing and threads](/docs/concepts/checkpointing-and-threads): Thread isolation and durability
- [Durability and concurrency](/docs/guides/durability-and-concurrency): Handling conflicts and concurrent access
- [Memory](/docs/concepts/memory): How PgCheckpointer stores state
- [Interrupts](/docs/concepts/interrupts): Pausing and resuming execution

## Frequently asked questions

### How do I retry a failed node?

Set retry_config on your Agent. LLM calls that fail with a transient error (429, 500, 502, 503, 529, or a connection or timeout error) retry with exponential back-off; other errors surface immediately.

### What causes a GraphRecursionError?

The default recursion limit is 25 steps. Each node execution counts as one step. Increase it when invoking, or redesign your graph to use fewer steps per run.

### What should I do about StaleStateError?

This is a write conflict when two runs update the same thread simultaneously. The API returns HTTP 409. Reload the latest state and retry, or serialize access to the thread.
