Errors and limits
In shortException taxonomy, execution limits, retry behavior, and error recovery in graph execution.
- 7 min read
- 9 sections
- Updated
- v0.10.0
- Markdown
The error model
Graph execution raises structured exceptions when nodes fail, storage operations conflict, or limits are exceeded. Rather than crash the system, 10xGraph distinguishes retryable errors (which the runtime can recover from automatically) from fatal errors (which are surfaced immediately for the caller to handle).
This lets you build resilient agents: brief network glitches are invisible, but genuine problems surface quickly with enough context to diagnose and fix them.
Exception taxonomy
Graph-level exceptions (GraphError, NodeError, NodeTimeoutError, GraphRecursionError) live in tenxgraph.core.exceptions. Storage exceptions derive from StorageError, a separate base. Both bases carry an error code, a human-readable message, and a structured context dictionary.
Graph and node errors
GraphError is the base exception for graph-related failures. It includes:
message: Human-readable descriptionerror_code: Unique code (e.g.,GRAPH_000)context: Dictionary with extra details (node name, input size, etc.)
NodeError is raised when a node fails to execute. This includes tool errors, user code exceptions inside a custom node, and LLM call failures that cannot be retried. The exception includes the node name in context so you can see which step failed.
GraphRecursionError occurs when the graph exceeds its recursion limit (default 25 steps). This is a safety mechanism to prevent infinite loops. Each node execution counts as one step, including a ToolNode run.
Timeout errors
NodeTimeoutError (a subclass of NodeError) is raised when a node or tool execution exceeds its timeout. Without a bound, a hanging tool (a half-open socket in an MCP server, a custom tool that never returns, a database query with no read timeout) blocks the entire graph forever. The recursion limit never trips because no step completes, and stop signals are only checked between nodes, never within them.
Timing out the node converts that indefinite hang into a reportable error. Timeouts are on by default; see Timeouts.
Storage and concurrency errors
StorageError is the base for persistence layer failures: database errors, serialization problems, schema version mismatches, and conflicts. It is not a subclass of GraphError.
TransientStorageError is a retryable storage error: a connection drop, timeout, or transient database unavailability. PgCheckpointer retries connection-level failures itself (up to 3 attempts with exponential back-off) and raises this when it still cannot store state.
SerializationError indicates that state or messages could not be serialized to or deserialized from the checkpointer. This is usually fatal (your state structure is corrupted or incompatible), so it does not retry.
SchemaVersionError occurs when the checkpointer detects a schema version mismatch, usually after a code upgrade changes the state structure. May require migration; the exact recovery depends on your checkpointer.
StaleStateError is raised during an optimistic-concurrency check: two graph runs tried to update the same thread simultaneously, and one write succeeded before the other. The losing write is rejected to prevent a lost update. The HTTP API returns 409 Conflict. Reload the latest state and retry. PgCheckpointer raises it.
ResourceNotFoundError indicates a resource (thread, file, stored message) does not exist. Non-retryable.
Other errors
GraphStopRequested is raised when a stop signal (from the API or client) is received while a node is running. This is control flow, not a failure: the execution loop catches it, marks the run stopped, and returns cleanly. It lets the stop check reach the caller even if a node is mid-execution.
MetricsError is the exception type for metrics emission failures (for example, OTEL export unavailable).
UnsupportedMediaInputError is raised before the provider call when a model cannot accept the given media type. For example, trying to send an image to a text-only model, or a document to a model that does not support documents. Includes a suggestion to switch models or change input format.
Limits
Recursion limit
The default recursion limit is 25 steps. Each node execution counts as one step. An Agent node and a ToolNode each count, so one model call plus one round of tool calls is two steps.
This prevents accidental infinite loops. Legitimate multi-turn flows (a ReAct loop with several tool rounds, for example) can exceed it. Increase it per call:
result = graph.invoke(
{"messages": [...]},
config={"thread_id": "t1", "recursion_limit": 50},
)When the limit is hit, the run raises GraphRecursionError.
Timeouts
Both timeouts are on by default:
| Config key | Default | Applies to |
|---|---|---|
node_timeout |
900 seconds | One node execution |
tool_timeout |
300 seconds | One tool call |
Override them per run, or disable one with None or 0. A timed-out node raises NodeTimeoutError:
result = graph.invoke(
{"messages": [...]},
config={"thread_id": "t1", "node_timeout": 120, "tool_timeout": 60},
)Defaults are in tenxgraph/utils/constants.py. Individual tools can also set their own timeouts with libraries that support them (for example, HTTP clients with read timeouts).
State and message size
There is no enforced maximum state size, but each state snapshot is persisted to the checkpointer. Storing very large state (gigabytes of uncompressed text) will stress your database. Trim context regularly with MessageContextManager or SummaryContextManager to keep state size reasonable. See Context management.
Retry and fallback behavior
The Agent class accepts retry_config and fallback_models parameters to control LLM call retry and fallback logic.
RetryConfig
The RetryConfig dataclass controls retry behavior:
from tenxgraph.core.graph import Agent, RetryConfig
retry = RetryConfig(
max_retries=3, # 3 retry attempts for the primary model
initial_delay=1.0, # Start with 1 second delay
max_delay=30.0, # Cap delay at 30 seconds
backoff_factor=2.0, # Double delay after each retry (1s, 2s, 4s, 8s...)
retryable_status_codes=frozenset({429, 500, 502, 503, 529}), # Retryable HTTP codes
circuit_breaker_enabled=False, # Optional: track failures per model
circuit_breaker_threshold=5, # Open circuit after 5 consecutive failures
circuit_breaker_reset_timeout=30.0, # Try again after 30 seconds
)
agent = Agent(
model="openai/gpt-4o",
retry_config=retry
)Retries are on by default with RetryConfig() (3 retries, 1s initial delay, 2x backoff). Pass retry_config=False to disable them.
What gets retried
LLM calls retry on:
- HTTP status codes 429 (rate limit), 500 (server error), 502/503 (gateway/service unavailable), 529 (provider overloaded)
- Connection-level errors (
ConnectionError,TimeoutError,OSError) and exceptions whose class name containstimeout,connectionorunavailable
These are temporary; the same request often succeeds on the next attempt.
What does not retry
Anything outside the retryable set surfaces immediately:
- 401/403 (authentication/authorization failure): Fix credentials or permissions
- 404 (model not found): Use a valid model name
- Serialization/schema errors in storage: The state structure is broken
- User code exceptions in a custom node: Fix the code
- GraphRecursionError: Design the graph to use fewer steps
Fallback models
If retries are exhausted on the primary model, try fallbacks:
agent = Agent(
model="openai/gpt-4o",
fallback_models=["gpt-4o-mini", ("gemini-2.0-flash", "google")],
retry_config=RetryConfig(max_retries=2),
)Fallback occurs after max_retries on the current model. Each entry is a model name or a (model, provider) tuple. Each fallback gets its own retry budget. This lets you degrade gracefully: use your fast/cheap primary model, but fall back to a more capable model if rate-limited or if the primary is unavailable.
Circuit breaker
The circuit breaker is off by default. When enabled, it tracks consecutive failures per (provider, model) pair. After circuit_breaker_threshold failures (default 5), the circuit opens and that model is skipped for circuit_breaker_reset_timeout (default 30s). This avoids repeatedly hammering a known-dead provider:
retry = RetryConfig(
circuit_breaker_enabled=True,
circuit_breaker_threshold=3,
circuit_breaker_reset_timeout=60.0,
)Enable this if you have flaky providers or your quota is exhausted and retries fail fast anyway.
Error recovery in streams
When a run fails while you stream it with astream(), the stream first yields a StreamChunk with event=StreamEvent.ERROR and the failure text in chunk.data["reason"], and then the original exception is raised. Wrap the loop in try/except to handle the exception type:
from tenxgraph.core.exceptions import GraphRecursionError, NodeError, StaleStateError
from tenxgraph.core.state import Message, StreamEvent
try:
async for chunk in app.astream(
{"messages": [Message.text_message("Hello")]},
config={"thread_id": "t1"},
):
if chunk.event == StreamEvent.ERROR:
print("run failed:", chunk.data["reason"])
except StaleStateError:
# Another run updated this thread. Reload the thread and retry.
...
except GraphRecursionError:
# Raise recursion_limit or simplify the graph.
...
except NodeError as e:
print(e.error_code, e.message)
raiseOn failure, the error is recorded in the state and the state is persisted before the exception propagates.
HTTP API error responses
The REST API maps exceptions to status codes. Notable ones:
- 401 Unauthorized and 403 Forbidden: missing or invalid authentication, or insufficient permissions
- 404 Not Found: thread or resource does not exist
- 409 Conflict:
StaleStateError; reload state and retry - 422 Unprocessable Entity: request or validation errors, and
SchemaVersionError - 429 Too Many Requests: rate limited
- 500 Internal Server Error:
GraphError,NodeError,GraphRecursionError,StorageErrorand other server-side failures; messages are sanitized whenMODE=production - 503 Service Unavailable:
TransientStorageError
Most errors use this body shape:
{
"error": {
"code": "NODE_000",
"message": "Node failed",
"details": []
},
"metadata": {}
}The 409 conflict response has its own shape:
{
"error": "state_conflict",
"detail": "This thread was updated by another run while yours was in flight. Reload the thread and retry.",
"thread_id": "abc123"
}Best practices
Design for transience. Assume any network-dependent operation (LLM call, database query, tool invocation) may fail once and succeed the next time. Set retry_config appropriately for your use case.
Handle StaleStateError. If your application has concurrent runs on the same thread, catch StaleStateError (or handle the 409) and reload state before retrying. Or serialize access to threads to avoid the conflict in the first place.
Log and monitor. Errors are logged with structured context; forward logs to your observability stack (Logfire, LangSmith, Datadog) to spot patterns.
Validate input. Catch errors early: validate user input before invoking the graph. Invalid input will fail all retries and delay error reporting.
Test error paths. Unit test your error handling: mock providers to return 429 or 503, raise TransientStorageError from a fake checkpointer, and check that fallback models are tried when the primary fails.
Related pages
- Checkpointing and threads: Thread isolation and durability
- Durability and concurrency: Handling conflicts and concurrent access
- Memory: How PgCheckpointer stores state
- Interrupts: Pausing and resuming execution
Frequently asked questions
- How do I retry a failed node?
- Set retry_config on your Agent. LLM calls that fail with a transient error (429, 500, 502, 503, 529, or a connection or timeout error) retry with exponential back-off; other errors surface immediately.
- What causes a GraphRecursionError?
- The default recursion limit is 25 steps. Each node execution counts as one step. Increase it when invoking, or redesign your graph to use fewer steps per run.
- What should I do about StaleStateError?
- This is a write conflict when two runs update the same thread simultaneously. The API returns HTTP 409. Reload the latest state and retry, or serialize access to the thread.