Your agent charged the card twice: making tool calls replay-safe
Why a crash mid-node re-runs tools that already fired, why the usual fixes fall short, and how 10xGraph's tool ledger and versioned writes close the gap.
A support agent receives “please charge my saved card for order 4821 and email me the receipt.” The model replies with three tool calls in a single assistant message: lookup_order, charge_card and send_receipt. The tool node starts all three in parallel. charge_card returns a successful payment. A deploy then sends SIGKILL to the container before the node finishes.
The orchestrator restarts the run from its last checkpoint, which points at the start of that tool node. The node runs again and so do all three tools. The customer is charged twice.
Neither the model nor the payment tool made an error here. The double charge comes from where the checkpoint boundary sits. This post explains why every durable agent runtime has this problem, why the common fixes are incomplete, and what 10xGraph does about it: a tool ledger keyed on the right identity, versioned state writes, and deadlines on the work that can hang.
Why does a resumed run repeat side effects?
10xGraph runs a graph as a loop. Before a node executes, the run’s state records it as the current_node. After the node completes, the loop picks the next node, advances the step counter and persists the state. That design gives you resumability: if the process dies, the next run on the same thread loads the state and continues from current_node.
The catch is that the unit of progress is the node, while side effects happen inside the node. Here is the failure laid out as a timeline:
| Time | Event | Durable state says |
|---|---|---|
| t0 | Model node finishes and emits three tool calls | current_node = tools |
| t1 | Tool node starts lookup_order, charge_card, send_receipt |
current_node = tools |
| t2 | lookup_order returns |
current_node = tools |
| t3 | charge_card returns: payment captured |
current_node = tools |
| t4 | Container is killed | current_node = tools |
| t5 | Run resumes on another worker | current_node = tools |
| t6 | Tool node runs from the start, charge_card fires again |
At t5 the runtime cannot tell “the tool never ran” apart from “the tool ran and the process died right after.” Both look the same in the checkpoint. Without more information the only safe choice is to run the node again, which means tools run at least once. Message queues and workflow engines provide the same guarantee for the same reason. The difference with agents is that the work inside the boundary is chosen by a model at runtime and often touches money, email, tickets or production systems.
The three fixes teams usually try first
1. “Make every tool idempotent”
This is correct advice and you should follow it. A charge endpoint that accepts an idempotency key, an upsert in place of an insert, and a “send if not already sent” check all make replays harmless.
As the only defense it has two problems. It puts the burden on every tool author, and one tool written without it reopens the hole. It also does not cover tools you do not control: a third-party SDK, a shell command, or a tool exposed by someone else’s MCP server.
2. “Checkpoint after every tool”
If the boundary is too coarse, make it finer. Persist the whole state after each tool call and the replay only covers the tool that was in flight.
This costs a full state write per tool call, and parallel tools would write concurrently to the same thread. It also does not fully solve the problem. The crash can still land after the side effect and before the checkpoint, which leaves the same window at a smaller scale. A full state snapshot is heavy machinery for answering a small question: did this specific call already complete?
3. “Keep a ledger keyed on tool_call_id”
This is close to the right answer. Record each completed call under its tool_call_id, and on replay look it up before running the tool.
It fails quietly because a tool_call_id is not unique over the life of a thread. Many models and providers generate short, sequential IDs such as call_1, and they reuse them on every turn. Consider:
- Turn 1: the model calls
charge_cardascall_1. The ledger recordscall_1. - Turn 9: the customer asks for a refund. The model calls
refund_payment, also ascall_1. - The ledger finds
call_1, returns the turn 1 result and skips the refund.
The refund never happens and the model is told it succeeded. Replaying a side effect twice is bad. Skipping one that never ran and reporting success is worse, because nobody finds out until a customer complains.
How 10xGraph keys the ledger
The question to ask is what uniquely identifies one tool invocation. The answer is the assistant message that issued it, combined with the call ID inside that message. That message is persisted with the thread, so on replay it has the same message_id and the key is stable. A different turn produces a different assistant message, so two turns never collide even when the model reuses call_1.
@staticmethod
def _ledger_key(origin_message_id: str | None, tool_call_id: str) -> str | None:
if not tool_call_id or not origin_message_id:
return None
return f"{origin_message_id}:{tool_call_id}"If either part is missing, the key is None and the ledger is skipped for that call. A wrong key could skip a tool that never ran, so in that case the runtime falls back to running the tool.
The lookup and record cycle
For each tool call in a node, the handler does the following:
- Builds the ledger key from the issuing message ID and the call ID.
- Asks the checkpointer for a recorded result under that key.
- On a hit, returns the recorded result and does not call the tool. A log line says the call was replayed from the ledger.
- On a miss, parses the arguments, runs the tool under its timeout, and records the result as soon as the tool returns, before the rest of the node continues.
ledger_key = self._ledger_key(origin_message_id, tool_call_id)
if checkpointer and ledger_key:
recorded = await self._get_recorded_tool_result(checkpointer, config, ledger_key)
if recorded is not None:
return recorded # already ran: replay the result, do not call the tool
timeout = resolve_timeout(config, "tool_timeout", DEFAULT_TOOL_TIMEOUT_SECONDS)
tool_result = await self._invoke_tool_guarded(
function_name, function_args, tool_call_id, state, config, timeout, tool_attrs
)
if checkpointer and ledger_key:
await self._record_tool_result(checkpointer, config, ledger_key, tool_result)Results are stored with a small envelope: {"__kind__": "message", "value": ...} for a tool message, or {"__kind__": "raw", "value": ...} for anything else. On a hit the envelope is turned back into a Message. Only a well-formed entry counts as a hit. A backend that returns an unexpected shape, or a message that no longer validates, is treated as “no record” and the tool runs. Running a tool a second time is recoverable. Silently skipping it because a ledger row was misread is not.
Where the ledger lives
In PgCheckpointer the ledger is a table added by schema migration 3. The migration runs automatically when the checkpointer sets up its schema.
CREATE TABLE IF NOT EXISTS tool_executions (
thread_id VARCHAR(255) NOT NULL
REFERENCES threads(thread_id) ON DELETE CASCADE,
tool_call_id VARCHAR(255) NOT NULL, -- stores "<message_id>:<tool_call_id>"
result JSONB NOT NULL,
created_at TIMESTAMPTZ DEFAULT NOW(),
PRIMARY KEY (thread_id, tool_call_id)
)The thread ID column type follows your configured ID type. Three details matter in production:
- The primary key is the idempotency key. Scoping by
thread_idmeans two threads can never share an entry. - Inserts use
ON CONFLICT DO NOTHING. If two attempts ever race to record the same call, the first recorded result stays authoritative. - Deleting a thread deletes its ledger through the foreign key cascade, so there is no separate cleanup job.
InMemoryCheckpointer implements the same interface with a dictionary, which is useful in tests but only protects replays inside one process. SqliteCheckpointer does not implement the ledger. It inherits the base class default, which always returns “no record”, so tools behave as they did before: at least once.
What happens when the ledger itself fails?
A safety mechanism needs defined behavior for its own failures. The read side and the write side fail differently, on purpose.
| Operation | What happens on failure | Reasoning |
|---|---|---|
Ledger read (aget_tool_result) |
Logged. Treated as “no record”. The tool runs. | A database blip should not stop the run. The fallback is the old at-least-once behavior, not a hard failure. |
Ledger write in PgCheckpointer (aput_tool_result) |
Raises to the node handler. | The checkpointer reports the failure instead of hiding it. |
| Ledger write, seen by the node handler | Logged at error level. The result still goes back to the model and the run continues. | The side effect has already happened. Failing the run would not undo it, and the model needs the result. |
The error log is deliberately specific: the call completed but could not be recorded, and a replay of the node may execute it again. In production, alert on that message. It marks the one case where the runtime knows its idempotency guarantee did not hold for a call.
The window the ledger cannot close
The record is written after the tool returns. If the process dies after the payment provider captured the charge but before the ledger row commits, there is no record, and the replay charges again. The window is small, typically a few milliseconds of network and database time, but it is not zero.
No orchestrator can close it alone. The side effect happens in another system, and making “do the thing” and “remember that we did it” atomic across two systems requires the other system to cooperate. Payment providers and many other APIs support this through idempotency keys.
The best idempotency key comes from your domain, not from the runtime. A charge for an order should happen once per order, whichever turn or tool call asks for it:
async def charge_card(order_id: str, amount_cents: int) -> str:
"""Charge the customer's saved card for an order."""
charge = await payments.charges.create(
amount=amount_cents,
currency="usd",
customer=await customer_for(order_id),
idempotency_key=f"order-{order_id}-charge", # one charge per order, ever
)
return f"Charged {amount_cents / 100:.2f} USD, charge {charge.id}"payments stands in for your provider’s client. The point is the key, which is derived from the business fact being enforced.
Parallel tools, failures and interrupts
Most real tool nodes run several calls at once, and 10xGraph runs them in parallel by default. The ledger is checked per call, which has some useful effects.
Each parallel call gets its own copy of the state. Tools can change injected state in place. If every parallel tool shared one state object, two tools writing different fields could overwrite each other. With more than one call, each branch gets a deep copy, and afterwards only the fields a branch actually changed are merged back against a baseline. A single tool call works on the shared state directly, so the common case has no copying cost.
One failing tool does not abandon its siblings. The handler waits for every call to settle (asyncio.gather with return_exceptions=True) and then turns each failure into a tool error message the model can read and react to. Earlier, a single exception could escape while sibling tasks kept running without being awaited. Now every successful sibling is in the ledger before the node returns.
Malformed arguments are reported to the model. If a model produces invalid JSON for a tool’s arguments, the call returns a tool error saying so, rather than failing the whole node with a decode error.
Human-in-the-loop resumes do not repeat finished work. A tool can call interrupt() to pause the graph for an approval. On resume the interrupted node runs again from the start. In a parallel tool node under invoke, calls that already finished are not run again: their results come from the ledger, and only the paused call continues. Code that runs before the interrupt() call inside the paused tool does run twice, so put side effects after the interrupt:
from tenxgraph.utils import interrupt
async def refund(order_id: str, amount: float) -> str:
decision = interrupt(
{"order_id": order_id, "amount": amount},
message=f"Refund ${amount:.2f} for order {order_id}?",
)
if not decision or not decision.get("approved"):
return "Refund declined"
return await issue_refund(order_id, amount) # side effect after the interruptTwo runs, one thread: versioned writes
Replay safety also depends on what happens when two executions touch the same thread. A client times out and retries while the original run is still going. A resumed run races a new message. Two API workers pick up the same thread. Without protection, the last writer wins and the other run’s work disappears without an error.
PgCheckpointer stores state as versioned rows and checks the version on every durable write:
- Reading a thread’s state records its version in the run config under
_checkpoint_version. - A write opens a transaction, makes sure the thread row exists, and takes a row lock on it with
SELECT ... FOR UPDATE. Concurrent writers on one thread are now serialized. - Under the lock it reads the current maximum version and compares it with the version the run started from.
- If they differ, another execution committed in between, and the write raises
StaleStateErrorinstead of overwriting that work. - Otherwise it inserts a new row at
version + 1and stores the new version back in the run config for the next step.
current_version = await self._lock_thread_for_write(conn, thread_id, user_id, config)
if expected_version is not None and int(expected_version) != current_version:
raise StaleStateError(
message=f"State was modified by another execution "
f"(expected version {expected_version}, found {current_version})",
error_code="STORAGE_CONFLICT_001",
...
)
new_version = current_version + 1A UNIQUE (thread_id, version) index backs this up at the database level, so two appends can never quietly produce the same version. State and the messages produced in that step are written in the same transaction, so a crash cannot leave state advanced without the messages that explain it.
The Redis cache has its own version guard and is cleared when a write conflicts. Without both, a stale cache entry could make a thread fail every later write. The companion post, Hot and cold agent memory, covers this in detail.
The other way a worker dies mid-node: hangs
Crashes are not the only way a node fails to finish. A tool blocked on a half-open socket, or an MCP server that accepts a connection and never answers, will hold a node open forever. The model SDK’s request timeout does not help, because it does not cover tool execution. The step counter never advances, so the recursion limit never trips. Eventually an operator kills the worker, and you are back in the crash-and-replay case.
10xGraph puts a deadline on both levels:
| Config key | Default | Bounds |
|---|---|---|
node_timeout |
900 seconds | A whole node, including every tool it runs |
tool_timeout |
300 seconds | A single tool call |
Both can be overridden per run in the config. None or any non-positive value, such as 0, disables the deadline. An invalid value logs a warning and falls back to the default. The node default deliberately sits above the default model request timeout of 600 seconds, so a slow model call fails with its own error instead of a generic node timeout.
config = {
"thread_id": "order-4821",
"user_id": "customer-118",
"tool_timeout": 60, # a stuck tool fails after one minute
"node_timeout": 300, # no node runs longer than five minutes
}
result = await app.ainvoke({"messages": [message]}, config)A timed-out tool raises NodeTimeoutError (error code NODE_TIMEOUT_002), which the node turns into a tool error message for the model. A timed-out node raises NodeTimeoutError with NODE_TIMEOUT_001.
The same mechanism makes stop requests actually stop work. Stopping used to be checked only between nodes, so a run hung inside a node never saw the request. The node now runs as a cancellable task. While it waits, the runtime checks once per second whether a stop has been requested and cancels the task if so. A stop() call now ends a hung run instead of only appearing to.
Each timeout, stop and error also increments a counter (tenxgraph.node.timeouts, tenxgraph.tool.timeouts, tenxgraph.tool.errors and related metrics), so you can alert on rates rather than reading logs.
How much can a crash replay?
The ledger makes replaying a node safe. Per-step durable checkpoints keep the replay short.
By default, invoke and ainvoke persist state and any new messages to the durable store after every completed node, not only when the run finishes, errors or pauses. Only messages not yet persisted are sent, so a long run does not rewrite its history on each step. A crash therefore replays at most the one node that was in flight, and the ledger makes that replay skip the tools that already finished.
You can turn this off with "durable_checkpoint_every_step": False in the run config. Between terminal points the loop then writes only the Redis cache, which trades durability for fewer database writes. With that setting, a crash after the cache entry expires, or after a Redis restart, resumes from the last terminal checkpoint, which can be the start of the run. Keep the default for any agent with real side effects.
Current limits
These are the limits we know about as of 0.10.0.
- Streaming runs are not covered yet. The ledger lookup, per-step durable checkpoints and node and tool timeouts are implemented in the
invoke/ainvokepath. The streaming path (stream/astream) runs tools without a ledger lookup, writes durably only at terminal points, and has no deadlines. If an agent has non-idempotent tools and runs over streaming, the provider idempotency key is your protection until streaming gets the same treatment. - At least once, not exactly once. The window between a side effect and its ledger record is small but real. See the section above.
- The ledger records outcomes, not intent. A tool that starts a long external job and crashes before returning has no record, so it is started again on replay. For long-running work, have the tool return a job handle quickly and track the job’s progress separately.
- SQLite has no ledger. Use
PgCheckpointeranywhere double execution matters.
Production checklist
- Use
PgCheckpointer(pip install "10xgraph[pg_checkpoint]") and point thecheckpointerkey in10xgraph.jsonat it, for example"checkpointer": "graph.agent:checkpointer". - Use
invoke/ainvokefor runs whose tools have side effects you cannot repeat. - Give every money-moving or message-sending tool an idempotency key derived from your domain.
- Put side effects after
interrupt()calls, never before them. - Set
tool_timeoutto what your slowest healthy tool needs plus some margin, not to the 300 second default. - Alert on the “COMPLETED but could not be recorded” log line and on
tenxgraph.tool.timeouts. - Always pass an explicit
thread_id. Without one the runtime generates a random ID, and the run can be neither resumed nor stopped. - Keep
durable_checkpoint_every_stepat its default oftrue.
from tenxgraph.storage.checkpointer import PgCheckpointer
checkpointer = PgCheckpointer(
postgres_dsn="postgresql://user:password@localhost/app",
redis_url="redis://localhost:6379/0",
)
app = graph.compile(checkpointer=checkpointer)Further reading
- Replay-safe tools is the reference for the ledger and which checkpointers support it.
- Hot and cold agent memory explains the storage underneath: versioned rows, the Redis cache and how they stay consistent.
- Checkpointing and threads covers thread lifecycle and resume.
- Production runtime covers timeouts, stop and metrics.
- Configuration lists the keys in
10xgraph.json.
Frequently asked questions
- Does 10xGraph guarantee a tool runs exactly once?
- No. The ledger records a call after it returns, so a crash in the short window between the side effect and the record can still re-run it. Pass an idempotency key to the downstream API to cover that window.
- Which checkpointers support the tool ledger?
- PgCheckpointer stores it durably in a tool_executions table. InMemoryCheckpointer keeps it in process memory. SqliteCheckpointer does not implement it, so tools there run at least once, as before.
- Does the ledger work for streaming runs?
- Not yet. As of 0.10.0 the ledger, per-step durable checkpoints and node and tool timeouts apply to invoke and ainvoke. Streaming runs persist state at terminal points and run tools without a ledger lookup.
- What happens if recording a tool result fails?
- The tool result is still returned to the model and the run continues. The failure is logged at error level with an explicit warning that a replay of that node may run the tool again.