Your agent charged the card twice: making tool calls replay-safe

Why a crash mid-node re-runs tools that already fired, why the usual fixes fall short, and how 10xGraph's tool ledger and versioned writes close the gap.

A support agent receives “please charge my saved card for order 4821 and email me the receipt.” The model replies with three tool calls in a single assistant message: lookup_order, charge_card and send_receipt. The tool node starts all three in parallel. charge_card returns a successful payment. A deploy then sends SIGKILL to the container before the node finishes.

The orchestrator restarts the run from its last checkpoint, which points at the start of that tool node. The node runs again and so do all three tools. The customer is charged twice.

Neither the model nor the payment tool made an error here. The double charge comes from where the checkpoint boundary sits. This post explains why every durable agent runtime has this problem, why the common fixes are incomplete, and what 10xGraph does about it: a tool ledger keyed on the right identity, versioned state writes, and deadlines on the work that can hang.

Why does a resumed run repeat side effects?

10xGraph runs a graph as a loop. Before a node executes, the run’s state records it as the current_node. After the node completes, the loop picks the next node, advances the step counter and persists the state. That design gives you resumability: if the process dies, the next run on the same thread loads the state and continues from current_node.

The catch is that the unit of progress is the node, while side effects happen inside the node. Here is the failure laid out as a timeline:

Time Event Durable state says
t0 Model node finishes and emits three tool calls current_node = tools
t1 Tool node starts lookup_order, charge_card, send_receipt current_node = tools
t2 lookup_order returns current_node = tools
t3 charge_card returns: payment captured current_node = tools
t4 Container is killed current_node = tools
t5 Run resumes on another worker current_node = tools
t6 Tool node runs from the start, charge_card fires again

At t5 the runtime cannot tell “the tool never ran” apart from “the tool ran and the process died right after.” Both look the same in the checkpoint. Without more information the only safe choice is to run the node again, which means tools run at least once. Message queues and workflow engines provide the same guarantee for the same reason. The difference with agents is that the work inside the boundary is chosen by a model at runtime and often touches money, email, tickets or production systems.

The three fixes teams usually try first

1. “Make every tool idempotent”

This is correct advice and you should follow it. A charge endpoint that accepts an idempotency key, an upsert in place of an insert, and a “send if not already sent” check all make replays harmless.

As the only defense it has two problems. It puts the burden on every tool author, and one tool written without it reopens the hole. It also does not cover tools you do not control: a third-party SDK, a shell command, or a tool exposed by someone else’s MCP server.

2. “Checkpoint after every tool”

If the boundary is too coarse, make it finer. Persist the whole state after each tool call and the replay only covers the tool that was in flight.

This costs a full state write per tool call, and parallel tools would write concurrently to the same thread. It also does not fully solve the problem. The crash can still land after the side effect and before the checkpoint, which leaves the same window at a smaller scale. A full state snapshot is heavy machinery for answering a small question: did this specific call already complete?

3. “Keep a ledger keyed on tool_call_id”

This is close to the right answer. Record each completed call under its tool_call_id, and on replay look it up before running the tool.

It fails quietly because a tool_call_id is not unique over the life of a thread. Many models and providers generate short, sequential IDs such as call_1, and they reuse them on every turn. Consider:

  • Turn 1: the model calls charge_card as call_1. The ledger records call_1.
  • Turn 9: the customer asks for a refund. The model calls refund_payment, also as call_1.
  • The ledger finds call_1, returns the turn 1 result and skips the refund.

The refund never happens and the model is told it succeeded. Replaying a side effect twice is bad. Skipping one that never ran and reporting success is worse, because nobody finds out until a customer complains.

How 10xGraph keys the ledger

The question to ask is what uniquely identifies one tool invocation. The answer is the assistant message that issued it, combined with the call ID inside that message. That message is persisted with the thread, so on replay it has the same message_id and the key is stable. A different turn produces a different assistant message, so two turns never collide even when the model reuses call_1.

tenxgraph/core/graph/utils/invoke_node_handler.py
@staticmethod
def _ledger_key(origin_message_id: str | None, tool_call_id: str) -> str | None:
    if not tool_call_id or not origin_message_id:
        return None
    return f"{origin_message_id}:{tool_call_id}"

If either part is missing, the key is None and the ledger is skipped for that call. A wrong key could skip a tool that never ran, so in that case the runtime falls back to running the tool.

The lookup and record cycle

For each tool call in a node, the handler does the following:

  1. Builds the ledger key from the issuing message ID and the call ID.
  2. Asks the checkpointer for a recorded result under that key.
  3. On a hit, returns the recorded result and does not call the tool. A log line says the call was replayed from the ledger.
  4. On a miss, parses the arguments, runs the tool under its timeout, and records the result as soon as the tool returns, before the rest of the node continues.
invoke_node_handler.py (abridged)
ledger_key = self._ledger_key(origin_message_id, tool_call_id)
if checkpointer and ledger_key:
    recorded = await self._get_recorded_tool_result(checkpointer, config, ledger_key)
    if recorded is not None:
        return recorded  # already ran: replay the result, do not call the tool

timeout = resolve_timeout(config, "tool_timeout", DEFAULT_TOOL_TIMEOUT_SECONDS)
tool_result = await self._invoke_tool_guarded(
    function_name, function_args, tool_call_id, state, config, timeout, tool_attrs
)

if checkpointer and ledger_key:
    await self._record_tool_result(checkpointer, config, ledger_key, tool_result)

Results are stored with a small envelope: {"__kind__": "message", "value": ...} for a tool message, or {"__kind__": "raw", "value": ...} for anything else. On a hit the envelope is turned back into a Message. Only a well-formed entry counts as a hit. A backend that returns an unexpected shape, or a message that no longer validates, is treated as “no record” and the tool runs. Running a tool a second time is recoverable. Silently skipping it because a ledger row was misread is not.

Where the ledger lives

In PgCheckpointer the ledger is a table added by schema migration 3. The migration runs automatically when the checkpointer sets up its schema.

pg_checkpointer.py, schema v3 (abridged)
CREATE TABLE IF NOT EXISTS tool_executions (
    thread_id   VARCHAR(255) NOT NULL
        REFERENCES threads(thread_id) ON DELETE CASCADE,
    tool_call_id VARCHAR(255) NOT NULL,   -- stores "<message_id>:<tool_call_id>"
    result      JSONB NOT NULL,
    created_at  TIMESTAMPTZ DEFAULT NOW(),
    PRIMARY KEY (thread_id, tool_call_id)
)

The thread ID column type follows your configured ID type. Three details matter in production:

  • The primary key is the idempotency key. Scoping by thread_id means two threads can never share an entry.
  • Inserts use ON CONFLICT DO NOTHING. If two attempts ever race to record the same call, the first recorded result stays authoritative.
  • Deleting a thread deletes its ledger through the foreign key cascade, so there is no separate cleanup job.

InMemoryCheckpointer implements the same interface with a dictionary, which is useful in tests but only protects replays inside one process. SqliteCheckpointer does not implement the ledger. It inherits the base class default, which always returns “no record”, so tools behave as they did before: at least once.

What happens when the ledger itself fails?

A safety mechanism needs defined behavior for its own failures. The read side and the write side fail differently, on purpose.

Operation What happens on failure Reasoning
Ledger read (aget_tool_result) Logged. Treated as “no record”. The tool runs. A database blip should not stop the run. The fallback is the old at-least-once behavior, not a hard failure.
Ledger write in PgCheckpointer (aput_tool_result) Raises to the node handler. The checkpointer reports the failure instead of hiding it.
Ledger write, seen by the node handler Logged at error level. The result still goes back to the model and the run continues. The side effect has already happened. Failing the run would not undo it, and the model needs the result.

The error log is deliberately specific: the call completed but could not be recorded, and a replay of the node may execute it again. In production, alert on that message. It marks the one case where the runtime knows its idempotency guarantee did not hold for a call.

The window the ledger cannot close

The record is written after the tool returns. If the process dies after the payment provider captured the charge but before the ledger row commits, there is no record, and the replay charges again. The window is small, typically a few milliseconds of network and database time, but it is not zero.

No orchestrator can close it alone. The side effect happens in another system, and making “do the thing” and “remember that we did it” atomic across two systems requires the other system to cooperate. Payment providers and many other APIs support this through idempotency keys.

The best idempotency key comes from your domain, not from the runtime. A charge for an order should happen once per order, whichever turn or tool call asks for it:

graph/tools.py
async def charge_card(order_id: str, amount_cents: int) -> str:
    """Charge the customer's saved card for an order."""
    charge = await payments.charges.create(
        amount=amount_cents,
        currency="usd",
        customer=await customer_for(order_id),
        idempotency_key=f"order-{order_id}-charge",  # one charge per order, ever
    )
    return f"Charged {amount_cents / 100:.2f} USD, charge {charge.id}"

payments stands in for your provider’s client. The point is the key, which is derived from the business fact being enforced.

Parallel tools, failures and interrupts

Most real tool nodes run several calls at once, and 10xGraph runs them in parallel by default. The ledger is checked per call, which has some useful effects.

Each parallel call gets its own copy of the state. Tools can change injected state in place. If every parallel tool shared one state object, two tools writing different fields could overwrite each other. With more than one call, each branch gets a deep copy, and afterwards only the fields a branch actually changed are merged back against a baseline. A single tool call works on the shared state directly, so the common case has no copying cost.

One failing tool does not abandon its siblings. The handler waits for every call to settle (asyncio.gather with return_exceptions=True) and then turns each failure into a tool error message the model can read and react to. Earlier, a single exception could escape while sibling tasks kept running without being awaited. Now every successful sibling is in the ledger before the node returns.

Malformed arguments are reported to the model. If a model produces invalid JSON for a tool’s arguments, the call returns a tool error saying so, rather than failing the whole node with a decode error.

Human-in-the-loop resumes do not repeat finished work. A tool can call interrupt() to pause the graph for an approval. On resume the interrupted node runs again from the start. In a parallel tool node under invoke, calls that already finished are not run again: their results come from the ledger, and only the paused call continues. Code that runs before the interrupt() call inside the paused tool does run twice, so put side effects after the interrupt:

graph/tools.py
from tenxgraph.utils import interrupt


async def refund(order_id: str, amount: float) -> str:
    decision = interrupt(
        {"order_id": order_id, "amount": amount},
        message=f"Refund ${amount:.2f} for order {order_id}?",
    )
    if not decision or not decision.get("approved"):
        return "Refund declined"
    return await issue_refund(order_id, amount)  # side effect after the interrupt

Two runs, one thread: versioned writes

Replay safety also depends on what happens when two executions touch the same thread. A client times out and retries while the original run is still going. A resumed run races a new message. Two API workers pick up the same thread. Without protection, the last writer wins and the other run’s work disappears without an error.

PgCheckpointer stores state as versioned rows and checks the version on every durable write:

  1. Reading a thread’s state records its version in the run config under _checkpoint_version.
  2. A write opens a transaction, makes sure the thread row exists, and takes a row lock on it with SELECT ... FOR UPDATE. Concurrent writers on one thread are now serialized.
  3. Under the lock it reads the current maximum version and compares it with the version the run started from.
  4. If they differ, another execution committed in between, and the write raises StaleStateError instead of overwriting that work.
  5. Otherwise it inserts a new row at version + 1 and stores the new version back in the run config for the next step.
pg_checkpointer.py (abridged)
current_version = await self._lock_thread_for_write(conn, thread_id, user_id, config)

if expected_version is not None and int(expected_version) != current_version:
    raise StaleStateError(
        message=f"State was modified by another execution "
                f"(expected version {expected_version}, found {current_version})",
        error_code="STORAGE_CONFLICT_001",
        ...
    )

new_version = current_version + 1

A UNIQUE (thread_id, version) index backs this up at the database level, so two appends can never quietly produce the same version. State and the messages produced in that step are written in the same transaction, so a crash cannot leave state advanced without the messages that explain it.

The Redis cache has its own version guard and is cleared when a write conflicts. Without both, a stale cache entry could make a thread fail every later write. The companion post, Hot and cold agent memory, covers this in detail.

The other way a worker dies mid-node: hangs

Crashes are not the only way a node fails to finish. A tool blocked on a half-open socket, or an MCP server that accepts a connection and never answers, will hold a node open forever. The model SDK’s request timeout does not help, because it does not cover tool execution. The step counter never advances, so the recursion limit never trips. Eventually an operator kills the worker, and you are back in the crash-and-replay case.

10xGraph puts a deadline on both levels:

Config key Default Bounds
node_timeout 900 seconds A whole node, including every tool it runs
tool_timeout 300 seconds A single tool call

Both can be overridden per run in the config. None or any non-positive value, such as 0, disables the deadline. An invalid value logs a warning and falls back to the default. The node default deliberately sits above the default model request timeout of 600 seconds, so a slow model call fails with its own error instead of a generic node timeout.

run_agent.py
config = {
    "thread_id": "order-4821",
    "user_id": "customer-118",
    "tool_timeout": 60,    # a stuck tool fails after one minute
    "node_timeout": 300,   # no node runs longer than five minutes
}
result = await app.ainvoke({"messages": [message]}, config)

A timed-out tool raises NodeTimeoutError (error code NODE_TIMEOUT_002), which the node turns into a tool error message for the model. A timed-out node raises NodeTimeoutError with NODE_TIMEOUT_001.

The same mechanism makes stop requests actually stop work. Stopping used to be checked only between nodes, so a run hung inside a node never saw the request. The node now runs as a cancellable task. While it waits, the runtime checks once per second whether a stop has been requested and cancels the task if so. A stop() call now ends a hung run instead of only appearing to.

Each timeout, stop and error also increments a counter (tenxgraph.node.timeouts, tenxgraph.tool.timeouts, tenxgraph.tool.errors and related metrics), so you can alert on rates rather than reading logs.

How much can a crash replay?

The ledger makes replaying a node safe. Per-step durable checkpoints keep the replay short.

By default, invoke and ainvoke persist state and any new messages to the durable store after every completed node, not only when the run finishes, errors or pauses. Only messages not yet persisted are sent, so a long run does not rewrite its history on each step. A crash therefore replays at most the one node that was in flight, and the ledger makes that replay skip the tools that already finished.

You can turn this off with "durable_checkpoint_every_step": False in the run config. Between terminal points the loop then writes only the Redis cache, which trades durability for fewer database writes. With that setting, a crash after the cache entry expires, or after a Redis restart, resumes from the last terminal checkpoint, which can be the start of the run. Keep the default for any agent with real side effects.

Current limits

These are the limits we know about as of 0.10.0.

  • Streaming runs are not covered yet. The ledger lookup, per-step durable checkpoints and node and tool timeouts are implemented in the invoke / ainvoke path. The streaming path (stream / astream) runs tools without a ledger lookup, writes durably only at terminal points, and has no deadlines. If an agent has non-idempotent tools and runs over streaming, the provider idempotency key is your protection until streaming gets the same treatment.
  • At least once, not exactly once. The window between a side effect and its ledger record is small but real. See the section above.
  • The ledger records outcomes, not intent. A tool that starts a long external job and crashes before returning has no record, so it is started again on replay. For long-running work, have the tool return a job handle quickly and track the job’s progress separately.
  • SQLite has no ledger. Use PgCheckpointer anywhere double execution matters.

Production checklist

  • Use PgCheckpointer (pip install "10xgraph[pg_checkpoint]") and point the checkpointer key in 10xgraph.json at it, for example "checkpointer": "graph.agent:checkpointer".
  • Use invoke / ainvoke for runs whose tools have side effects you cannot repeat.
  • Give every money-moving or message-sending tool an idempotency key derived from your domain.
  • Put side effects after interrupt() calls, never before them.
  • Set tool_timeout to what your slowest healthy tool needs plus some margin, not to the 300 second default.
  • Alert on the “COMPLETED but could not be recorded” log line and on tenxgraph.tool.timeouts.
  • Always pass an explicit thread_id. Without one the runtime generates a random ID, and the run can be neither resumed nor stopped.
  • Keep durable_checkpoint_every_step at its default of true.
graph/agent.py
from tenxgraph.storage.checkpointer import PgCheckpointer

checkpointer = PgCheckpointer(
    postgres_dsn="postgresql://user:password@localhost/app",
    redis_url="redis://localhost:6379/0",
)
app = graph.compile(checkpointer=checkpointer)

Further reading

Frequently asked questions

Does 10xGraph guarantee a tool runs exactly once?
No. The ledger records a call after it returns, so a crash in the short window between the side effect and the record can still re-run it. Pass an idempotency key to the downstream API to cover that window.
Which checkpointers support the tool ledger?
PgCheckpointer stores it durably in a tool_executions table. InMemoryCheckpointer keeps it in process memory. SqliteCheckpointer does not implement it, so tools there run at least once, as before.
Does the ledger work for streaming runs?
Not yet. As of 0.10.0 the ledger, per-step durable checkpoints and node and tool timeouts apply to invoke and ainvoke. Streaming runs persist state at terminal points and run tools without a ledger lookup.
What happens if recording a tool result fails?
The tool result is still returned to the model and the run continues. The failure is logged at error level with an explicit warning that a replay of that node may run the tool again.

Written by

Shudipto Trafder

Maintainer, 10xGraph

Keep reading

Subscribe via RSS

3 min read

Agentflow is now 10xGraph

Agentflow is now 10xGraph. Why we renamed it, what the framework is, what stays the same in your code, what changes in package names, and the honest gaps.

  • announcement
  • rename

Blog home Archive