Checkpointing

In shortHow to choose, configure, and troubleshoot checkpointers for development and production 10xGraph deployments.

  • 3 min read
  • 10 sections
  • Updated
  • v0.9.2
  • Markdown

Checkpointing is what makes threads durable. Without it, every request is effectively stateless. With it, the API can remember prior messages, state snapshots, and thread metadata across requests.

Development vs production choice

Use:

  • InMemoryCheckpointer for local development and tests
  • PgCheckpointer for production or any multi-instance deployment

Checkpoint lifecycle

sequenceDiagram
    participant Client
    participant API
    participant Checkpointer

    Client->>API: invoke(thread_id=t1)
    API->>Checkpointer: load thread state
    Checkpointer-->>API: prior state or empty state
    API->>API: run graph
    API->>Checkpointer: save updated state/messages/thread info
    API-->>Client: response

Development setup

For local work:

Python
from tenxgraph.storage.checkpointer import InMemoryCheckpointer

my_checkpointer = InMemoryCheckpointer()

Pass it to compile() in the module that 10xgraph.json’s agent points at:

Python
# graph/react.py
from graph.dependencies import my_checkpointer

app = state_graph.compile(checkpointer=my_checkpointer)

The API server uses the checkpointer the compiled graph carries. The checkpointer key in 10xgraph.json is recognised but not applied yet, so do not rely on it.

This is perfect when you want:

  • quick startup
  • no external services
  • disposable state

Do not use it in production because all state disappears on process restart and cannot be shared across workers.

Production setup

For durable, shared state use PgCheckpointer.

Example shape:

Python
from tenxgraph.storage.checkpointer import PgCheckpointer

my_checkpointer = PgCheckpointer(
    postgres_dsn="postgresql://user:password@db/agentflow",
    redis_url="redis://redis:6379/0",
)

Then compile the graph with it, as above:

Python
app = state_graph.compile(checkpointer=my_checkpointer)

Why this is the production choice:

  • survives restarts
  • supports multiple app instances
  • gives shared thread/message/state storage
  • separates fast access and durable storage concerns

PgCheckpointer keeps a bounded, per-thread history of state snapshots and prunes older ones automatically. Tune how many are retained with state_history_limit (default 20; set 1 to keep only the current state) - see State history retention.

Deployment topology

flowchart LR
    A[Load balancer] --> B[10xGraph API instance 1]
    A --> C[10xGraph API instance 2]
    A --> D[10xGraph API instance 3]
    B --> E[(PostgreSQL)]
    C --> E
    D --> E
    B --> F[(Redis)]
    C --> F
    D --> F

If you want multiple API instances, they must share the same durable checkpointer backend.

Verification steps

After enabling checkpointing, verify actual thread persistence.

1. Invoke the graph with a thread ID

Terminal
curl -X POST http://127.0.0.1:8000/v1/graph/invoke \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "Remember my name is Alice"}], "config": {"thread_id": "demo-thread"}}'

2. Send a follow-up request with the same thread ID

Terminal
curl -X POST http://127.0.0.1:8000/v1/graph/invoke \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "What is my name?"}], "config": {"thread_id": "demo-thread"}}'

3. Inspect thread endpoints

Terminal
curl http://127.0.0.1:8000/v1/threads
curl http://127.0.0.1:8000/v1/threads/demo-thread/messages
curl http://127.0.0.1:8000/v1/threads/demo-thread/state

If checkpointing is working, these endpoints should return persisted data instead of empty results.

Production recommendations

  1. never rely on in-memory persistence for public or shared deployments
  2. use one shared durable backend across all workers
  3. test restart behavior before release
  4. monitor database and Redis connectivity as first-class dependencies
  5. back up your database if thread history matters operationally or legally

Common failure modes

Symptom Cause Fix
thread history disappears after restart using InMemoryCheckpointer switch to PgCheckpointer
one worker sees state and another does not workers are not sharing the same backend point all instances at the same Postgres/Redis
/v1/threads returns errors database or checkpointer config issue validate DSN, connectivity, and startup logs
request works but no thread data appears later missing or inconsistent thread_id use a stable thread_id per conversation

What you learned

  • When in-memory checkpointing is enough and when it is not.
  • How durable checkpointing supports production thread persistence.
  • How to verify checkpointing with real API calls instead of assumptions.
Last updated for v0.9.2Edit this page on GitHubReport an issue