Checkpointing
In shortHow to choose, configure, and troubleshoot checkpointers for development and production 10xGraph deployments.
- 3 min read
- 10 sections
- Updated
- v0.9.2
- Markdown
Checkpointing is what makes threads durable. Without it, every request is effectively stateless. With it, the API can remember prior messages, state snapshots, and thread metadata across requests.
Development vs production choice
Use:
InMemoryCheckpointerfor local development and testsPgCheckpointerfor production or any multi-instance deployment
Checkpoint lifecycle
sequenceDiagram
participant Client
participant API
participant Checkpointer
Client->>API: invoke(thread_id=t1)
API->>Checkpointer: load thread state
Checkpointer-->>API: prior state or empty state
API->>API: run graph
API->>Checkpointer: save updated state/messages/thread info
API-->>Client: response
Development setup
For local work:
from tenxgraph.storage.checkpointer import InMemoryCheckpointer
my_checkpointer = InMemoryCheckpointer()Pass it to compile() in the module that 10xgraph.json’s agent points at:
# graph/react.py
from graph.dependencies import my_checkpointer
app = state_graph.compile(checkpointer=my_checkpointer)The API server uses the checkpointer the compiled graph carries. The checkpointer key in
10xgraph.json is recognised but not applied yet, so do not rely on it.
This is perfect when you want:
- quick startup
- no external services
- disposable state
Do not use it in production because all state disappears on process restart and cannot be shared across workers.
Production setup
For durable, shared state use PgCheckpointer.
Example shape:
from tenxgraph.storage.checkpointer import PgCheckpointer
my_checkpointer = PgCheckpointer(
postgres_dsn="postgresql://user:password@db/agentflow",
redis_url="redis://redis:6379/0",
)Then compile the graph with it, as above:
app = state_graph.compile(checkpointer=my_checkpointer)Why this is the production choice:
- survives restarts
- supports multiple app instances
- gives shared thread/message/state storage
- separates fast access and durable storage concerns
PgCheckpointer keeps a bounded, per-thread history of state snapshots and
prunes older ones automatically. Tune how many are retained with
state_history_limit (default 20; set 1 to keep only the current state) -
see State history retention.
Deployment topology
flowchart LR
A[Load balancer] --> B[10xGraph API instance 1]
A --> C[10xGraph API instance 2]
A --> D[10xGraph API instance 3]
B --> E[(PostgreSQL)]
C --> E
D --> E
B --> F[(Redis)]
C --> F
D --> F
If you want multiple API instances, they must share the same durable checkpointer backend.
Verification steps
After enabling checkpointing, verify actual thread persistence.
1. Invoke the graph with a thread ID
curl -X POST http://127.0.0.1:8000/v1/graph/invoke \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "Remember my name is Alice"}], "config": {"thread_id": "demo-thread"}}'2. Send a follow-up request with the same thread ID
curl -X POST http://127.0.0.1:8000/v1/graph/invoke \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "What is my name?"}], "config": {"thread_id": "demo-thread"}}'3. Inspect thread endpoints
curl http://127.0.0.1:8000/v1/threads
curl http://127.0.0.1:8000/v1/threads/demo-thread/messages
curl http://127.0.0.1:8000/v1/threads/demo-thread/stateIf checkpointing is working, these endpoints should return persisted data instead of empty results.
Production recommendations
- never rely on in-memory persistence for public or shared deployments
- use one shared durable backend across all workers
- test restart behavior before release
- monitor database and Redis connectivity as first-class dependencies
- back up your database if thread history matters operationally or legally
Common failure modes
| Symptom | Cause | Fix |
|---|---|---|
| thread history disappears after restart | using InMemoryCheckpointer |
switch to PgCheckpointer |
| one worker sees state and another does not | workers are not sharing the same backend | point all instances at the same Postgres/Redis |
/v1/threads returns errors |
database or checkpointer config issue | validate DSN, connectivity, and startup logs |
| request works but no thread data appears later | missing or inconsistent thread_id |
use a stable thread_id per conversation |
Related docs
What you learned
- When in-memory checkpointing is enough and when it is not.
- How durable checkpointing supports production thread persistence.
- How to verify checkpointing with real API calls instead of assumptions.