# Checkpointing

> How to choose, configure, and troubleshoot checkpointers for development and production 10xGraph deployments.

Source: https://10xgraph.com/docs/how-to/production/checkpointing
Last updated: 2026-09-29

Checkpointing is what makes threads durable. Without it, every request is effectively stateless. With it, the API can remember prior messages, state snapshots, and thread metadata across requests.

## Development vs production choice

Use:

- `InMemoryCheckpointer` for local development and tests
- `PgCheckpointer` for production or any multi-instance deployment

## Checkpoint lifecycle

```mermaid
sequenceDiagram
    participant Client
    participant API
    participant Checkpointer

    Client->>API: invoke(thread_id=t1)
    API->>Checkpointer: load thread state
    Checkpointer-->>API: prior state or empty state
    API->>API: run graph
    API->>Checkpointer: save updated state/messages/thread info
    API-->>Client: response
```

## Development setup

For local work:

```python
from tenxgraph.storage.checkpointer import InMemoryCheckpointer

my_checkpointer = InMemoryCheckpointer()
```

Pass it to `compile()` in the module that `10xgraph.json`'s `agent` points at:

```python
# graph/react.py
from graph.dependencies import my_checkpointer

app = state_graph.compile(checkpointer=my_checkpointer)
```

The API server uses the checkpointer the compiled graph carries. The `checkpointer` key in
`10xgraph.json` is recognised but not applied yet, so do not rely on it.

This is perfect when you want:

- quick startup
- no external services
- disposable state

Do not use it in production because all state disappears on process restart and cannot be shared across workers.

## Production setup

For durable, shared state use `PgCheckpointer`.

Example shape:

```python
from tenxgraph.storage.checkpointer import PgCheckpointer

my_checkpointer = PgCheckpointer(
    postgres_dsn="postgresql://user:password@db/agentflow",
    redis_url="redis://redis:6379/0",
)
```

Then compile the graph with it, as above:

```python
app = state_graph.compile(checkpointer=my_checkpointer)
```

Why this is the production choice:

- survives restarts
- supports multiple app instances
- gives shared thread/message/state storage
- separates fast access and durable storage concerns

`PgCheckpointer` keeps a bounded, per-thread history of state snapshots and
prunes older ones automatically. Tune how many are retained with
`state_history_limit` (default `20`; set `1` to keep only the current state) -
see [State history retention](/docs/how-to/python/set-up-checkpointing#state-history-retention-state_history_limit).

## Deployment topology

```mermaid
flowchart LR
    A[Load balancer] --> B[10xGraph API instance 1]
    A --> C[10xGraph API instance 2]
    A --> D[10xGraph API instance 3]
    B --> E[(PostgreSQL)]
    C --> E
    D --> E
    B --> F[(Redis)]
    C --> F
    D --> F
```

If you want multiple API instances, they must share the same durable checkpointer backend.

## Verification steps

After enabling checkpointing, verify actual thread persistence.

### 1. Invoke the graph with a thread ID

```bash
curl -X POST http://127.0.0.1:8000/v1/graph/invoke \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "Remember my name is Alice"}], "config": {"thread_id": "demo-thread"}}'
```

### 2. Send a follow-up request with the same thread ID

```bash
curl -X POST http://127.0.0.1:8000/v1/graph/invoke \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "What is my name?"}], "config": {"thread_id": "demo-thread"}}'
```

### 3. Inspect thread endpoints

```bash
curl http://127.0.0.1:8000/v1/threads
curl http://127.0.0.1:8000/v1/threads/demo-thread/messages
curl http://127.0.0.1:8000/v1/threads/demo-thread/state
```

If checkpointing is working, these endpoints should return persisted data instead of empty results.

## Production recommendations

1. never rely on in-memory persistence for public or shared deployments
2. use one shared durable backend across all workers
3. test restart behavior before release
4. monitor database and Redis connectivity as first-class dependencies
5. back up your database if thread history matters operationally or legally

## Common failure modes

| Symptom | Cause | Fix |
|---|---|---|
| thread history disappears after restart | using `InMemoryCheckpointer` | switch to `PgCheckpointer` |
| one worker sees state and another does not | workers are not sharing the same backend | point all instances at the same Postgres/Redis |
| `/v1/threads` returns errors | database or checkpointer config issue | validate DSN, connectivity, and startup logs |
| request works but no thread data appears later | missing or inconsistent `thread_id` | use a stable `thread_id` per conversation |

## Related docs

- [Checkpointing and Threads](/docs/concepts/checkpointing-and-threads)
- [Configure 10xgraph.json](/docs/how-to/api-cli/configure-agentflow-json)
- [Deployment](/docs/how-to/production/deployment)

## What you learned

- When in-memory checkpointing is enough and when it is not.
- How durable checkpointing supports production thread persistence.
- How to verify checkpointing with real API calls instead of assumptions.
