# Hot and cold agent memory: why 10xGraph puts Redis in front of Postgres

> How PgCheckpointer splits agent state between a Redis hot cache and versioned Postgres history, keeps them consistent, and behaves when either one fails.

Source: https://10xgraph.com/blog/hot-and-cold-agent-memory
Last updated: 2026-10-06

Every agent framework has to decide where conversation state lives between steps. The decision looks like plumbing, but it sets your per-step latency, what a crash loses, how many workers can serve one thread, and what happens on the day Redis restarts.

10xGraph's production checkpointer, `PgCheckpointer`, uses two stores with separate jobs. **Redis is the hot layer**: the latest state of each active thread, read first and refreshed on every step, with a TTL. **Postgres is the cold layer**: versioned state snapshots, the full message history, thread metadata and the tool-call ledger. It is the source of truth.

Running two stores is easy. Keeping them consistent under crashes, retries and concurrent runs takes more care, and most of this post is about that.

## What an agent run asks of storage

The access pattern of a single run looks like this:

1. **Load the thread** at the start: the current state, including the conversation context and where execution stopped.
2. **Execute a node**: a model call, a tool node, or your own function.
3. **Save the thread** after the node: new state, plus any messages the node produced.
4. Repeat steps 2 and 3 until the graph ends, pauses for a human, or fails.

A twenty-step run reads once and writes twenty times. A chat product with many active threads does this across thousands of threads at once, and the same thread is often picked up by a different API worker on the next turn. The storage layer needs fast reads of the latest state from any worker, cheap writes on every step, durability for anything a user or auditor might need later, and correct behavior when two executions touch one thread.

No single common store does all of that well.

## Why not Postgres alone?

Postgres can store everything, and in 10xGraph it does. The problem is putting it on the read path for every load. Every run that starts or resumes, and every "what is the state of this thread?" request from a UI, becomes a query that finds the newest snapshot row and deserializes a JSON payload that grows with the conversation.

Each query is fast on its own. Across thousands of active threads and several API workers, they add up to steady load on the primary database and latency on the critical path of every turn. The latest state of an active thread is also the most frequently read and most cacheable data in the system. Serving all of it from the system of record is wasteful.

## Why not Redis alone?

Redis is the right home for state you read constantly and can rebuild. It is the wrong home for the only copy.

- **TTL expiry.** Thread state needs a TTL, or memory grows without bound as conversations accumulate. A thread idle longer than the TTL disappears.
- **Eviction.** Under memory pressure, Redis evicts keys according to its `maxmemory-policy`. An active customer's thread can be evicted to make room for another.
- **Restarts.** Depending on persistence settings, a restart can lose recent writes or the whole dataset.
- **No history.** Redis holds the current value. Audits, debugging and "what did the agent see at step 7?" need history.

An agent that forgets a customer's open order after a quiet weekend has a bug, not a cache miss.

## The split

Each store gets the job it is good at:

| Data | Redis (hot) | Postgres (cold) |
|---|---|---|
| Latest thread state | JSON value with a TTL, key `state_cache:<thread_id>:<user_id>` | Versioned rows in `states` |
| State history | None | The last `state_history_limit` snapshots per thread |
| Messages | Only inside the cached state payload | One row per message in `messages` |
| Thread metadata and owner | None | `threads` table |
| Tool-call ledger | None | `tool_executions` table |
| Version | Embedded in the cached payload | `version` column, unique per thread |

Two rules govern the whole design:

1. **Postgres is the source of truth.** Anything in Redis can be thrown away and rebuilt from Postgres.
2. **The cache must never get ahead of the truth.** It must never hold state that was not persisted, and never move backwards in a way that misleads the next writer.

The write path, the read path and the version guards all follow from these two rules.

## The write path

After each completed step, the run loop saves the thread. The order is the important part:

```python title="tenxgraph/core/graph/utils/utils.py (abridged)"
if checkpointer:
    # Durable store is the source of truth: state + messages, atomically, first.
    await checkpointer.aput_checkpoint(config, new_state, messages)
    # Then refresh the realtime cache with the same state we just persisted.
    await checkpointer.aput_state_cache(config, new_state)
```

**Durable first, cache second.** If the Postgres write fails, the exception surfaces before the cache is touched, so the cache cannot hold state that was never persisted. The cache is then written with the same state object that was saved durably, so the two copies cannot differ in shape. This matters when a context manager trims the conversation before saving.

**State and messages in one transaction.** `aput_checkpoint` writes the new state row and that step's messages in a single Postgres transaction. A crash cannot leave the state advanced without the messages that justify it, or the reverse.

**Only new messages are sent.** The run loop tracks how many messages have already been persisted and sends only the rest. A long run does not rewrite its history on every step. Message inserts are upserts keyed by `message_id`, restricted to the same thread. A message ID that already belongs to another thread is rejected and rolls the transaction back, so one thread cannot overwrite another thread's messages.

**State is an append-only log with a retention window.** Each write inserts a new row at the next version number. In the same transaction, rows older than the newest `state_history_limit` (default 20) are deleted. You get recent snapshots for debugging and rollback without unbounded growth. Messages are not affected by this limit.

**Connection errors are retried.** A dropped connection or interface error retries the whole transaction on a fresh pooled connection, up to three attempts with exponential backoff (one second, then two). Other errors, and connection errors that outlast the retries, are raised as a `StorageError`, so the run fails visibly instead of continuing on unsaved state.

You can disable the per-step durable write with `"durable_checkpoint_every_step": False` in the run config. Steps then write only the cache, and Postgres is written when the run completes, errors, pauses or stops. That saves database writes but brings back the risk described above: a crash after the cache expires resumes from the last terminal checkpoint. The default is per-step durability, and you should keep it for any agent whose work matters.

## The read path

At the start of a run, the loop calls `aget_state_cache`:

1. **Try Redis.** On a hit, deserialize the state and restore its embedded version into the run config under `_checkpoint_version`, so the next durable write can still perform its version check.
2. **On a miss, read Postgres.** Select the highest-version snapshot for the thread, record its version in the run config, and seed the cache so the next read is a hit.
3. **If Redis raises**, log the error and read Postgres. A Redis outage makes loads slower and does not make them fail.

If the cache returns nothing, the loop also tries the durable store directly before deciding the thread is new.

The rule throughout is that a cache problem must never become a correctness problem. Cache writes are best-effort too. A failed `aput_state_cache` is logged and returns `None`, and the run continues.

> **Failure handling is deliberately uneven**
>
> Cache failures are swallowed and durable failures are raised. A lost cache entry costs one slower read. A lost durable write loses work, so it must surface.

## The hard part: keeping the cache honest under concurrency

Two executions can touch the same thread: a client retry racing the original run, or a resumed run racing a new message. Postgres handles this with optimistic concurrency. Each run remembers the version it started from, and a write fails with `StaleStateError` if the thread moved on in the meantime. The companion post, [Your agent charged the card twice](/blog/your-agent-charged-the-card-twice), covers that mechanism.

A cache in front of a versioned store adds a less obvious way to fail. Here it is step by step:

1. Run A and run B both load thread T at version 5.
2. Run B finishes first. Postgres moves to version 6 and the cache holds version 6.
3. Run A is still working. After its next step it writes the cache with its own state, still based on version 5. A plain `SETEX` accepts it, and the cache now holds stale state stamped version 5.
4. Run A's durable write fails its version check, correctly.
5. The next run on T reads the cache, which is preferred, picks up version 5 as its expected version, and fails its version check too.
6. So does every run after that, until the TTL expires. With the default TTL, the thread is stuck for up to a day.

Nothing crashed and every component did what it was told. The thread is still unusable. 10xGraph prevents this with two guards.

### Guard 1: the cache never moves backwards

Cache writes go through a small Lua script that Redis runs atomically. It reads the version embedded in the current entry and only writes if the incoming version is at least as new:

```lua title="pg_checkpointer.py: _CACHE_CAS_LUA"
local existing = redis.call('GET', KEYS[1])
if existing then
    local ok, decoded = pcall(cjson.decode, existing)
    if ok and type(decoded) == 'table' and decoded[ARGV[4]] then
        if tonumber(decoded[ARGV[4]]) > tonumber(ARGV[2]) then
            return 0
        end
    end
end
redis.call('SETEX', KEYS[1], ARGV[3], ARGV[1])
return 1
```

Running the compare and the set inside Redis matters. A `GET` followed by a `SETEX` from Python would still race between coroutines and between workers. Equal versions are allowed on purpose: within a single run, state changes on every step (progress, stop flags) while the durable version stays the same until the next durable write. A brand-new thread with no durable version yet writes with a sentinel of `-1`, which is lower than any real version and so can never overwrite a newer entry.

In the scenario above, step 3 now returns `0`. Run A's stale write is skipped and the cache keeps version 6.

### Guard 2: a conflict clears the cache entry

If a durable write raises `StaleStateError`, the checkpointer deletes that thread's cache entry before re-raising. The cache might hold state built on the version that just lost. Deleting it forces the next read back to Postgres, which returns the true latest version, and the thread recovers on its next run instead of staying stuck until the TTL expires.

Either guard covers most cases on its own. Together they make sure a stale entry cannot drive later writes.

## Isolation is part of the key

With authentication enabled, a thread belongs to a user, and knowing a thread ID must not be enough to read it. `PgCheckpointer` treats `user_id` as an ownership boundary by default (`enforce_user_isolation=True`):

- The cache key includes both IDs, `state_cache:<thread_id>:<user_id>`, so one user's cache lookups can never return another user's entry.
- Durable reads, writes and deletes are scoped to the owning user. On write, the thread row is locked and its owner checked, and a write to another user's thread raises a `STORAGE_FORBIDDEN_001` error.

For single-tenant deployments, or setups with no real user identity, set `enforce_user_isolation=False` and the thread ID alone is the key. Runs without a `user_id` share one `"anonymous"` identity, which matches what the API server uses when auth is off.

## Failure matrix

What each failure costs, with the default settings:

| Event | What the agent sees | Data lost |
|---|---|---|
| Redis unreachable | Loads read Postgres. Cache writes are logged and skipped. | None |
| Redis restarts empty | First load of each thread misses, reads Postgres and re-seeds the cache | None |
| Cache TTL expires on an idle thread | Same as a restart, for that thread | None |
| Key evicted under memory pressure | Same as a TTL expiry | None |
| Postgres connection drops briefly | Transaction retried on a fresh connection, up to three attempts | None |
| Postgres unavailable | Durable write raises and the run fails visibly | The in-flight step, which is replayed on resume |
| Worker killed mid-node | Next run resumes at that node. Completed tools come from the ledger. | At most the in-flight node |
| Two runs race on a thread | The loser gets `StaleStateError` and its cache entry is cleared | None. The losing write is rejected rather than merged. |

The only row that loses work is the one where the source of truth is down, which is the trade-off you want.

## Tuning it

```python title="graph/agent.py"
from tenxgraph.storage.checkpointer import PgCheckpointer

checkpointer = PgCheckpointer(
    postgres_dsn="postgresql://user:password@localhost/app",
    redis_url="redis://localhost:6379/0",
    cache_ttl=3600,          # default 86400 (24 hours)
    state_history_limit=20,  # default 20
)
app = graph.compile(checkpointer=checkpointer)
```

Point the `checkpointer` key in `10xgraph.json` at this object, for example `"checkpointer": "graph.agent:checkpointer"`, and the API server builds and uses it.

**`cache_ttl`** should cover the typical gap between turns in an active conversation. For live chat, an hour or a few hours keeps working sets small. For agents users come back to over days, keep the 24-hour default or raise it. A short TTL never loses data. It only means more reads go to Postgres.

**`state_history_limit`** sets how far back you can inspect or roll back a thread's state. Twenty snapshots fits most debugging. Raise it for agents you audit step by step, and lower it if state payloads are large.

**Redis memory policy.** Every thread key has a TTL, so a `volatile-*` eviction policy such as `volatile-lru` evicts thread state under pressure and leaves keys without a TTL alone. An eviction is only a cache miss, so this is safe. The `noeviction` policy is the one to avoid. Under memory pressure it makes cache writes fail. 10xGraph tolerates that, but every read then goes to Postgres.

**Pool sizes.** Both pools can be configured (`pool_config` for asyncpg, `redis_pool_config` for Redis), or you can pass in pools your application already owns. The checkpointer only closes resources it created, so a shared pool is never closed from under your app.

## Local development: the same layers in one file

You do not need Redis and Postgres running to build an agent. `SqliteCheckpointer` keeps durable state, the state cache, messages and threads in one local SQLite file, with WAL mode for concurrent reads. `InMemoryCheckpointer` keeps everything in process, which suits tests. Graph code stays the same. You change the checkpointer and move to `PgCheckpointer` when you deploy.

One difference matters before production: `SqliteCheckpointer` does not implement the tool ledger, so replay safety requires `PgCheckpointer`.

## The same pattern in the API server

The API server uses the same idea for a different question: who owns this thread? Every authenticated request on a thread needs that answer, and asking Postgres every time would put a query in front of every call.

The ownership resolver uses two tiers. L1 is a bounded in-process LRU per worker (10,000 entries, 30-second TTL by default). L2 is an optional shared Redis cache (one-hour TTL by default) so workers do not each query the database. Only positive answers are cached. Caching "this thread has no owner yet" would let a later caller be treated as the creator of a thread someone else had just created. Deleting a thread evicts it from the local L1 and from L2, and the short L1 TTL limits how long another worker can serve a stale answer.

The structure is the same as for state: a fast cache in front, the database as the authority, and TTLs that bound how long a stale answer can live.

## Current limits

- **The cache stores the whole state.** Each cached value is the full serialized state, including the conversation context. Very long contexts make large Redis values. Use a context manager to trim what the agent carries, or a shorter TTL, if values grow large.
- **History is snapshots, not diffs.** Each retained version is a full state copy, so `state_history_limit` multiplies storage per thread.
- **No cross-region story.** The design assumes one Postgres primary and one Redis deployment near your workers. Multi-region replication is left to your infrastructure.
- **Concurrent runs are rejected, not merged.** When two runs race, one wins and the other gets `StaleStateError`. Your client or API layer decides whether to retry on the new state.

## Further reading

- [Memory](/docs/concepts/memory) explains the memory layers and when each one is written.
- [Checkpointing and threads](/docs/concepts/checkpointing-and-threads) covers thread lifecycle, resume and history.
- [Your agent charged the card twice](/blog/your-agent-charged-the-card-twice) covers the tool ledger and versioned writes built on this storage.
- [Installation](/docs/get-started/installation) lists the `pg_checkpoint` and `sqlite_checkpoint` extras.
- [Deploy with Docker](/docs/how-to/api-cli/generate-docker-files) shows running the API server with Redis and Postgres together.

## Frequently asked questions

### How long does 10xGraph keep agent state in Redis?

PgCheckpointer caches thread state for 86400 seconds (24 hours) by default. Change it with the cache_ttl option. Postgres keeps the durable copy whatever the TTL is.

### What happens if Redis is down?

Runs continue on Postgres. A failed cache read falls back to Postgres and a failed cache write is logged without raising. Runs are slower but nothing is lost.

### How much history does PgCheckpointer keep?

The last 20 state snapshots per thread by default, set with state_history_limit. Messages are stored separately in a messages table and are not pruned by that limit.

### Can I run this without Redis or Postgres?

Yes. SqliteCheckpointer keeps durable state and its cache in one local SQLite file, and InMemoryCheckpointer keeps everything in process. Both suit development and tests.
