Hot and cold agent memory: why 10xGraph puts Redis in front of Postgres
How PgCheckpointer splits agent state between a Redis hot cache and versioned Postgres history, keeps them consistent, and behaves when either one fails.
Every agent framework has to decide where conversation state lives between steps. The decision looks like plumbing, but it sets your per-step latency, what a crash loses, how many workers can serve one thread, and what happens on the day Redis restarts.
10xGraph’s production checkpointer, PgCheckpointer, uses two stores with separate jobs. Redis is the hot layer: the latest state of each active thread, read first and refreshed on every step, with a TTL. Postgres is the cold layer: versioned state snapshots, the full message history, thread metadata and the tool-call ledger. It is the source of truth.
Running two stores is easy. Keeping them consistent under crashes, retries and concurrent runs takes more care, and most of this post is about that.
What an agent run asks of storage
The access pattern of a single run looks like this:
- Load the thread at the start: the current state, including the conversation context and where execution stopped.
- Execute a node: a model call, a tool node, or your own function.
- Save the thread after the node: new state, plus any messages the node produced.
- Repeat steps 2 and 3 until the graph ends, pauses for a human, or fails.
A twenty-step run reads once and writes twenty times. A chat product with many active threads does this across thousands of threads at once, and the same thread is often picked up by a different API worker on the next turn. The storage layer needs fast reads of the latest state from any worker, cheap writes on every step, durability for anything a user or auditor might need later, and correct behavior when two executions touch one thread.
No single common store does all of that well.
Why not Postgres alone?
Postgres can store everything, and in 10xGraph it does. The problem is putting it on the read path for every load. Every run that starts or resumes, and every “what is the state of this thread?” request from a UI, becomes a query that finds the newest snapshot row and deserializes a JSON payload that grows with the conversation.
Each query is fast on its own. Across thousands of active threads and several API workers, they add up to steady load on the primary database and latency on the critical path of every turn. The latest state of an active thread is also the most frequently read and most cacheable data in the system. Serving all of it from the system of record is wasteful.
Why not Redis alone?
Redis is the right home for state you read constantly and can rebuild. It is the wrong home for the only copy.
- TTL expiry. Thread state needs a TTL, or memory grows without bound as conversations accumulate. A thread idle longer than the TTL disappears.
- Eviction. Under memory pressure, Redis evicts keys according to its
maxmemory-policy. An active customer’s thread can be evicted to make room for another. - Restarts. Depending on persistence settings, a restart can lose recent writes or the whole dataset.
- No history. Redis holds the current value. Audits, debugging and “what did the agent see at step 7?” need history.
An agent that forgets a customer’s open order after a quiet weekend has a bug, not a cache miss.
The split
Each store gets the job it is good at:
| Data | Redis (hot) | Postgres (cold) |
|---|---|---|
| Latest thread state | JSON value with a TTL, key state_cache:<thread_id>:<user_id> |
Versioned rows in states |
| State history | None | The last state_history_limit snapshots per thread |
| Messages | Only inside the cached state payload | One row per message in messages |
| Thread metadata and owner | None | threads table |
| Tool-call ledger | None | tool_executions table |
| Version | Embedded in the cached payload | version column, unique per thread |
Two rules govern the whole design:
- Postgres is the source of truth. Anything in Redis can be thrown away and rebuilt from Postgres.
- The cache must never get ahead of the truth. It must never hold state that was not persisted, and never move backwards in a way that misleads the next writer.
The write path, the read path and the version guards all follow from these two rules.
The write path
After each completed step, the run loop saves the thread. The order is the important part:
if checkpointer:
# Durable store is the source of truth: state + messages, atomically, first.
await checkpointer.aput_checkpoint(config, new_state, messages)
# Then refresh the realtime cache with the same state we just persisted.
await checkpointer.aput_state_cache(config, new_state)Durable first, cache second. If the Postgres write fails, the exception surfaces before the cache is touched, so the cache cannot hold state that was never persisted. The cache is then written with the same state object that was saved durably, so the two copies cannot differ in shape. This matters when a context manager trims the conversation before saving.
State and messages in one transaction. aput_checkpoint writes the new state row and that step’s messages in a single Postgres transaction. A crash cannot leave the state advanced without the messages that justify it, or the reverse.
Only new messages are sent. The run loop tracks how many messages have already been persisted and sends only the rest. A long run does not rewrite its history on every step. Message inserts are upserts keyed by message_id, restricted to the same thread. A message ID that already belongs to another thread is rejected and rolls the transaction back, so one thread cannot overwrite another thread’s messages.
State is an append-only log with a retention window. Each write inserts a new row at the next version number. In the same transaction, rows older than the newest state_history_limit (default 20) are deleted. You get recent snapshots for debugging and rollback without unbounded growth. Messages are not affected by this limit.
Connection errors are retried. A dropped connection or interface error retries the whole transaction on a fresh pooled connection, up to three attempts with exponential backoff (one second, then two). Other errors, and connection errors that outlast the retries, are raised as a StorageError, so the run fails visibly instead of continuing on unsaved state.
You can disable the per-step durable write with "durable_checkpoint_every_step": False in the run config. Steps then write only the cache, and Postgres is written when the run completes, errors, pauses or stops. That saves database writes but brings back the risk described above: a crash after the cache expires resumes from the last terminal checkpoint. The default is per-step durability, and you should keep it for any agent whose work matters.
The read path
At the start of a run, the loop calls aget_state_cache:
- Try Redis. On a hit, deserialize the state and restore its embedded version into the run config under
_checkpoint_version, so the next durable write can still perform its version check. - On a miss, read Postgres. Select the highest-version snapshot for the thread, record its version in the run config, and seed the cache so the next read is a hit.
- If Redis raises, log the error and read Postgres. A Redis outage makes loads slower and does not make them fail.
If the cache returns nothing, the loop also tries the durable store directly before deciding the thread is new.
The rule throughout is that a cache problem must never become a correctness problem. Cache writes are best-effort too. A failed aput_state_cache is logged and returns None, and the run continues.
The hard part: keeping the cache honest under concurrency
Two executions can touch the same thread: a client retry racing the original run, or a resumed run racing a new message. Postgres handles this with optimistic concurrency. Each run remembers the version it started from, and a write fails with StaleStateError if the thread moved on in the meantime. The companion post, Your agent charged the card twice, covers that mechanism.
A cache in front of a versioned store adds a less obvious way to fail. Here it is step by step:
- Run A and run B both load thread T at version 5.
- Run B finishes first. Postgres moves to version 6 and the cache holds version 6.
- Run A is still working. After its next step it writes the cache with its own state, still based on version 5. A plain
SETEXaccepts it, and the cache now holds stale state stamped version 5. - Run A’s durable write fails its version check, correctly.
- The next run on T reads the cache, which is preferred, picks up version 5 as its expected version, and fails its version check too.
- So does every run after that, until the TTL expires. With the default TTL, the thread is stuck for up to a day.
Nothing crashed and every component did what it was told. The thread is still unusable. 10xGraph prevents this with two guards.
Guard 1: the cache never moves backwards
Cache writes go through a small Lua script that Redis runs atomically. It reads the version embedded in the current entry and only writes if the incoming version is at least as new:
local existing = redis.call('GET', KEYS[1])
if existing then
local ok, decoded = pcall(cjson.decode, existing)
if ok and type(decoded) == 'table' and decoded[ARGV[4]] then
if tonumber(decoded[ARGV[4]]) > tonumber(ARGV[2]) then
return 0
end
end
end
redis.call('SETEX', KEYS[1], ARGV[3], ARGV[1])
return 1Running the compare and the set inside Redis matters. A GET followed by a SETEX from Python would still race between coroutines and between workers. Equal versions are allowed on purpose: within a single run, state changes on every step (progress, stop flags) while the durable version stays the same until the next durable write. A brand-new thread with no durable version yet writes with a sentinel of -1, which is lower than any real version and so can never overwrite a newer entry.
In the scenario above, step 3 now returns 0. Run A’s stale write is skipped and the cache keeps version 6.
Guard 2: a conflict clears the cache entry
If a durable write raises StaleStateError, the checkpointer deletes that thread’s cache entry before re-raising. The cache might hold state built on the version that just lost. Deleting it forces the next read back to Postgres, which returns the true latest version, and the thread recovers on its next run instead of staying stuck until the TTL expires.
Either guard covers most cases on its own. Together they make sure a stale entry cannot drive later writes.
Isolation is part of the key
With authentication enabled, a thread belongs to a user, and knowing a thread ID must not be enough to read it. PgCheckpointer treats user_id as an ownership boundary by default (enforce_user_isolation=True):
- The cache key includes both IDs,
state_cache:<thread_id>:<user_id>, so one user’s cache lookups can never return another user’s entry. - Durable reads, writes and deletes are scoped to the owning user. On write, the thread row is locked and its owner checked, and a write to another user’s thread raises a
STORAGE_FORBIDDEN_001error.
For single-tenant deployments, or setups with no real user identity, set enforce_user_isolation=False and the thread ID alone is the key. Runs without a user_id share one "anonymous" identity, which matches what the API server uses when auth is off.
Failure matrix
What each failure costs, with the default settings:
| Event | What the agent sees | Data lost |
|---|---|---|
| Redis unreachable | Loads read Postgres. Cache writes are logged and skipped. | None |
| Redis restarts empty | First load of each thread misses, reads Postgres and re-seeds the cache | None |
| Cache TTL expires on an idle thread | Same as a restart, for that thread | None |
| Key evicted under memory pressure | Same as a TTL expiry | None |
| Postgres connection drops briefly | Transaction retried on a fresh connection, up to three attempts | None |
| Postgres unavailable | Durable write raises and the run fails visibly | The in-flight step, which is replayed on resume |
| Worker killed mid-node | Next run resumes at that node. Completed tools come from the ledger. | At most the in-flight node |
| Two runs race on a thread | The loser gets StaleStateError and its cache entry is cleared |
None. The losing write is rejected rather than merged. |
The only row that loses work is the one where the source of truth is down, which is the trade-off you want.
Tuning it
from tenxgraph.storage.checkpointer import PgCheckpointer
checkpointer = PgCheckpointer(
postgres_dsn="postgresql://user:password@localhost/app",
redis_url="redis://localhost:6379/0",
cache_ttl=3600, # default 86400 (24 hours)
state_history_limit=20, # default 20
)
app = graph.compile(checkpointer=checkpointer)Point the checkpointer key in 10xgraph.json at this object, for example "checkpointer": "graph.agent:checkpointer", and the API server builds and uses it.
cache_ttl should cover the typical gap between turns in an active conversation. For live chat, an hour or a few hours keeps working sets small. For agents users come back to over days, keep the 24-hour default or raise it. A short TTL never loses data. It only means more reads go to Postgres.
state_history_limit sets how far back you can inspect or roll back a thread’s state. Twenty snapshots fits most debugging. Raise it for agents you audit step by step, and lower it if state payloads are large.
Redis memory policy. Every thread key has a TTL, so a volatile-* eviction policy such as volatile-lru evicts thread state under pressure and leaves keys without a TTL alone. An eviction is only a cache miss, so this is safe. The noeviction policy is the one to avoid. Under memory pressure it makes cache writes fail. 10xGraph tolerates that, but every read then goes to Postgres.
Pool sizes. Both pools can be configured (pool_config for asyncpg, redis_pool_config for Redis), or you can pass in pools your application already owns. The checkpointer only closes resources it created, so a shared pool is never closed from under your app.
Local development: the same layers in one file
You do not need Redis and Postgres running to build an agent. SqliteCheckpointer keeps durable state, the state cache, messages and threads in one local SQLite file, with WAL mode for concurrent reads. InMemoryCheckpointer keeps everything in process, which suits tests. Graph code stays the same. You change the checkpointer and move to PgCheckpointer when you deploy.
One difference matters before production: SqliteCheckpointer does not implement the tool ledger, so replay safety requires PgCheckpointer.
The same pattern in the API server
The API server uses the same idea for a different question: who owns this thread? Every authenticated request on a thread needs that answer, and asking Postgres every time would put a query in front of every call.
The ownership resolver uses two tiers. L1 is a bounded in-process LRU per worker (10,000 entries, 30-second TTL by default). L2 is an optional shared Redis cache (one-hour TTL by default) so workers do not each query the database. Only positive answers are cached. Caching “this thread has no owner yet” would let a later caller be treated as the creator of a thread someone else had just created. Deleting a thread evicts it from the local L1 and from L2, and the short L1 TTL limits how long another worker can serve a stale answer.
The structure is the same as for state: a fast cache in front, the database as the authority, and TTLs that bound how long a stale answer can live.
Current limits
- The cache stores the whole state. Each cached value is the full serialized state, including the conversation context. Very long contexts make large Redis values. Use a context manager to trim what the agent carries, or a shorter TTL, if values grow large.
- History is snapshots, not diffs. Each retained version is a full state copy, so
state_history_limitmultiplies storage per thread. - No cross-region story. The design assumes one Postgres primary and one Redis deployment near your workers. Multi-region replication is left to your infrastructure.
- Concurrent runs are rejected, not merged. When two runs race, one wins and the other gets
StaleStateError. Your client or API layer decides whether to retry on the new state.
Further reading
- Memory explains the memory layers and when each one is written.
- Checkpointing and threads covers thread lifecycle, resume and history.
- Your agent charged the card twice covers the tool ledger and versioned writes built on this storage.
- Installation lists the
pg_checkpointandsqlite_checkpointextras. - Deploy with Docker shows running the API server with Redis and Postgres together.
Frequently asked questions
- How long does 10xGraph keep agent state in Redis?
- PgCheckpointer caches thread state for 86400 seconds (24 hours) by default. Change it with the cache_ttl option. Postgres keeps the durable copy whatever the TTL is.
- What happens if Redis is down?
- Runs continue on Postgres. A failed cache read falls back to Postgres and a failed cache write is logged without raising. Runs are slower but nothing is lost.
- How much history does PgCheckpointer keep?
- The last 20 state snapshots per thread by default, set with state_history_limit. Messages are stored separately in a messages table and are not pruned by that limit.
- Can I run this without Redis or Postgres?
- Yes. SqliteCheckpointer keeps durable state and its cache in one local SQLite file, and InMemoryCheckpointer keeps everything in process. Both suit development and tests.