RAGAgent
In shortRAGAgent retrieves relevant documents from a knowledge base, optionally reranks them, then answers questions grounded in that context.
- 11 min read
- 16 sections
- Updated
- v0.10.0
- Markdown
RAGAgent implements Retrieval-Augmented Generation: it retrieves relevant documents from a knowledge base, optionally reranks them by relevance, and synthesizes the LLM’s answer grounded in that context. Use it for question-answering over custom documents, long documents, or knowledge bases that grow over time.
Import path: tenxgraph.prebuilt.agent
Why RAGAgent
A language model’s answers are limited to its training data, which is static and quickly becomes outdated. RAGAgent solves this by injecting retrieval into the flow: the LLM no longer guesses, it reads. This makes answers more accurate, traceable, and verifiable because the retrieved documents are available in state for every answer.
The pattern is especially effective for:
- Support tickets: Answer questions using your knowledge base, API docs, or FAQ.
- Research: Ground claims in papers, reports, or research documents.
- Legal/compliance: Ensure decisions reference actual policy documents, not memorized rules.
- Customer data: Answer questions about specific customers’ historical interactions or documents.
You retain full control: the LLM’s system prompt, the retrieval strategy, how documents are formatted, and whether to rerank are all configurable.
How it works
RAGAgent runs a three-phase pipeline (or two phases if you skip reranking):
flowchart LR
START([START]) --> RETRIEVE["RETRIEVE\n(vector search)"]
RETRIEVE --> RERANK["RERANK\n(optional: score and trim)"]
RERANK --> SYNTHESIZE["SYNTHESIZE\n(inject docs + call LLM)"]
SYNTHESIZE --> END_NODE([END])
RETRIEVE, The agent extracts the user’s latest question from the conversation and searches the vector store for the top k most similar documents. These are stored internally so subsequent nodes can access them.
RERANK (optional), If you provide a reranker, it re-scores the candidates using a cross-encoder or other ranking function. This filters out false positives from vector similarity and returns only the top n most relevant results. Skipped if no reranker is configured.
SYNTHESIZE, The agent formats the retrieved documents into a <context> block and prepends it to the user’s question. This augmented message is then passed to the underlying LLM (typically a faster or cheaper model), which generates the answer. The agent’s own system prompt remains untouched.
The retrieved documents are always available in state.execution_meta.internal_data["rag_docs"] if you need to inspect or post-process them.
When to use RAGAgent vs other building blocks
Use RAGAgent when:
- Your knowledge base is too large to fit in the system prompt or context window.
- Documents are added, updated, or deleted over time.
- You need traceability: you want to inspect which documents backed each answer.
- The user’s question is about specific documents, not general reasoning.
Use ReactAgent instead if:
- You only need general reasoning or arithmetic, not document lookup.
- Your knowledge base is tiny (fits in the context window).
- The user’s question is unlikely to match documents (e.g., creative brainstorming).
Use a custom StateGraph if:
- You need to combine RAG with tools (e.g., retrieve docs, then call an API).
- You want conditional retrieval (e.g., skip retrieval if the question is about current time).
- You need to chunk documents on-the-fly rather than indexing them upfront.
Reranking: optional but powerful
Vector embeddings measure distance, not relevance. A query might match many documents at similar distances, but only a few are truly relevant. Reranking solves this using a cross-encoder that directly scores (query, document) pairs.
The typical workflow when using a reranker:
- Retrieve many candidates with vector search (
top_k=20): fast, cheap. - Rerank and keep the top few (
top_n=5): slower, more accurate. - Synthesize with only the best candidates.
This two-stage approach keeps the expensive scoring to a small candidate set.
Reranker options
| Name | Backend | Pros | Cons | Install |
|---|---|---|---|---|
| None (default) | Vector similarity only | Fastest, simplest, no API calls | May include irrelevant docs if embeddings are weak | (included) |
CohereReranker |
Cohere Rerank API | Highly accurate, supports any language, fast | Requires API key and network call | pip install cohere |
CrossEncoderReranker |
Local sentence-transformers | No API key, works offline, fully private | CPU-bound (runs in executor to avoid blocking), slower than API | pip install sentence-transformers |
| Custom | Your own class | Integrate your ranking logic | You maintain it | Implement BaseReranker protocol |
Constructor parameters
When you create a RAGAgent, you configure both the retrieval behavior and the underlying graph infrastructure.
Core retrieval parameters
| Parameter | Type | Default | What it does |
|---|---|---|---|
store |
BaseStore |
required | The vector store holding your knowledge base (e.g., QdrantStore). The RETRIEVE node calls store.asearch() to find similar documents. |
agent |
BaseAgent |
required | The LLM that generates the final answer. Typically Agent(model="gpt-4o-mini") or similar. The SYNTHESIZE node calls agent.execute() after injecting documents. |
reranker |
BaseReranker | None |
None |
Optional reranker (CohereReranker, CrossEncoderReranker, or your custom class). If provided, a RERANK node is inserted between RETRIEVE and SYNTHESIZE. |
top_k |
int |
5 |
Number of candidates retrieved from the store. Increase to 15-20 if you use a reranker (it needs more candidates to filter). |
top_n |
int |
3 |
Number of documents forwarded to the LLM after reranking. Ignored if no reranker is provided; all top_k docs go to the LLM instead. |
retrieval_strategy |
RetrievalStrategy |
SIMILARITY |
Vector search strategy (e.g., RetrievalStrategy.SIMILARITY). Passed to store.asearch(). |
score_threshold |
float | None |
None |
Minimum similarity score. Documents below this are filtered out. Useful to exclude near-misses. |
store_config |
dict | None |
None |
Extra key-value pairs passed to every store.asearch() call. Example: {"user_id": "u42"} to filter documents by user. |
Graph infrastructure parameters
| Parameter | Type | Default | What it does |
|---|---|---|---|
state |
AgentState | None |
None |
Optional custom state subclass. If you need extra fields beyond messages and context, subclass AgentState and pass it here. |
context_manager |
BaseContextManager | None |
None |
Strategy for trimming or summarizing context when it grows too large. See /docs/guides/use-context-manager. |
publisher |
BasePublisher | None |
None |
Event publisher for streaming or observability (e.g., ConsolePublisher). Can be a single publisher or a list of publishers. |
id_generator |
BaseIDGenerator |
DefaultIDGenerator() |
Generates run and message IDs. Replace to use UUIDs, Snowflake IDs, or custom formats. |
container |
InjectQ | None |
None |
Dependency injection container. If your tools need injected parameters, pass the InjectQ container here. |
Compile parameters
After constructing RAGAgent, you call .compile(...) to finalize the graph and get a runnable CompiledGraph.
| Parameter | Type | Default | What it does |
|---|---|---|---|
checkpointer |
BaseCheckpointer | None |
None |
Persists conversation state. None uses an InMemoryCheckpointer; use a persistent checkpointer for conversations that survive restarts. Examples: InMemoryCheckpointer(), PgCheckpointer(). |
store |
BaseStore | None |
None |
Long-term memory store for the agent (separate from the retrieval store). Used by memory tools to store facts across conversations. See /docs/guides/use-memory-store. |
interrupt_before |
list[str] | None |
None |
Node names to pause before. Useful for human-in-the-loop: pause before RETRIEVE, review the query, then resume. Valid names: "RETRIEVE", "RERANK", "SYNTHESIZE". |
interrupt_after |
list[str] | None |
None |
Node names to pause after. Example: pause after RETRIEVE to inspect retrieved documents before the LLM sees them. |
callback_manager |
CallbackManager |
CallbackManager() |
Lifecycle hooks (before/after invoke, on error, etc.). Used for logging, monitoring, or side effects. |
media_store |
BaseMediaStore | None |
None |
Storage for images, audio, or documents passed through the conversation. Required if your LLM calls see multimodal content. |
shutdown_timeout |
float |
30.0 |
Seconds to wait for graceful shutdown. |
Getting started: minimal example
The simplest RAGAgent only needs a store and an agent. Vector retrieval alone works well for most use cases. The example assumes the ./knowledge_base collection is already indexed with the same embedding model (see /docs/guides/use-memory-store).
import asyncio
from tenxgraph.core.graph import Agent
from tenxgraph.prebuilt.agent import RAGAgent
from tenxgraph.storage import create_local_qdrant_store
from tenxgraph.storage.store.embedding import OpenAIEmbedding
from tenxgraph.core.state import Message
# Set up the vector store with your knowledge base.
store = create_local_qdrant_store(
path="./knowledge_base",
embedding=OpenAIEmbedding(model="text-embedding-3-small"),
)
# Create the agent.
rag = RAGAgent(
store=store,
agent=Agent(
model="gpt-4o-mini",
provider="openai",
system_prompt=[{"role": "system", "content": "Answer using only the provided context. If not found, say so."}],
),
top_k=5, # Retrieve 5 documents.
)
# Compile and run.
app = rag.compile()
async def main():
result = await app.ainvoke(
{"messages": [Message.text_message("What is the refund policy?")]},
config={"thread_id": "customer-1"},
)
print(result["messages"][-1].text())
asyncio.run(main())Run this with:
pip install "10xgraph[openai,qdrant]"
export OPENAI_API_KEY=sk-...
python script.pyUsing a reranker for higher accuracy
If your vector embeddings are noisy or you need to filter many candidates, add a reranker. Retrieve many, rerank to keep the best few, then answer.
Option 1: Cohere Rerank (API-based)
Cohere’s Rerank API is fast and highly accurate. You retrieve 20 candidates, rerank, and keep the top 5 for the LLM.
import asyncio
from tenxgraph.core.graph import Agent
from tenxgraph.prebuilt.agent import RAGAgent
from tenxgraph.prebuilt.agent.rag import CohereReranker
from tenxgraph.storage import create_local_qdrant_store
from tenxgraph.storage.store.embedding import OpenAIEmbedding
from tenxgraph.core.state import Message
store = create_local_qdrant_store(
path="./knowledge_base",
embedding=OpenAIEmbedding(model="text-embedding-3-small"),
)
rag = RAGAgent(
store=store,
agent=Agent(
model="gpt-4o",
provider="openai",
system_prompt=[{"role": "system", "content": "Use only the provided context."}],
),
reranker=CohereReranker(api_key="YOUR_COHERE_KEY", model="rerank-v4.0-pro"),
top_k=20, # Retrieve 20 candidates.
top_n=5, # Keep top 5 for the LLM.
)
app = rag.compile()
async def main():
result = await app.ainvoke(
{"messages": [Message.text_message("Summarize the warranty terms.")]},
config={"thread_id": "customer-2"},
)
print(result["messages"][-1].text())
asyncio.run(main())Install and run:
pip install "10xgraph[openai,qdrant]" cohere
export OPENAI_API_KEY=sk-...
export COHERE_API_KEY=...
python script.pyOption 2: CrossEncoder (local, no API key)
The sentence-transformers library provides local cross-encoder models. This runs on your machine, no API calls, suitable for private data.
from tenxgraph.core.graph import Agent
from tenxgraph.prebuilt.agent import RAGAgent
from tenxgraph.prebuilt.agent.rag import CrossEncoderReranker
from tenxgraph.storage import create_local_qdrant_store
from tenxgraph.storage.store.embedding import OpenAIEmbedding
store = create_local_qdrant_store(
path="./knowledge_base",
embedding=OpenAIEmbedding(model="text-embedding-3-small"),
)
rag = RAGAgent(
store=store,
agent=Agent(model="gpt-4o-mini", provider="openai"),
reranker=CrossEncoderReranker("cross-encoder/ms-marco-MiniLM-L-6-v2"),
top_k=15,
top_n=4,
)
app = rag.compile()Install:
pip install "10xgraph[openai,qdrant]" sentence-transformersPersistent conversations with checkpointing
To support multi-turn conversations where the LLM remembers previous questions and answers, use a checkpointer.
import asyncio
from tenxgraph.core.graph import Agent
from tenxgraph.prebuilt.agent import RAGAgent
from tenxgraph.storage import create_local_qdrant_store
from tenxgraph.storage.store.embedding import OpenAIEmbedding
from tenxgraph.storage.checkpointer import PgCheckpointer
from tenxgraph.core.state import Message
store = create_local_qdrant_store(
path="./knowledge_base",
embedding=OpenAIEmbedding(model="text-embedding-3-small"),
)
rag = RAGAgent(
store=store,
agent=Agent(
model="gpt-4o-mini",
provider="openai",
system_prompt=[{"role": "system", "content": "Answer questions using only the provided context."}],
),
top_k=5,
)
# Persist state to Postgres + Redis.
checkpointer = PgCheckpointer(
postgres_dsn="postgresql://user:pass@localhost/db",
redis_url="redis://localhost:6379",
)
app = rag.compile(checkpointer=checkpointer)
async def main():
# First turn.
result1 = await app.ainvoke(
{"messages": [Message.text_message("What is the refund policy?")]},
config={"thread_id": "customer-session-1", "user_id": "customer-1"},
)
print("First answer:", result1["messages"][-1].text())
# Second turn in the same thread, the agent remembers the first exchange.
result2 = await app.ainvoke(
{"messages": [Message.text_message("Does it apply to digital products?")]},
config={"thread_id": "customer-session-1", "user_id": "customer-1"},
)
print("Follow-up answer:", result2["messages"][-1].text())
asyncio.run(main())Install and run:
pip install "10xgraph[openai,qdrant,pg_checkpoint]"
export OPENAI_API_KEY=sk-...
python script.pyCustom reranker
Implement the BaseReranker protocol to plug in your own ranking logic.
from tenxgraph.core.graph import Agent
from tenxgraph.prebuilt.agent import RAGAgent
# Any object with an async `arerank` method satisfies the protocol.
class MyCustomReranker:
"""Score documents by length (longest first)."""
async def arerank(self, query: str, documents: list[str], top_n: int) -> list[str]:
# Your ranking logic here.
ranked = sorted(documents, key=len, reverse=True)
return ranked[:top_n]
rag = RAGAgent(
store=store,
agent=Agent(model="gpt-4o-mini", provider="openai"),
reranker=MyCustomReranker(),
top_k=10,
top_n=3,
)Running interactively with 10xgraph play
The 10xgraph play command starts a web-based playground where you can test your agent interactively.
Create these two files:
graph.py
from tenxgraph.core.graph import Agent
from tenxgraph.prebuilt.agent import RAGAgent
from tenxgraph.storage import create_local_qdrant_store
from tenxgraph.storage.store.embedding import OpenAIEmbedding
store = create_local_qdrant_store(
path="./knowledge_base",
embedding=OpenAIEmbedding(model="text-embedding-3-small"),
)
rag = RAGAgent(
store=store,
agent=Agent(
model="gpt-4o-mini",
provider="openai",
system_prompt=[{"role": "system", "content": "Answer questions using the provided context only."}],
),
top_k=5,
)
app = rag.compile()10xgraph.json
{
"agent": "graph:app",
"env": ".env"
}.env
OPENAI_API_KEY=sk-...Then run:
10xgraph playThe command starts the API server and opens the hosted playground. Type your question to see the answer.
Customization: tuning retrieval quality
RAGAgent retrieval quality depends on three factors: embeddings, top_k, and optional reranking.
Improving embeddings
The quality of your vector store’s embeddings directly affects recall. Good embeddings separate relevant documents from irrelevant ones in embedding space.
- Use a model matched to your domain. OpenAI’s
text-embedding-3-smallis a solid default; trytext-embedding-3-largeif you have small, dense documents. - Index documents at the right granularity: chunks of 200-400 tokens typically work well. Too small = missing context; too large = noise.
- Store document metadata (source, date, author) so you can filter or post-process results.
Tuning top_k and top_n
- Without reranker:
top_k=5ortop_k=10retrieves just enough to answer the question. Increase if documents are repetitive or if your embedding model is weak. - With reranker:
top_k=20, top_n=5is a good starting point. The reranker filters out false positives from vector similarity. - Low-latency apps: Minimize
top_k(e.g., 3) and skip the reranker. Vector similarity alone is usually fast enough. - High-accuracy apps: Increase
top_k(20-30) and add a reranker. Spend the extra latency to get the best documents.
Adding a score threshold
If your store supports score_threshold, use it to exclude low-confidence matches:
rag = RAGAgent(
store=store,
agent=agent,
top_k=10,
score_threshold=0.7, # Only docs with similarity >= 0.7
)Per-user or per-session filtering
Use store_config to pass extra parameters to retrieval, e.g., to restrict documents by user:
rag = RAGAgent(
store=store,
agent=agent,
store_config={"user_id": "u42"}, # Passed to every store.asearch() call
)Introspecting retrieved documents
The retrieved documents are always available in state if you need to inspect, log, or post-process them.
import asyncio
from tenxgraph.core.state import Message
from tenxgraph.utils import ResponseGranularity
async def main():
result = await app.ainvoke(
{"messages": [Message.text_message("What is the refund policy?")]},
config={"thread_id": "customer-1"},
response_granularity=ResponseGranularity.FULL,
)
# Inspect the retrieved documents (result["state"] needs FULL granularity).
docs = result["state"].execution_meta.internal_data.get("rag_docs", [])
print(f"Retrieved {len(docs)} documents:")
for i, doc in enumerate(docs, 1):
print(f"[{i}] {doc[:100]}...")
print(f"\nAnswer: {result['messages'][-1].text()}")
asyncio.run(main())Common patterns
RAG + tools
If you need to combine retrieval with tool calls (e.g., retrieve docs, then call an API), use a custom StateGraph instead. See /docs/guides/build-a-graph.
RAG + long context
For very long documents, retrieve chunks and let the LLM synthesize. Avoid putting all documents in the system prompt upfront; RAG handles that dynamically.
Multi-turn RAG
Each turn retrieves fresh documents based on the latest message. The checkpointer remembers the full conversation, so the LLM can refer back to earlier turns while always retrieving the most relevant documents for the current question.
Evaluating RAG quality
Use evaluation sets to measure retrieval and generation quality. See /docs/reference/python/evaluation for the evaluation tools.
Related pages
/docs/guides/prebuilt-agents: Overview of all prebuilt agents and when to use each./docs/concepts/choosing-a-building-block: Decision guide: RAGAgent vs ReactAgent vs custom graph./docs/guides/use-memory-store: Add long-term memory (separate from the retrieval store)./docs/guides/set-up-checkpointing: Configure persistence for multi-turn conversations./docs/reference/python/prebuilt-agents: Full API reference for RAGAgent and all constructor/compile parameters.
Frequently asked questions
- When should I use RAGAgent instead of a plain Agent?
- Use RAGAgent when you need the LLM to answer questions based on a specific knowledge base rather than its training data. It ensures answers are grounded in your documents.
- What's the difference between top_k and top_n?
- top_k is the number of candidates retrieved from the vector store; top_n is how many reach the LLM after reranking. Use top_k=20, top_n=5 to retrieve many, keep the best few.
- Do I need a reranker?
- Only if retrieval accuracy matters more than speed. For most applications, good embeddings and tuned top_k work well. Add a reranker when you need to filter noisy candidates.