Build a Data Extraction AI Agent in Python
In shortBuild a data extraction agent in Python that turns unstructured text into validated Pydantic records and retries when validation fails, with 10xGraph.
- 6 min read
- 10 sections
- Updated
- v0.10.0
- Markdown
A data extraction agent reads unstructured text, such as an invoice, and saves it as a typed record. It calls a submit tool that validates the data with Pydantic. If validation fails, the error goes back to the model, which fixes the record and resubmits, so only valid data is stored.
This page builds that loop with one agent node, one tool node and a routing function. It is a runnable script, and the in-memory list stands in for your database.
How the validation loop works
The agent never writes data directly. It calls submit_invoice, and that tool validates the arguments against a Pydantic model. A valid record is saved. An invalid one returns the error text as the tool result, and the graph sends control back to the agent to try again.
document text
|
v
MAIN (agent) --- no tool call ---> END
|
| tool call
v
TOOL (submit_invoice) --- always returns to MAINThe retry needs no special edge. The tool result lands in the message history, the agent reads the error on its next turn, and calls the tool again. The run ends when the agent replies without a tool call. A recursion limit stops a model that can never produce a valid record.
Prerequisites
Install the core package with the Anthropic extra. Pydantic is a dependency of the core package. Set the API key your provider SDK expects (for Anthropic, ANTHROPIC_API_KEY) in your shell. To use OpenAI or Google instead, install that extra and change the model string.
pip install "10xgraph[anthropic]"Define the schema
The Pydantic model is both the contract and the validator. Field descriptions tell the model what each field means, and Decimal keeps money exact where float would not.
# schema.py
from datetime import date
from decimal import Decimal
from pydantic import BaseModel, Field, model_validator
class InvoiceLineItem(BaseModel):
description: str
quantity: int
unit_price: Decimal
total: Decimal
class Invoice(BaseModel):
invoice_number: str = Field(description="The invoice ID printed on the document")
issue_date: date
due_date: date
vendor_name: str
total_amount: Decimal
line_items: list[InvoiceLineItem]
@model_validator(mode="after")
def check_business_rules(self) -> "Invoice":
# Business rules the model must satisfy, reported back as errors
if self.due_date < self.issue_date:
raise ValueError("due_date must not be before issue_date")
if sum(item.total for item in self.line_items) != self.total_amount:
raise ValueError("line item totals must add up to total_amount")
return selfThe model_validator shows the real value of the loop: a rule such as “line items add up to the total” produces an error message the model can act on.
Build the extraction graph
The graph has two nodes. MAIN is an Agent whose tool_node="TOOL" points at the tool node by name, and route sends the run to TOOL when the agent asked for a tool, or to END when it answered in text.
The system prompt is passed through Python string formatting at run time, so literal braces in the JSON schema are doubled. Without that, the schema would be treated as placeholders and the prompt would not be interpolated.
# extraction_agent.py
import json
from tenxgraph.core import Agent, AgentState, Message, StateGraph, ToolNode
from tenxgraph.utils import END
from schema import Invoice
# Stand-in for a database table
extracted_invoices: list[dict] = []
def submit_invoice(invoice: dict) -> str:
"""Submit the extracted invoice. Call this with the full structured data."""
try:
validated = Invoice.model_validate(invoice)
except Exception as exc:
# Returned to the model as the tool result so it can fix the data
return f"Validation error: {exc}. Fix the issues and resubmit."
extracted_invoices.append(validated.model_dump(mode="json"))
return f"Saved invoice {validated.invoice_number}."
# Double the braces so Agent's prompt formatting leaves the JSON schema intact
schema_text = json.dumps(Invoice.model_json_schema()).replace("{", "{{").replace("}", "}}")
extractor = Agent(
model="claude-sonnet-4-5", # use any model id you have access to
system_prompt=[
{
"role": "system",
"content": (
"Extract invoice data from the document the user sends. "
"Call the submit_invoice tool with an object matching this JSON schema: "
f"{schema_text}. "
"If submit_invoice returns a validation error, fix the data and call it again. "
"Do not summarize and do not chat."
),
}
],
tool_node="TOOL",
)
def route(state: AgentState) -> str:
"""Go to the tool node when the agent requested a tool, otherwise finish."""
last = state.context[-1] if state.context else None
if last and last.role == "assistant" and last.tools_calls:
return "TOOL"
return END
graph = StateGraph()
graph.add_node("MAIN", extractor)
graph.add_node("TOOL", ToolNode([submit_invoice]))
graph.add_conditional_edges("MAIN", route, {"TOOL": "TOOL", END: END})
graph.add_edge("TOOL", "MAIN") # tool results always return to the agent
graph.set_entry_point("MAIN")
app = graph.compile()Run the agent on a sample invoice
Call app.invoke with the document text as a user message. The recursion_limit counts every node execution, so a first-try success takes three steps (agent, tool, agent) and each retry adds two more. A limit of 9 allows about three retries, and the run raises GraphRecursionError if it is exceeded.
# run.py
from tenxgraph.core import Message
from extraction_agent import app, extracted_invoices
invoice_text = """
INVOICE #INV-2024-001
Issue Date: 2024-01-15
Due Date: 2024-02-15
Vendor: Acme Corporation
Items:
- Widget A x 10 @ $5.00 = $50.00
- Widget B x 5 @ $10.00 = $50.00
Total: $100.00
"""
result = app.invoke(
{"messages": [Message.text_message(invoice_text)]},
config={"thread_id": "extract-inv-001", "recursion_limit": 9},
)
# The last message is the agent's closing reply; the data is in your store
print(result["messages"][-1].text())
print(extracted_invoices[-1])Check the result by inspecting extracted_invoices, not the agent’s closing text. The model’s wording varies between runs, but a stored record has always passed Invoice.model_validate. For this input it contains invoice_number INV-2024-001, the two line items and a total_amount of 100.00.
Operational considerations
Treat these as the checklist before you point the agent at real documents.
- Model choice. Use a strong tool-calling model. Smaller models miss nested fields and produce invalid arguments more often, which costs retries.
- Preprocess the text. Strip headers, footers, page numbers and decorative text. The model spends attention on them and may extract them as data.
- Validate business rules. Type checks catch format errors. Custom validators, like the totals check above, catch values that are well-formed but wrong.
- Log failures with their input. When a run hits the recursion limit, store the document and the last error (redact personal data first). Those cases show where to improve the prompt or schema.
- Use idempotent thread IDs. Set
thread_idto a hash of the document so the same document is not processed twice under different threads. - Cap retries. If the model fails three or four times, it is usually stuck on a structural problem, not a typo, and more attempts only cost money.
Adapt the pattern to other documents
The graph shape stays the same and only the schema, tool name and prompt change.
| Task | Schema focus |
|---|---|
| Form filling | Required fields collected from a chat, with validators for each |
| Email triage | Sender, intent, action items, and an Enum for urgency |
| Contract review | Parties, dates and obligations as nested models |
| Scraped page cleanup | Typed records from semi-structured text |
When to use this pattern
Use it when the output goes into a database or downstream code and you need to know it is valid. The loop gives you a hard guarantee that stored records passed validation, and a bounded cost when the model cannot comply.
Do not use it when a person will only read the result (summarize instead), or when a single call with a schema is enough. For one-shot typed output without retries, see structured output. If the agent must also look things up, use a ReAct agent or a RAG agent.
See also
- ReAct agent with validation: input validators and callbacks on a tool-using agent
- Custom state: extend
AgentStatefor richer workflows - Structured output: typed output from a single model call
- Get started: build your first agent
Frequently asked questions
- How do I make extraction more reliable?
- Validate every result with a Pydantic schema, return the validation error to the model as the tool result, and let it resubmit. Add custom validators for business rules and cap retries with recursion_limit.
- When should I extract instead of summarize?
- Extract when you need structured fields in a database or downstream code. Summarize when a person will read the result. Extraction with validation is checkable, a summary is not.
- Can I use this for contracts, emails or scraped pages?
- Yes. The graph stays the same and only the schema and the system prompt change. If a single call is enough and you do not need retries, use the structured output guide instead.