OpenAI
In shortConfigure the OpenAI provider in 10xGraph to run GPT and reasoning models, including the install extra, API key, and model settings.
- 6 min read
- 14 sections
- Updated
- v0.10.0
- Markdown
Run GPT-class models (gpt-4o, gpt-4o-mini) and reasoning models (o3-mini, o4-mini) through the OpenAI API. 10xGraph handles response conversion, tool calling, and prompt caching automatically.
Prerequisites
Install the OpenAI provider extra:
pip install "10xgraph[openai]"Get an API key from platform.openai.com and set it as an environment variable:
export OPENAI_API_KEY="sk-..."Or load it from a .env file:
OPENAI_API_KEY=sk-...Pass api_key= to the Agent instead if you prefer. Without a key, 10xGraph logs a warning and the call fails unless your base_url needs no authentication. 10xGraph detects the provider from the model name or an explicit provider="openai" parameter.
Basic usage
Agent is a graph node, so put it in a StateGraph:
from tenxgraph import Agent, END, Message, StateGraph
agent = Agent(
model="gpt-4o",
provider="openai",
system_prompt=[{"role": "system", "content": "You are a helpful assistant."}],
)
graph = StateGraph()
graph.add_node("MAIN", agent)
graph.add_edge("MAIN", END)
graph.set_entry_point("MAIN")
app = graph.compile()
result = app.invoke({"messages": [Message.text_message("What is 2 + 2?")]})
for message in result["messages"]:
print(f"{message.role}: {message.text()}")The provider="openai" parameter is optional if your model name starts with gpt-, o1-, o3- or o4-. Both Chat Completions (the default) and the Responses API are supported; see API Style below.
Full example with tools
Here is a complete example using ReactAgent with a weather tool:
from dotenv import load_dotenv
from tenxgraph.core.state import Message
from tenxgraph.prebuilt.agent import ReactAgent
load_dotenv()
def get_weather(location: str) -> str:
"""Get the current weather for a location."""
return f"The weather in {location} is sunny."
# Create a ReactAgent with OpenAI
react_agent = ReactAgent(
model="gpt-4o",
provider="openai",
system_prompt=[
{
"role": "system",
"content": "You are a helpful assistant. Use tools when they help answer the user.",
}
],
tools=[get_weather],
trim_context=True,
)
if __name__ == "__main__":
app = react_agent.compile()
# Invoke the agent
result = app.invoke(
{"messages": [Message.text_message("What is the weather in New York City?")]},
config={"thread_id": "openai-demo", "recursion_limit": 10},
)
# Print the messages in the result
for message in result["messages"]:
print(f"{message.role}: {message.content}")Run the example:
python your_script.pyYou should see the agent call the tool and summarize the weather.
Verify it worked
After running, you will see:
- A user message with your question.
- An assistant message with a tool call to
get_weather. - A tool result message with the weather data.
- A final assistant message with the answer.
This confirms the agent is routing tool calls correctly.
API Style
OpenAI exposes two distinct APIs for text generation. 10xGraph supports both.
api_style |
Underlying call | When to use |
|---|---|---|
"chat" (default) |
client.chat.completions.create |
All GPT and O-series models. Default for the Agent class. |
"responses" |
client.responses.create |
Newer Responses API. Default for SummaryContextManager and the evaluation judge. |
Agent
from tenxgraph import Agent
# Default, Chat Completions
agent = Agent(model="gpt-4o")
# Opt into the Responses API
agent = Agent(model="gpt-4o", api_style="responses")SummaryContextManager
from tenxgraph.core.state import SummaryContextManager
# Default is "responses" for the context manager
manager = SummaryContextManager(model="gpt-4o-mini", token_budget=8000)
# Older or third-party-hosted models that only support Chat Completions
manager = SummaryContextManager(
model="gpt-4o-mini",
api_style="chat",
token_budget=8000,
)Evaluation judge
The evaluation judge reads api_style from CriterionConfig.api_style (defaults to
"responses"). Set it per-criterion when using a model that only supports Chat Completions:
from tenxgraph.qa.evaluation import CriterionConfig, EvalConfig, CriteriaConfig
config = EvalConfig(
criteria=CriteriaConfig(
llm_judge=CriterionConfig(
judge_model="gpt-4o",
api_style="chat", # override if needed
)
)
)Prompt Caching
OpenAI caches the prompt prefix automatically, no code changes required. Cache hits are
reported back in usage.input_tokens_details.cached_tokens and logged at DEBUG level
by 10xGraph. You only need to act if you want to improve hit rates.
How it works: OpenAI hashes the first N tokens of your request (system prompt + conversation history + tool definitions). Requests sharing an identical prefix are routed to a server that already has the KV cache in GPU memory. You always send the full prompt; OpenAI serves the cached computation.
Minimum size: 1,024 tokens. Requests below this threshold report zero cached tokens.
TTL: set by OpenAI (short, in-memory). Some models support extended retention through
prompt_cache_retention="24h"; check OpenAI’s documentation for which.
prompt_cache_key
When multiple Agent instances (or multiple processes) share the same long system prompt, pass a stable key to colocate them on the same cached server and raise the hit rate.
agent = Agent(
model="gpt-4o",
system_prompt=[{"role": "system", "content": very_long_prompt}],
prompt_cache_key="legal-analyst-v2", # passed through llm_kwargs
)This is an OpenAI request-level parameter forwarded directly through llm_kwargs. It is
not in CALL_EXCLUDED_KWARGS so it reaches the API unchanged.
prompt_cache_retention
Requests extended cache retention. OpenAI only honours it on models that support it.
agent = Agent(
model="gpt-4o",
system_prompt=[...],
prompt_cache_key="assistant-v1",
prompt_cache_retention="24h",
)SummaryContextManager with caching
SummaryContextManager does not accept extra LLM keyword arguments, so it cannot
send prompt_cache_key. The implicit cache still applies when its prompt prefix is stable.
Evaluation judge with caching
Extra kwargs are not forwarded from CriterionConfig to the judge call. The implicit
cache still fires automatically when the judge prompt prefix is stable (same rubric, same model).
Reasoning Models
Reasoning models such as o4-mini are controlled via reasoning_config. The default is
{"effort": "medium"}.
# Enable with defaults (effort="medium")
agent = Agent(model="o4-mini", reasoning_config=True)
# Set effort level
agent = Agent(
model="o4-mini",
reasoning_config={"effort": "high"}, # "low" | "medium" | "high"
)
# Disable
agent = Agent(model="o4-mini", reasoning_config=False)On Chat Completions, effort is sent as reasoning_effort; on the Responses API the whole
dict is sent as reasoning (so summary is also allowed). Because the default is on, pass
reasoning_config=None when using a model that does not accept reasoning parameters.
Structured Output
Force the model to return a Pydantic model by passing output_schema. The Agent routes
this through beta.chat.completions.parse regardless of api_style.
from pydantic import BaseModel
class MyOutput(BaseModel):
answer: str
confidence: float
agent = Agent(
model="gpt-4o",
output_schema=MyOutput,
system_prompt=[...],
)Caching still applies to the beta.chat.completions.parse path, cache hits are logged
the same way.
OpenAI-Compatible Endpoints
Any OpenAI-compatible server (ollama, vllm, LM Studio, etc.) can be used by setting
base_url:
agent = Agent(
model="llama3.2",
base_url="http://localhost:11434/v1",
api_style="chat", # most local servers only support Chat Completions
)When base_url is set and api_style="responses", the Agent tries the Responses API
first and falls back to Chat Completions automatically if the server does not support it.
Prompt caching params (prompt_cache_key, prompt_cache_retention) have no effect on
local servers unless the server explicitly implements them.
llm_kwargs Reference
All unrecognised keyword arguments passed to Agent(...) land in self.llm_kwargs and
are forwarded to the underlying API call.
| kwarg | Type | Applies to | Notes |
|---|---|---|---|
prompt_cache_key |
str |
Chat + Responses | Improves cross-request cache hit rate |
prompt_cache_retention |
"in_memory" / "24h" |
Chat + Responses | 24h only on supporting models |
temperature |
float |
Chat + Responses | Sampling temperature (0.0-2.0) |
max_tokens |
int |
Chat | Max output tokens |
max_output_tokens |
int |
Responses | Max output tokens (Responses API name) |
reasoning_effort |
str |
Chat | Normally derived from reasoning_config; dropped on the Responses API |
top_p |
float |
Chat + Responses | Nucleus sampling |
frequency_penalty |
float |
Chat | Penalise repeated tokens |
presence_penalty |
float |
Chat | Penalise already-seen tokens |
Keys in CALL_EXCLUDED_KWARGS (organization, project, timeout, max_retries,
default_headers, default_query, http_client, api_key, base_url) are stripped
before the request is sent and must be passed to the client constructor instead.
Environment Variables
| Variable | Required | Description |
|---|---|---|
OPENAI_API_KEY |
yes | API key from platform.openai.com |
Common Errors
| Error | Fix |
|---|---|
AuthenticationError |
OPENAI_API_KEY missing or invalid |
RateLimitError |
You hit a rate limit. Retries are on by default; tune them with retry_config=RetryConfig(max_retries=5) (from tenxgraph.core import RetryConfig) |
Model not found |
Check the model name; some models require tier-gated access |
Next steps
- See Models for a full capability matrix comparing OpenAI with Google and Anthropic.
- Read Configure an Agent to learn about fallbacks, retries, and other provider-agnostic settings.
- Explore Structured output for typed responses with
output_schema.
Frequently asked questions
- Which OpenAI models work with 10xGraph?
- Any model your OpenAI account can call, for example gpt-4o, gpt-4o-mini and o4-mini. Names starting with gpt-, o1-, o3- or o4- are detected as OpenAI automatically. See platform.openai.com for the current list.
- How do I enable caching for OpenAI?
- Prompt caching is automatic on OpenAI. Cache hits are logged at DEBUG level. Pass prompt_cache_key for stable cross-request hits.
- Do I need to change code to use the Responses API?
- No. Chat Completions is the default. Pass api_style='responses' to your Agent if you want the Responses API.