OpenAI

In shortConfigure the OpenAI provider in 10xGraph to run GPT and reasoning models, including the install extra, API key, and model settings.

  • 6 min read
  • 14 sections
  • Updated
  • v0.10.0
  • Markdown

Run GPT-class models (gpt-4o, gpt-4o-mini) and reasoning models (o3-mini, o4-mini) through the OpenAI API. 10xGraph handles response conversion, tool calling, and prompt caching automatically.

Prerequisites

Install the OpenAI provider extra:

Terminal
pip install "10xgraph[openai]"

Get an API key from platform.openai.com and set it as an environment variable:

Terminal
export OPENAI_API_KEY="sk-..."

Or load it from a .env file:

Terminal
OPENAI_API_KEY=sk-...

Pass api_key= to the Agent instead if you prefer. Without a key, 10xGraph logs a warning and the call fails unless your base_url needs no authentication. 10xGraph detects the provider from the model name or an explicit provider="openai" parameter.

Basic usage

Agent is a graph node, so put it in a StateGraph:

Python
from tenxgraph import Agent, END, Message, StateGraph

agent = Agent(
    model="gpt-4o",
    provider="openai",
    system_prompt=[{"role": "system", "content": "You are a helpful assistant."}],
)

graph = StateGraph()
graph.add_node("MAIN", agent)
graph.add_edge("MAIN", END)
graph.set_entry_point("MAIN")
app = graph.compile()

result = app.invoke({"messages": [Message.text_message("What is 2 + 2?")]})
for message in result["messages"]:
    print(f"{message.role}: {message.text()}")

The provider="openai" parameter is optional if your model name starts with gpt-, o1-, o3- or o4-. Both Chat Completions (the default) and the Responses API are supported; see API Style below.

Full example with tools

Here is a complete example using ReactAgent with a weather tool:

Python
from dotenv import load_dotenv

from tenxgraph.core.state import Message
from tenxgraph.prebuilt.agent import ReactAgent

load_dotenv()

def get_weather(location: str) -> str:
    """Get the current weather for a location."""
    return f"The weather in {location} is sunny."

# Create a ReactAgent with OpenAI
react_agent = ReactAgent(
    model="gpt-4o",
    provider="openai",
    system_prompt=[
        {
            "role": "system",
            "content": "You are a helpful assistant. Use tools when they help answer the user.",
        }
    ],
    tools=[get_weather],
    trim_context=True,
)

if __name__ == "__main__":
    app = react_agent.compile()

    # Invoke the agent
    result = app.invoke(
        {"messages": [Message.text_message("What is the weather in New York City?")]},
        config={"thread_id": "openai-demo", "recursion_limit": 10},
    )

    # Print the messages in the result
    for message in result["messages"]:
        print(f"{message.role}: {message.content}")

Run the example:

Terminal
python your_script.py

You should see the agent call the tool and summarize the weather.

Verify it worked

After running, you will see:

  1. A user message with your question.
  2. An assistant message with a tool call to get_weather.
  3. A tool result message with the weather data.
  4. A final assistant message with the answer.

This confirms the agent is routing tool calls correctly.


API Style

OpenAI exposes two distinct APIs for text generation. 10xGraph supports both.

api_style Underlying call When to use
"chat" (default) client.chat.completions.create All GPT and O-series models. Default for the Agent class.
"responses" client.responses.create Newer Responses API. Default for SummaryContextManager and the evaluation judge.

Agent

Python
from tenxgraph import Agent

# Default, Chat Completions
agent = Agent(model="gpt-4o")

# Opt into the Responses API
agent = Agent(model="gpt-4o", api_style="responses")

SummaryContextManager

Python
from tenxgraph.core.state import SummaryContextManager

# Default is "responses" for the context manager
manager = SummaryContextManager(model="gpt-4o-mini", token_budget=8000)

# Older or third-party-hosted models that only support Chat Completions
manager = SummaryContextManager(
    model="gpt-4o-mini",
    api_style="chat",
    token_budget=8000,
)

Evaluation judge

The evaluation judge reads api_style from CriterionConfig.api_style (defaults to "responses"). Set it per-criterion when using a model that only supports Chat Completions:

Python
from tenxgraph.qa.evaluation import CriterionConfig, EvalConfig, CriteriaConfig

config = EvalConfig(
    criteria=CriteriaConfig(
        llm_judge=CriterionConfig(
            judge_model="gpt-4o",
            api_style="chat",   # override if needed
        )
    )
)

Prompt Caching

OpenAI caches the prompt prefix automatically, no code changes required. Cache hits are reported back in usage.input_tokens_details.cached_tokens and logged at DEBUG level by 10xGraph. You only need to act if you want to improve hit rates.

How it works: OpenAI hashes the first N tokens of your request (system prompt + conversation history + tool definitions). Requests sharing an identical prefix are routed to a server that already has the KV cache in GPU memory. You always send the full prompt; OpenAI serves the cached computation.

Minimum size: 1,024 tokens. Requests below this threshold report zero cached tokens.

TTL: set by OpenAI (short, in-memory). Some models support extended retention through prompt_cache_retention="24h"; check OpenAI’s documentation for which.

prompt_cache_key

When multiple Agent instances (or multiple processes) share the same long system prompt, pass a stable key to colocate them on the same cached server and raise the hit rate.

Python
agent = Agent(
    model="gpt-4o",
    system_prompt=[{"role": "system", "content": very_long_prompt}],
    prompt_cache_key="legal-analyst-v2",  # passed through llm_kwargs
)

This is an OpenAI request-level parameter forwarded directly through llm_kwargs. It is not in CALL_EXCLUDED_KWARGS so it reaches the API unchanged.

prompt_cache_retention

Requests extended cache retention. OpenAI only honours it on models that support it.

Python
agent = Agent(
    model="gpt-4o",
    system_prompt=[...],
    prompt_cache_key="assistant-v1",
    prompt_cache_retention="24h",
)

SummaryContextManager with caching

SummaryContextManager does not accept extra LLM keyword arguments, so it cannot send prompt_cache_key. The implicit cache still applies when its prompt prefix is stable.

Evaluation judge with caching

Extra kwargs are not forwarded from CriterionConfig to the judge call. The implicit cache still fires automatically when the judge prompt prefix is stable (same rubric, same model).


Reasoning Models

Reasoning models such as o4-mini are controlled via reasoning_config. The default is {"effort": "medium"}.

Python
# Enable with defaults (effort="medium")
agent = Agent(model="o4-mini", reasoning_config=True)

# Set effort level
agent = Agent(
    model="o4-mini",
    reasoning_config={"effort": "high"},  # "low" | "medium" | "high"
)

# Disable
agent = Agent(model="o4-mini", reasoning_config=False)

On Chat Completions, effort is sent as reasoning_effort; on the Responses API the whole dict is sent as reasoning (so summary is also allowed). Because the default is on, pass reasoning_config=None when using a model that does not accept reasoning parameters.


Structured Output

Force the model to return a Pydantic model by passing output_schema. The Agent routes this through beta.chat.completions.parse regardless of api_style.

Python
from pydantic import BaseModel

class MyOutput(BaseModel):
    answer: str
    confidence: float

agent = Agent(
    model="gpt-4o",
    output_schema=MyOutput,
    system_prompt=[...],
)

Caching still applies to the beta.chat.completions.parse path, cache hits are logged the same way.


OpenAI-Compatible Endpoints

Any OpenAI-compatible server (ollama, vllm, LM Studio, etc.) can be used by setting base_url:

Python
agent = Agent(
    model="llama3.2",
    base_url="http://localhost:11434/v1",
    api_style="chat",  # most local servers only support Chat Completions
)

When base_url is set and api_style="responses", the Agent tries the Responses API first and falls back to Chat Completions automatically if the server does not support it.

Prompt caching params (prompt_cache_key, prompt_cache_retention) have no effect on local servers unless the server explicitly implements them.


llm_kwargs Reference

All unrecognised keyword arguments passed to Agent(...) land in self.llm_kwargs and are forwarded to the underlying API call.

kwarg Type Applies to Notes
prompt_cache_key str Chat + Responses Improves cross-request cache hit rate
prompt_cache_retention "in_memory" / "24h" Chat + Responses 24h only on supporting models
temperature float Chat + Responses Sampling temperature (0.0-2.0)
max_tokens int Chat Max output tokens
max_output_tokens int Responses Max output tokens (Responses API name)
reasoning_effort str Chat Normally derived from reasoning_config; dropped on the Responses API
top_p float Chat + Responses Nucleus sampling
frequency_penalty float Chat Penalise repeated tokens
presence_penalty float Chat Penalise already-seen tokens

Keys in CALL_EXCLUDED_KWARGS (organization, project, timeout, max_retries, default_headers, default_query, http_client, api_key, base_url) are stripped before the request is sent and must be passed to the client constructor instead.


Environment Variables

Variable Required Description
OPENAI_API_KEY yes API key from platform.openai.com

Common Errors

Error Fix
AuthenticationError OPENAI_API_KEY missing or invalid
RateLimitError You hit a rate limit. Retries are on by default; tune them with retry_config=RetryConfig(max_retries=5) (from tenxgraph.core import RetryConfig)
Model not found Check the model name; some models require tier-gated access

Next steps

  • See Models for a full capability matrix comparing OpenAI with Google and Anthropic.
  • Read Configure an Agent to learn about fallbacks, retries, and other provider-agnostic settings.
  • Explore Structured output for typed responses with output_schema.

Frequently asked questions

Which OpenAI models work with 10xGraph?
Any model your OpenAI account can call, for example gpt-4o, gpt-4o-mini and o4-mini. Names starting with gpt-, o1-, o3- or o4- are detected as OpenAI automatically. See platform.openai.com for the current list.
How do I enable caching for OpenAI?
Prompt caching is automatic on OpenAI. Cache hits are logged at DEBUG level. Pass prompt_cache_key for stable cross-request hits.
Do I need to change code to use the Responses API?
No. Chat Completions is the default. Pass api_style='responses' to your Agent if you want the Responses API.
Last updated for v0.10.0Edit this page on GitHubReport an issue