Serving Agents
In shortHow 10xgraph.json wires a compiled graph to the API server, plus authentication, authorization, and publisher configuration for production.
- 9 min read
- 12 sections
- Updated
- v0.9.2
- Markdown
This page covers how the API/CLI layer exposes your compiled graph over HTTP, how authentication and authorization protect it, how publishers route execution events to external systems, and what a production deployment looks like.
10xgraph.json — the project config
10xgraph.json is the single file that wires everything together. The CLI and API server read it at startup.
Import paths are dotted module paths (module.path:attribute), resolved with importlib —
not file paths.
{
"agent": "graph.agent:get_compiled_graph",
"auth": { "method": "custom", "path": "auth.agent_auth:MyAuth" },
"injectq": "graph.agent:container",
"evaluation": {
"directory": "evals",
"threshold": 0.8
}
}| Key | Purpose |
|---|---|
agent |
module:callable that returns a CompiledGraph |
auth |
null, "jwt", or {"method": "custom", "path": "module:attr"} for a BaseAuth subclass |
injectq |
Services registered in the DI container |
evaluation |
Eval directory and pass threshold |
Starting the server
agentflow api # starts with auto-reload (development default)
agentflow api --host 0.0.0.0 --port 8000 # bind address
agentflow api --config 10xgraph.json # explicit config path
agentflow play # API + hosted playground in browserThe server loads the compiled graph once at startup and keeps it in memory. All requests share the same graph instance; per-request isolation comes from thread_id. In development --reload is on by default — any change to your source files restarts the server automatically. In production, run with multiple workers (see Production deployment) and omit --reload.
REST endpoints
flowchart TB
subgraph "agentflow api process"
UV[Uvicorn ASGI]
FA[FastAPI]
AUTH[BaseAuth middleware]
AUTHZ[AuthorizationBackend]
RATE[BaseRateLimitBackend]
SVC[GraphService]
GRAPH["Compiled Graph\n(loaded once at startup)"]
end
subgraph "Publishers (BasePublisher)"
PUB_C[ConsolePublisher]
PUB_R[RedisPublisher]
PUB_K[KafkaPublisher]
PUB_Q[RabbitMQPublisher]
PUB_O[OtelPublisher]
end
REQ[HTTP Request] --> UV --> FA --> AUTH --> AUTHZ --> RATE --> SVC --> GRAPH
GRAPH -->|EventModel| PUB_C & PUB_R & PUB_K & PUB_Q & PUB_O
| Router | Prefix | Key endpoints |
|---|---|---|
| Graph | /v1/graph |
POST /invoke, POST /stream, WebSocket /ws, POST /stop, GET / |
| Checkpointer | /v1/threads |
Thread state CRUD, message CRUD |
| Store | /v1/store |
Memory store, search, get, update, delete, list, forget |
| Media | /v1/media |
File upload / download |
| Health | /ping |
Health check |
There is no Agent-to-Agent (A2A) endpoint. The unmounted a2a routers were
removed from the CLI package; agents compose in-process through handoffs, or
across processes over the normal REST API. See the
roadmap.
Authentication
Authentication is pluggable via BaseAuth. The framework ships with JwtAuth; you can replace it with any backend.
flowchart LR
REQ[HTTP Request] --> AUTH[BaseAuth\nauthenticate]
AUTH -->|returns None| R401[401 Unauthorized]
AUTH -->|returns user context| AUTHZ[AuthorizationBackend\ncheck permission]
AUTHZ -->|denied| R403[403 Forbidden]
AUTHZ -->|allowed| ROUTE[Route handler]
Built-in: JwtAuth
Point to the built-in class in 10xgraph.json using its importable path:
{
"auth": "agentflow_cli.src.app.core.auth.jwt_auth:JwtAuth"
}Then set the required environment variables:
export JWT_SECRET_KEY="your-secret"
export JWT_ALGORITHM="HS256" # default; optionalCustom auth — subclass BaseAuth and point 10xgraph.json to your class:
authenticate is synchronous and takes (request, response, credential); the bearer token
arrives as credential. Declaring it async def returns an un-awaited coroutine and breaks auth.
# auth/agent_auth.py
from typing import Any
from fastapi import Request, Response
from fastapi.security import HTTPAuthorizationCredentials
from agentflow_cli import BaseAuth
class FirebaseAuth(BaseAuth):
def authenticate(
self, request: Request, response: Response,
credential: HTTPAuthorizationCredentials | None,
) -> dict[str, Any] | None:
if credential is None:
return None # → 401
try:
claims = firebase_admin.auth.verify_id_token(credential.credentials)
return {"user_id": claims["uid"], **claims}
except Exception:
return None # returning None → 401{
"auth": { "method": "custom", "path": "auth.agent_auth:FirebaseAuth" }
}Authorization
Authorization is a separate extension point from authentication. After a user is identified, AuthorizationBackend decides whether they can perform a specific operation on a specific resource.
# auth/agent_auth.py
from typing import Any
from agentflow_cli.src.app.core.auth.authorization import AuthorizationBackend
class TenantAuthorizationBackend(AuthorizationBackend):
async def authorize(
self, user: dict[str, Any], resource: str, action: str,
resource_id: str | None = None, **context: Any,
) -> bool:
# resource: "graph" | "checkpointer" | "store" | "files" | "config"
# action: "invoke" | "stream" | "read" | "write" | "delete" | ...
# resource_id: thread_id / memory_id when the path carries one
return user.get("tenant_id") == context.get("tenant"){
"authorization": "auth.agent_auth:TenantAuthorizationBackend"
}Without an authorization key the default is mode-based: "ownership" (owner-only threads) in
production, "allow_all" in development. Set "ownership" explicitly, an RBAC config block
({"backend": "rbac", "roles": {...}}), or your own backend to override. See the
Authentication reference for scopes and the isolation policy.
Rate limiting
Rate limiting is pluggable via BaseRateLimitBackend. Two backends are built in; swap or extend via dependency injection.
| Backend | When to use |
|---|---|
| In-memory | Single-process development |
| Redis | Multi-worker production — set REDIS_URL |
| Custom | Subclass BaseRateLimitBackend and register via injectq |
# services/rate_limit.py
from agentflow_cli.src.app.core.middleware.rate_limit.base import BaseRateLimitBackend
class CustomRateLimitBackend(BaseRateLimitBackend):
async def check(self, key: str, limit: int, window: int) -> bool:
# return True to allow, False to rate-limit (→ 429)
...
async def close(self) -> None:
...{
"injectq": {
"BaseRateLimitBackend": "services/rate_limit.py:CustomRateLimitBackend"
}
}Publishers
BasePublisher emits an EventModel on every execution event — node start/end, tool calls, state updates, errors. Wire one or more publishers at StateGraph initialization; they compose automatically.
from tenxgraph.runtime.publisher import RedisPublisher, KafkaPublisher, CompositePublisher
from tenxgraph.core.graph import StateGraph
publisher = CompositePublisher([
RedisPublisher(url="redis://localhost:6379", channel="tenxgraph.events"),
KafkaPublisher(bootstrap_servers="kafka:9092", topic="agentflow"),
])
graph = StateGraph(publisher=publisher)
# ... add nodes and edges ...
compiled = graph.compile()| Publisher | Transport | Use case |
|---|---|---|
ConsolePublisher |
stdout | Development / debugging |
RedisPublisher |
Redis pub/sub | Real-time dashboards, fan-out |
KafkaPublisher |
Kafka topic | High-throughput event pipelines |
RabbitMQPublisher |
RabbitMQ exchange | Queue-based workflows, notifications |
OtelPublisher |
OpenTelemetry | Distributed tracing (Jaeger, Honeycomb, Langfuse) |
Custom publisher — subclass BasePublisher:
from tenxgraph.runtime.publisher.base_publisher import BasePublisher
from tenxgraph.runtime.publisher.events import EventModel
class DatadogPublisher(BasePublisher):
async def publish(self, event: EventModel) -> None:
datadog.send_event(event.dict())
async def close(self) -> None:
passDependency injection
InjectQ is the DI container shipped with 10xgraph. Register service instances into it once, pass it to StateGraph, and node functions receive their dependencies automatically.
Registering services
# graph/agent.py
from injectq import InjectQ
from services.db import DatabaseService
container = InjectQ.get_instance()
container.bind_instance(DatabaseService, DatabaseService())
# Named scalar values (retrieved by key, not by type)
container["api_version"] = "v2"Pass the container to StateGraph at init time:
graph = StateGraph(container=container)Consuming injected dependencies in nodes
Declare dependencies as default parameters using Inject[T]:
from injectq import Inject
from services.db import DatabaseService
async def my_node(
state: AgentState,
config: dict,
db: DatabaseService = Inject[DatabaseService],
) -> Message:
result = await db.query("SELECT ...")
return Message.text_message(str(result), role="assistant")To read named scalar values inside a node:
from injectq import InjectQ
async def my_node(state: AgentState, config: dict) -> Message:
inq = InjectQ.get_instance()
api_version = inq.get("api_version") # raises if missing
request_id = inq.try_get("request_id", "default-id") # returns default if missing
...Always-injected parameters — no annotation needed:
| Parameter name | Value |
|---|---|
state |
Current AgentState |
config |
Run config dict (thread_id, user_id, etc.) |
tool_call_id |
ID of the tool call (inside ToolNode only) |
Wiring the container via 10xgraph.json
When using agentflow api, point injectq to the exported InjectQ instance in your graph module. The server loads that object and activates it as the global singleton.
{
"injectq": "graph.agent:container"
}The value is a dotted module:attribute path that resolves to an InjectQ instance — not a class, not a dict.
Thread name generator
By default the API generates an AI-powered name for each new thread. Override it by subclassing ThreadNameGenerator and registering it via injectq:
# services/naming.py
from agentflow_cli.src.app.utils.thread_name_generator import ThreadNameGenerator
class SlugThreadNameGenerator(ThreadNameGenerator):
async def generate_name(self, messages: list) -> str:
return slugify(messages[0].text[:40]){
"thread_name_generator": "graph.thread_name_generator:SlugThreadNameGenerator"
}Production deployment
flowchart LR
LB[Load Balancer] --> W1[Worker 1]
LB --> W2[Worker 2]
LB --> W3[Worker 3]
W1 & W2 & W3 --> REDIS[(Redis\nhot cache + event bus)]
W1 & W2 & W3 --> PG[(Postgres\ndurable state)]
W1 & W2 & W3 --> QD[(Qdrant\nlong-term memory)]
agentflow build generates a production-ready Dockerfile (and optional docker-compose.yml):
agentflow build # Dockerfile only
agentflow build --docker-compose # + docker-compose.yml
agentflow build --python-version 3.13Key environment variables — set them in a .env file, via export, or as Docker ENV / --env-file:
# .env (or export VAR=value, or Docker ENV in Dockerfile)
MODE=production # enables production guards (warns on ORIGINS=*, etc.)
REDIS_URL=redis://redis:6379
JWT_SECRET_KEY=your-secret-here
SENTRY_DSN=https://[email protected]/123
OTEL_ENABLED=true
OTEL_SERVICE_NAME=my-agent
OTEL_EXPORTER_OTLP_ENDPOINT=http://collector:4317
OTEL_LEVEL=standard| Variable | Default | Purpose |
|---|---|---|
MODE |
development |
Set to production to enable security guards |
REDIS_URL |
None |
Redis for state cache, rate limiter, pub/sub |
JWT_SECRET_KEY |
None |
Required for JwtAuth |
JWT_ALGORITHM |
HS256 |
JWT signing algorithm |
SENTRY_DSN |
None |
Sentry error tracking |
OTEL_ENABLED |
false |
Enable OpenTelemetry tracing (see OpenTelemetry) |
OTEL_SERVICE_NAME |
agentflow-api |
Service name reported in all traces |
OTEL_EXPORTER_OTLP_ENDPOINT |
None |
OTLP collector URL — omit to print spans to console |
OTEL_LEVEL |
standard |
Span detail level: spans | standard | full |
ORIGINS |
* |
CORS allowed origins — restrict in production |
OpenTelemetry
10xGraph has first-class OpenTelemetry support at two independent layers. You can use either or both.
flowchart TB
subgraph "API layer (FastAPI + HTTP)"
FI[FastAPIInstrumentor\nHTTP spans — latency, status, route]
end
subgraph "Graph layer (OtelPublisher)"
GS[tenxgraph.graph span]
NS[tenxgraph.node span]
LS[tenxgraph.llm span\ntoken counts, model, finish reason]
TS[tenxgraph.tool span\ntool name, type]
GS --> NS --> LS
NS --> TS
end
FI -.->|parent| GS
API layer — automatic when OTEL_ENABLED=true
Setting OTEL_ENABLED=true in your environment is all that’s required. The API server automatically:
- Creates a
TracerProviderwith yourOTEL_SERVICE_NAME - Instruments the FastAPI app with
FastAPIInstrumentor(HTTP-level spans) - Wires
OtelPublisherinto the graph so every LLM call, tool call, and node transition becomes a child span - Exports via OTLP when
OTEL_EXPORTER_OTLP_ENDPOINTis set; falls back to console output in non-production
OTEL_ENABLED=true
OTEL_SERVICE_NAME=my-agent
OTEL_EXPORTER_OTLP_ENDPOINT=http://collector:4317 # omit to print spans to console
OTEL_LEVEL=standard # spans | standard | fullNo code changes are needed. The SDK does not need to be configured separately — the API configures OtelPublisher automatically and merges it with any existing publisher (such as RedisPublisher) without replacing it.
Graph layer — OtelPublisher and ObservabilityLevel
When running the graph directly (without agentflow api), pass OtelPublisher to StateGraph at init time:
from tenxgraph.core.graph import StateGraph
from tenxgraph.runtime.publisher import OtelPublisher
from tenxgraph.runtime.publisher.otel_publisher import ObservabilityLevel
graph = StateGraph(publisher=OtelPublisher(level=ObservabilityLevel.STANDARD))
# ... add nodes and edges ...
compiled = graph.compile()ObservabilityLevel controls how much data is emitted as span attributes:
| Level | What it includes |
|---|---|
STANDARD |
Token counts, model name, request params, finish reason (default) |
FULL |
All of STANDARD + prompt messages, completions, tool I/O — may contain PII |
With an explicit TracerProvider (e.g. to export to Jaeger or Honeycomb):
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry import trace
from tenxgraph.core.graph import StateGraph
from tenxgraph.runtime.publisher import OtelPublisher
from tenxgraph.runtime.publisher.otel_publisher import ObservabilityLevel
provider = TracerProvider()
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter(endpoint="http://collector:4317")))
trace.set_tracer_provider(provider)
graph = StateGraph(publisher=OtelPublisher(level=ObservabilityLevel.FULL))
# ... add nodes and edges ...
compiled = graph.compile()Span hierarchy
Every graph run produces a consistent span tree:
tenxgraph.graph ← one per ainvoke / astream call
tenxgraph.node ← one per node execution (e.g. "MAIN", "TOOL")
tenxgraph.llm ← one per LLM call (tokens, model, finish reason)
tenxgraph.tool ← one per tool call (name, type: local | mcp)The tenxgraph.graph span carries thread_id as session.id so tools like Langfuse automatically group multi-turn conversations.
Install:
pip install "10xgraph[otel]" # graph-level spans (OtelPublisher)
pip install "10xscale-agentflow-cli[otel]" # API layer (FastAPIInstrumentor + OTLP exporter)What’s next
| Page | What it covers |
|---|---|
| Connecting Clients | TypeScript SDK, streaming, remote tools |
| Memory | PgCheckpointer, Redis cache, long-term vector store |
| Extensibility | BaseAuth, AuthorizationBackend, BasePublisher and all other ABCs |
| Quality & Observability | GraphLifecycleHook with OpenTelemetry, evaluation, testing |