Serving Agents

In shortHow 10xgraph.json wires a compiled graph to the API server, plus authentication, authorization, and publisher configuration for production.

  • 9 min read
  • 12 sections
  • Updated
  • v0.9.2
  • Markdown

This page covers how the API/CLI layer exposes your compiled graph over HTTP, how authentication and authorization protect it, how publishers route execution events to external systems, and what a production deployment looks like.


10xgraph.json — the project config

10xgraph.json is the single file that wires everything together. The CLI and API server read it at startup.

Import paths are dotted module paths (module.path:attribute), resolved with importlib — not file paths.

JSON
{
  "agent": "graph.agent:get_compiled_graph",
  "auth": { "method": "custom", "path": "auth.agent_auth:MyAuth" },
  "injectq": "graph.agent:container",
  "evaluation": {
    "directory": "evals",
    "threshold": 0.8
  }
}
Key Purpose
agent module:callable that returns a CompiledGraph
auth null, "jwt", or {"method": "custom", "path": "module:attr"} for a BaseAuth subclass
injectq Services registered in the DI container
evaluation Eval directory and pass threshold

Starting the server

Terminal
agentflow api                                    # starts with auto-reload (development default)
agentflow api --host 0.0.0.0 --port 8000        # bind address
agentflow api --config 10xgraph.json            # explicit config path
agentflow play                                   # API + hosted playground in browser

The server loads the compiled graph once at startup and keeps it in memory. All requests share the same graph instance; per-request isolation comes from thread_id. In development --reload is on by default — any change to your source files restarts the server automatically. In production, run with multiple workers (see Production deployment) and omit --reload.


REST endpoints

flowchart TB
  subgraph "agentflow api process"
    UV[Uvicorn ASGI]
    FA[FastAPI]
    AUTH[BaseAuth middleware]
    AUTHZ[AuthorizationBackend]
    RATE[BaseRateLimitBackend]
    SVC[GraphService]
    GRAPH["Compiled Graph\n(loaded once at startup)"]
  end
  subgraph "Publishers  (BasePublisher)"
    PUB_C[ConsolePublisher]
    PUB_R[RedisPublisher]
    PUB_K[KafkaPublisher]
    PUB_Q[RabbitMQPublisher]
    PUB_O[OtelPublisher]
  end
  REQ[HTTP Request] --> UV --> FA --> AUTH --> AUTHZ --> RATE --> SVC --> GRAPH
  GRAPH -->|EventModel| PUB_C & PUB_R & PUB_K & PUB_Q & PUB_O
Router Prefix Key endpoints
Graph /v1/graph POST /invoke, POST /stream, WebSocket /ws, POST /stop, GET /
Checkpointer /v1/threads Thread state CRUD, message CRUD
Store /v1/store Memory store, search, get, update, delete, list, forget
Media /v1/media File upload / download
Health /ping Health check

There is no Agent-to-Agent (A2A) endpoint. The unmounted a2a routers were removed from the CLI package; agents compose in-process through handoffs, or across processes over the normal REST API. See the roadmap.


Authentication

Authentication is pluggable via BaseAuth. The framework ships with JwtAuth; you can replace it with any backend.

flowchart LR
  REQ[HTTP Request] --> AUTH[BaseAuth\nauthenticate]
  AUTH -->|returns None| R401[401 Unauthorized]
  AUTH -->|returns user context| AUTHZ[AuthorizationBackend\ncheck permission]
  AUTHZ -->|denied| R403[403 Forbidden]
  AUTHZ -->|allowed| ROUTE[Route handler]

Built-in: JwtAuth

Point to the built-in class in 10xgraph.json using its importable path:

JSON
{
  "auth": "agentflow_cli.src.app.core.auth.jwt_auth:JwtAuth"
}

Then set the required environment variables:

Terminal
export JWT_SECRET_KEY="your-secret"
export JWT_ALGORITHM="HS256"      # default; optional

Custom auth — subclass BaseAuth and point 10xgraph.json to your class:

authenticate is synchronous and takes (request, response, credential); the bearer token arrives as credential. Declaring it async def returns an un-awaited coroutine and breaks auth.

Python
# auth/agent_auth.py
from typing import Any

from fastapi import Request, Response
from fastapi.security import HTTPAuthorizationCredentials

from agentflow_cli import BaseAuth

class FirebaseAuth(BaseAuth):
    def authenticate(
        self, request: Request, response: Response,
        credential: HTTPAuthorizationCredentials | None,
    ) -> dict[str, Any] | None:
        if credential is None:
            return None   # → 401
        try:
            claims = firebase_admin.auth.verify_id_token(credential.credentials)
            return {"user_id": claims["uid"], **claims}
        except Exception:
            return None   # returning None → 401
JSON
{
  "auth": { "method": "custom", "path": "auth.agent_auth:FirebaseAuth" }
}

Authorization

Authorization is a separate extension point from authentication. After a user is identified, AuthorizationBackend decides whether they can perform a specific operation on a specific resource.

Python
# auth/agent_auth.py
from typing import Any

from agentflow_cli.src.app.core.auth.authorization import AuthorizationBackend

class TenantAuthorizationBackend(AuthorizationBackend):
    async def authorize(
        self, user: dict[str, Any], resource: str, action: str,
        resource_id: str | None = None, **context: Any,
    ) -> bool:
        # resource: "graph" | "checkpointer" | "store" | "files" | "config"
        # action:   "invoke" | "stream" | "read" | "write" | "delete" | ...
        # resource_id: thread_id / memory_id when the path carries one
        return user.get("tenant_id") == context.get("tenant")
JSON
{
  "authorization": "auth.agent_auth:TenantAuthorizationBackend"
}

Without an authorization key the default is mode-based: "ownership" (owner-only threads) in production, "allow_all" in development. Set "ownership" explicitly, an RBAC config block ({"backend": "rbac", "roles": {...}}), or your own backend to override. See the Authentication reference for scopes and the isolation policy.


Rate limiting

Rate limiting is pluggable via BaseRateLimitBackend. Two backends are built in; swap or extend via dependency injection.

Backend When to use
In-memory Single-process development
Redis Multi-worker production — set REDIS_URL
Custom Subclass BaseRateLimitBackend and register via injectq
Python
# services/rate_limit.py
from agentflow_cli.src.app.core.middleware.rate_limit.base import BaseRateLimitBackend

class CustomRateLimitBackend(BaseRateLimitBackend):
    async def check(self, key: str, limit: int, window: int) -> bool:
        # return True to allow, False to rate-limit (→ 429)
        ...

    async def close(self) -> None:
        ...
JSON
{
  "injectq": {
    "BaseRateLimitBackend": "services/rate_limit.py:CustomRateLimitBackend"
  }
}

Publishers

BasePublisher emits an EventModel on every execution event — node start/end, tool calls, state updates, errors. Wire one or more publishers at StateGraph initialization; they compose automatically.

Python
from tenxgraph.runtime.publisher import RedisPublisher, KafkaPublisher, CompositePublisher
from tenxgraph.core.graph import StateGraph

publisher = CompositePublisher([
    RedisPublisher(url="redis://localhost:6379", channel="tenxgraph.events"),
    KafkaPublisher(bootstrap_servers="kafka:9092", topic="agentflow"),
])

graph = StateGraph(publisher=publisher)
# ... add nodes and edges ...
compiled = graph.compile()
Publisher Transport Use case
ConsolePublisher stdout Development / debugging
RedisPublisher Redis pub/sub Real-time dashboards, fan-out
KafkaPublisher Kafka topic High-throughput event pipelines
RabbitMQPublisher RabbitMQ exchange Queue-based workflows, notifications
OtelPublisher OpenTelemetry Distributed tracing (Jaeger, Honeycomb, Langfuse)

Custom publisher — subclass BasePublisher:

Python
from tenxgraph.runtime.publisher.base_publisher import BasePublisher
from tenxgraph.runtime.publisher.events import EventModel

class DatadogPublisher(BasePublisher):
    async def publish(self, event: EventModel) -> None:
        datadog.send_event(event.dict())

    async def close(self) -> None:
        pass

Dependency injection

InjectQ is the DI container shipped with 10xgraph. Register service instances into it once, pass it to StateGraph, and node functions receive their dependencies automatically.

Registering services

Python
# graph/agent.py
from injectq import InjectQ
from services.db import DatabaseService

container = InjectQ.get_instance()
container.bind_instance(DatabaseService, DatabaseService())

# Named scalar values (retrieved by key, not by type)
container["api_version"] = "v2"

Pass the container to StateGraph at init time:

Python
graph = StateGraph(container=container)

Consuming injected dependencies in nodes

Declare dependencies as default parameters using Inject[T]:

Python
from injectq import Inject
from services.db import DatabaseService

async def my_node(
    state: AgentState,
    config: dict,
    db: DatabaseService = Inject[DatabaseService],
) -> Message:
    result = await db.query("SELECT ...")
    return Message.text_message(str(result), role="assistant")

To read named scalar values inside a node:

Python
from injectq import InjectQ

async def my_node(state: AgentState, config: dict) -> Message:
    inq = InjectQ.get_instance()
    api_version = inq.get("api_version")                    # raises if missing
    request_id = inq.try_get("request_id", "default-id")   # returns default if missing
    ...

Always-injected parameters — no annotation needed:

Parameter name Value
state Current AgentState
config Run config dict (thread_id, user_id, etc.)
tool_call_id ID of the tool call (inside ToolNode only)

Wiring the container via 10xgraph.json

When using agentflow api, point injectq to the exported InjectQ instance in your graph module. The server loads that object and activates it as the global singleton.

JSON
{
  "injectq": "graph.agent:container"
}

The value is a dotted module:attribute path that resolves to an InjectQ instance — not a class, not a dict.


Thread name generator

By default the API generates an AI-powered name for each new thread. Override it by subclassing ThreadNameGenerator and registering it via injectq:

Python
# services/naming.py
from agentflow_cli.src.app.utils.thread_name_generator import ThreadNameGenerator

class SlugThreadNameGenerator(ThreadNameGenerator):
    async def generate_name(self, messages: list) -> str:
        return slugify(messages[0].text[:40])
JSON
{
  "thread_name_generator": "graph.thread_name_generator:SlugThreadNameGenerator"
}

Production deployment

flowchart LR
  LB[Load Balancer] --> W1[Worker 1]
  LB --> W2[Worker 2]
  LB --> W3[Worker 3]
  W1 & W2 & W3 --> REDIS[(Redis\nhot cache + event bus)]
  W1 & W2 & W3 --> PG[(Postgres\ndurable state)]
  W1 & W2 & W3 --> QD[(Qdrant\nlong-term memory)]

agentflow build generates a production-ready Dockerfile (and optional docker-compose.yml):

Terminal
agentflow build                          # Dockerfile only
agentflow build --docker-compose         # + docker-compose.yml
agentflow build --python-version 3.13

Key environment variables — set them in a .env file, via export, or as Docker ENV / --env-file:

Terminal
# .env  (or export VAR=value, or Docker ENV in Dockerfile)
MODE=production           # enables production guards (warns on ORIGINS=*, etc.)
REDIS_URL=redis://redis:6379
JWT_SECRET_KEY=your-secret-here
SENTRY_DSN=https://[email protected]/123
OTEL_ENABLED=true
OTEL_SERVICE_NAME=my-agent
OTEL_EXPORTER_OTLP_ENDPOINT=http://collector:4317
OTEL_LEVEL=standard
Variable Default Purpose
MODE development Set to production to enable security guards
REDIS_URL None Redis for state cache, rate limiter, pub/sub
JWT_SECRET_KEY None Required for JwtAuth
JWT_ALGORITHM HS256 JWT signing algorithm
SENTRY_DSN None Sentry error tracking
OTEL_ENABLED false Enable OpenTelemetry tracing (see OpenTelemetry)
OTEL_SERVICE_NAME agentflow-api Service name reported in all traces
OTEL_EXPORTER_OTLP_ENDPOINT None OTLP collector URL — omit to print spans to console
OTEL_LEVEL standard Span detail level: spans | standard | full
ORIGINS * CORS allowed origins — restrict in production

OpenTelemetry

10xGraph has first-class OpenTelemetry support at two independent layers. You can use either or both.

flowchart TB
  subgraph "API layer  (FastAPI + HTTP)"
    FI[FastAPIInstrumentor\nHTTP spans — latency, status, route]
  end
  subgraph "Graph layer  (OtelPublisher)"
    GS[tenxgraph.graph span]
    NS[tenxgraph.node span]
    LS[tenxgraph.llm span\ntoken counts, model, finish reason]
    TS[tenxgraph.tool span\ntool name, type]
    GS --> NS --> LS
    NS --> TS
  end
  FI -.->|parent| GS

API layer — automatic when OTEL_ENABLED=true

Setting OTEL_ENABLED=true in your environment is all that’s required. The API server automatically:

  • Creates a TracerProvider with your OTEL_SERVICE_NAME
  • Instruments the FastAPI app with FastAPIInstrumentor (HTTP-level spans)
  • Wires OtelPublisher into the graph so every LLM call, tool call, and node transition becomes a child span
  • Exports via OTLP when OTEL_EXPORTER_OTLP_ENDPOINT is set; falls back to console output in non-production
Terminal
OTEL_ENABLED=true
OTEL_SERVICE_NAME=my-agent
OTEL_EXPORTER_OTLP_ENDPOINT=http://collector:4317   # omit to print spans to console
OTEL_LEVEL=standard                                  # spans | standard | full

No code changes are needed. The SDK does not need to be configured separately — the API configures OtelPublisher automatically and merges it with any existing publisher (such as RedisPublisher) without replacing it.

Graph layer — OtelPublisher and ObservabilityLevel

When running the graph directly (without agentflow api), pass OtelPublisher to StateGraph at init time:

Python
from tenxgraph.core.graph import StateGraph
from tenxgraph.runtime.publisher import OtelPublisher
from tenxgraph.runtime.publisher.otel_publisher import ObservabilityLevel

graph = StateGraph(publisher=OtelPublisher(level=ObservabilityLevel.STANDARD))
# ... add nodes and edges ...
compiled = graph.compile()

ObservabilityLevel controls how much data is emitted as span attributes:

Level What it includes
STANDARD Token counts, model name, request params, finish reason (default)
FULL All of STANDARD + prompt messages, completions, tool I/O — may contain PII

With an explicit TracerProvider (e.g. to export to Jaeger or Honeycomb):

Python
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry import trace

from tenxgraph.core.graph import StateGraph
from tenxgraph.runtime.publisher import OtelPublisher
from tenxgraph.runtime.publisher.otel_publisher import ObservabilityLevel

provider = TracerProvider()
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter(endpoint="http://collector:4317")))
trace.set_tracer_provider(provider)

graph = StateGraph(publisher=OtelPublisher(level=ObservabilityLevel.FULL))
# ... add nodes and edges ...
compiled = graph.compile()

Span hierarchy

Every graph run produces a consistent span tree:

plaintext
tenxgraph.graph          ← one per ainvoke / astream call
  tenxgraph.node         ← one per node execution (e.g. "MAIN", "TOOL")
    tenxgraph.llm        ← one per LLM call (tokens, model, finish reason)
    tenxgraph.tool       ← one per tool call (name, type: local | mcp)

The tenxgraph.graph span carries thread_id as session.id so tools like Langfuse automatically group multi-turn conversations.

Install:

Terminal
pip install "10xgraph[otel]"           # graph-level spans (OtelPublisher)
pip install "10xscale-agentflow-cli[otel]"       # API layer (FastAPIInstrumentor + OTLP exporter)

What’s next

Page What it covers
Connecting Clients TypeScript SDK, streaming, remote tools
Memory PgCheckpointer, Redis cache, long-term vector store
Extensibility BaseAuth, AuthorizationBackend, BasePublisher and all other ABCs
Quality & Observability GraphLifecycleHook with OpenTelemetry, evaluation, testing
Last updated for v0.9.2Edit this page on GitHubReport an issue