# Serving Agents

> How 10xgraph.json wires a compiled graph to the API server, plus authentication, authorization, and publisher configuration for production.

Source: https://10xgraph.com/docs/concepts/serving-agents
Last updated: 2026-07-21

This page covers how the API/CLI layer exposes your compiled graph over HTTP, how authentication and authorization protect it, how publishers route execution events to external systems, and what a production deployment looks like.

---

## `10xgraph.json` — the project config

`10xgraph.json` is the single file that wires everything together. The CLI and API server read it at startup.

Import paths are **dotted module paths** (`module.path:attribute`), resolved with `importlib` —
not file paths.

```json
{
  "agent": "graph.agent:get_compiled_graph",
  "auth": { "method": "custom", "path": "auth.agent_auth:MyAuth" },
  "injectq": "graph.agent:container",
  "evaluation": {
    "directory": "evals",
    "threshold": 0.8
  }
}
```

| Key | Purpose |
|---|---|
| `agent` | `module:callable` that returns a `CompiledGraph` |
| `auth` | `null`, `"jwt"`, or `{"method": "custom", "path": "module:attr"}` for a `BaseAuth` subclass |
| `injectq` | Services registered in the DI container |
| `evaluation` | Eval directory and pass threshold |

---

## Starting the server

```bash
agentflow api                                    # starts with auto-reload (development default)
agentflow api --host 0.0.0.0 --port 8000        # bind address
agentflow api --config 10xgraph.json            # explicit config path
agentflow play                                   # API + hosted playground in browser
```

The server loads the compiled graph once at startup and keeps it in memory. All requests share the same graph instance; per-request isolation comes from `thread_id`. In development `--reload` is on by default — any change to your source files restarts the server automatically. In production, run with multiple workers (see [Production deployment](#production-deployment)) and omit `--reload`.

---

## REST endpoints

```mermaid
flowchart TB
  subgraph "agentflow api process"
    UV[Uvicorn ASGI]
    FA[FastAPI]
    AUTH[BaseAuth middleware]
    AUTHZ[AuthorizationBackend]
    RATE[BaseRateLimitBackend]
    SVC[GraphService]
    GRAPH["Compiled Graph\n(loaded once at startup)"]
  end
  subgraph "Publishers  (BasePublisher)"
    PUB_C[ConsolePublisher]
    PUB_R[RedisPublisher]
    PUB_K[KafkaPublisher]
    PUB_Q[RabbitMQPublisher]
    PUB_O[OtelPublisher]
  end
  REQ[HTTP Request] --> UV --> FA --> AUTH --> AUTHZ --> RATE --> SVC --> GRAPH
  GRAPH -->|EventModel| PUB_C & PUB_R & PUB_K & PUB_Q & PUB_O
```

| Router | Prefix | Key endpoints |
|---|---|---|
| Graph | `/v1/graph` | `POST /invoke`, `POST /stream`, `WebSocket /ws`, `POST /stop`, `GET /` |
| Checkpointer | `/v1/threads` | Thread state CRUD, message CRUD |
| Store | `/v1/store` | Memory store, search, get, update, delete, list, forget |
| Media | `/v1/media` | File upload / download |
| Health | `/ping` | Health check |

There is no Agent-to-Agent (A2A) endpoint. The unmounted `a2a` routers were
removed from the CLI package; agents compose in-process through handoffs, or
across processes over the normal REST API. See the
[roadmap](/docs/project/roadmap).

---

## Authentication

Authentication is pluggable via `BaseAuth`. The framework ships with `JwtAuth`; you can replace it with any backend.

```mermaid
flowchart LR
  REQ[HTTP Request] --> AUTH[BaseAuth\nauthenticate]
  AUTH -->|returns None| R401[401 Unauthorized]
  AUTH -->|returns user context| AUTHZ[AuthorizationBackend\ncheck permission]
  AUTHZ -->|denied| R403[403 Forbidden]
  AUTHZ -->|allowed| ROUTE[Route handler]
```

**Built-in: `JwtAuth`**

Point to the built-in class in `10xgraph.json` using its importable path:

```json
{
  "auth": "agentflow_cli.src.app.core.auth.jwt_auth:JwtAuth"
}
```

Then set the required environment variables:

```bash
export JWT_SECRET_KEY="your-secret"
export JWT_ALGORITHM="HS256"      # default; optional
```

**Custom auth** — subclass `BaseAuth` and point `10xgraph.json` to your class:

`authenticate` is **synchronous** and takes `(request, response, credential)`; the bearer token
arrives as `credential`. Declaring it `async def` returns an un-awaited coroutine and breaks auth.

```python
# auth/agent_auth.py
from typing import Any

from fastapi import Request, Response
from fastapi.security import HTTPAuthorizationCredentials

from agentflow_cli import BaseAuth

class FirebaseAuth(BaseAuth):
    def authenticate(
        self, request: Request, response: Response,
        credential: HTTPAuthorizationCredentials | None,
    ) -> dict[str, Any] | None:
        if credential is None:
            return None   # → 401
        try:
            claims = firebase_admin.auth.verify_id_token(credential.credentials)
            return {"user_id": claims["uid"], **claims}
        except Exception:
            return None   # returning None → 401
```

```json
{
  "auth": { "method": "custom", "path": "auth.agent_auth:FirebaseAuth" }
}
```

---

## Authorization

Authorization is a separate extension point from authentication. After a user is identified, `AuthorizationBackend` decides whether they can perform a specific operation on a specific resource.

```python
# auth/agent_auth.py
from typing import Any

from agentflow_cli.src.app.core.auth.authorization import AuthorizationBackend

class TenantAuthorizationBackend(AuthorizationBackend):
    async def authorize(
        self, user: dict[str, Any], resource: str, action: str,
        resource_id: str | None = None, **context: Any,
    ) -> bool:
        # resource: "graph" | "checkpointer" | "store" | "files" | "config"
        # action:   "invoke" | "stream" | "read" | "write" | "delete" | ...
        # resource_id: thread_id / memory_id when the path carries one
        return user.get("tenant_id") == context.get("tenant")
```

```json
{
  "authorization": "auth.agent_auth:TenantAuthorizationBackend"
}
```

Without an `authorization` key the default is mode-based: `"ownership"` (owner-only threads) in
production, `"allow_all"` in development. Set `"ownership"` explicitly, an RBAC config block
(`{"backend": "rbac", "roles": {...}}`), or your own backend to override. See the
[Authentication reference](/docs/reference/api-cli/auth) for scopes and the isolation policy.

---

## Rate limiting

Rate limiting is pluggable via `BaseRateLimitBackend`. Two backends are built in; swap or extend via dependency injection.

| Backend | When to use |
|---|---|
| In-memory | Single-process development |
| Redis | Multi-worker production — set `REDIS_URL` |
| Custom | Subclass `BaseRateLimitBackend` and register via `injectq` |

```python
# services/rate_limit.py
from agentflow_cli.src.app.core.middleware.rate_limit.base import BaseRateLimitBackend

class CustomRateLimitBackend(BaseRateLimitBackend):
    async def check(self, key: str, limit: int, window: int) -> bool:
        # return True to allow, False to rate-limit (→ 429)
        ...

    async def close(self) -> None:
        ...
```

```json
{
  "injectq": {
    "BaseRateLimitBackend": "services/rate_limit.py:CustomRateLimitBackend"
  }
}
```

---

## Publishers

`BasePublisher` emits an `EventModel` on every execution event — node start/end, tool calls, state updates, errors. Wire one or more publishers at `StateGraph` initialization; they compose automatically.

```python
from tenxgraph.runtime.publisher import RedisPublisher, KafkaPublisher, CompositePublisher
from tenxgraph.core.graph import StateGraph

publisher = CompositePublisher([
    RedisPublisher(url="redis://localhost:6379", channel="tenxgraph.events"),
    KafkaPublisher(bootstrap_servers="kafka:9092", topic="agentflow"),
])

graph = StateGraph(publisher=publisher)
# ... add nodes and edges ...
compiled = graph.compile()
```

| Publisher | Transport | Use case |
|---|---|---|
| `ConsolePublisher` | stdout | Development / debugging |
| `RedisPublisher` | Redis pub/sub | Real-time dashboards, fan-out |
| `KafkaPublisher` | Kafka topic | High-throughput event pipelines |
| `RabbitMQPublisher` | RabbitMQ exchange | Queue-based workflows, notifications |
| `OtelPublisher` | OpenTelemetry | Distributed tracing (Jaeger, Honeycomb, Langfuse) |

Custom publisher — subclass `BasePublisher`:

```python
from tenxgraph.runtime.publisher.base_publisher import BasePublisher
from tenxgraph.runtime.publisher.events import EventModel

class DatadogPublisher(BasePublisher):
    async def publish(self, event: EventModel) -> None:
        datadog.send_event(event.dict())

    async def close(self) -> None:
        pass
```

---

## Dependency injection

`InjectQ` is the DI container shipped with `10xgraph`. Register service instances into it once, pass it to `StateGraph`, and node functions receive their dependencies automatically.

### Registering services

```python
# graph/agent.py
from injectq import InjectQ
from services.db import DatabaseService

container = InjectQ.get_instance()
container.bind_instance(DatabaseService, DatabaseService())

# Named scalar values (retrieved by key, not by type)
container["api_version"] = "v2"
```

Pass the container to `StateGraph` at init time:

```python
graph = StateGraph(container=container)
```

### Consuming injected dependencies in nodes

Declare dependencies as default parameters using `Inject[T]`:

```python
from injectq import Inject
from services.db import DatabaseService

async def my_node(
    state: AgentState,
    config: dict,
    db: DatabaseService = Inject[DatabaseService],
) -> Message:
    result = await db.query("SELECT ...")
    return Message.text_message(str(result), role="assistant")
```

To read named scalar values inside a node:

```python
from injectq import InjectQ

async def my_node(state: AgentState, config: dict) -> Message:
    inq = InjectQ.get_instance()
    api_version = inq.get("api_version")                    # raises if missing
    request_id = inq.try_get("request_id", "default-id")   # returns default if missing
    ...
```

Always-injected parameters — no annotation needed:

| Parameter name | Value |
|---|---|
| `state` | Current `AgentState` |
| `config` | Run config dict (`thread_id`, `user_id`, etc.) |
| `tool_call_id` | ID of the tool call (inside `ToolNode` only) |

### Wiring the container via `10xgraph.json`

When using `agentflow api`, point `injectq` to the exported `InjectQ` instance in your graph module. The server loads that object and activates it as the global singleton.

```json
{
  "injectq": "graph.agent:container"
}
```

The value is a dotted `module:attribute` path that resolves to an `InjectQ` instance — not a class, not a dict.

---

## Thread name generator

By default the API generates an AI-powered name for each new thread. Override it by subclassing `ThreadNameGenerator` and registering it via `injectq`:

```python
# services/naming.py
from agentflow_cli.src.app.utils.thread_name_generator import ThreadNameGenerator

class SlugThreadNameGenerator(ThreadNameGenerator):
    async def generate_name(self, messages: list) -> str:
        return slugify(messages[0].text[:40])
```

```json
{
  "thread_name_generator": "graph.thread_name_generator:SlugThreadNameGenerator"
}
```

---

## Production deployment

```mermaid
flowchart LR
  LB[Load Balancer] --> W1[Worker 1]
  LB --> W2[Worker 2]
  LB --> W3[Worker 3]
  W1 & W2 & W3 --> REDIS[(Redis\nhot cache + event bus)]
  W1 & W2 & W3 --> PG[(Postgres\ndurable state)]
  W1 & W2 & W3 --> QD[(Qdrant\nlong-term memory)]
```

`agentflow build` generates a production-ready `Dockerfile` (and optional `docker-compose.yml`):

```bash
agentflow build                          # Dockerfile only
agentflow build --docker-compose         # + docker-compose.yml
agentflow build --python-version 3.13
```

Key environment variables — set them in a `.env` file, via `export`, or as Docker `ENV` / `--env-file`:

```bash
# .env  (or export VAR=value, or Docker ENV in Dockerfile)
MODE=production           # enables production guards (warns on ORIGINS=*, etc.)
REDIS_URL=redis://redis:6379
JWT_SECRET_KEY=your-secret-here
SENTRY_DSN=https://...@sentry.io/123
OTEL_ENABLED=true
OTEL_SERVICE_NAME=my-agent
OTEL_EXPORTER_OTLP_ENDPOINT=http://collector:4317
OTEL_LEVEL=standard
```

| Variable | Default | Purpose |
|---|---|---|
| `MODE` | `development` | Set to `production` to enable security guards |
| `REDIS_URL` | `None` | Redis for state cache, rate limiter, pub/sub |
| `JWT_SECRET_KEY` | `None` | Required for `JwtAuth` |
| `JWT_ALGORITHM` | `HS256` | JWT signing algorithm |
| `SENTRY_DSN` | `None` | Sentry error tracking |
| `OTEL_ENABLED` | `false` | Enable OpenTelemetry tracing (see [OpenTelemetry](#opentelemetry)) |
| `OTEL_SERVICE_NAME` | `agentflow-api` | Service name reported in all traces |
| `OTEL_EXPORTER_OTLP_ENDPOINT` | `None` | OTLP collector URL — omit to print spans to console |
| `OTEL_LEVEL` | `standard` | Span detail level: `spans` \| `standard` \| `full` |
| `ORIGINS` | `*` | CORS allowed origins — restrict in production |

---

## OpenTelemetry

10xGraph has first-class OpenTelemetry support at two independent layers. You can use either or both.

```mermaid
flowchart TB
  subgraph "API layer  (FastAPI + HTTP)"
    FI[FastAPIInstrumentor\nHTTP spans — latency, status, route]
  end
  subgraph "Graph layer  (OtelPublisher)"
    GS[tenxgraph.graph span]
    NS[tenxgraph.node span]
    LS[tenxgraph.llm span\ntoken counts, model, finish reason]
    TS[tenxgraph.tool span\ntool name, type]
    GS --> NS --> LS
    NS --> TS
  end
  FI -.->|parent| GS
```

### API layer — automatic when `OTEL_ENABLED=true`

Setting `OTEL_ENABLED=true` in your environment is all that's required. The API server automatically:

- Creates a `TracerProvider` with your `OTEL_SERVICE_NAME`
- Instruments the FastAPI app with `FastAPIInstrumentor` (HTTP-level spans)
- Wires `OtelPublisher` into the graph so every LLM call, tool call, and node transition becomes a child span
- Exports via OTLP when `OTEL_EXPORTER_OTLP_ENDPOINT` is set; falls back to console output in non-production

```bash
OTEL_ENABLED=true
OTEL_SERVICE_NAME=my-agent
OTEL_EXPORTER_OTLP_ENDPOINT=http://collector:4317   # omit to print spans to console
OTEL_LEVEL=standard                                  # spans | standard | full
```

No code changes are needed. The SDK does not need to be configured separately — the API configures `OtelPublisher` automatically and merges it with any existing publisher (such as `RedisPublisher`) without replacing it.

### Graph layer — `OtelPublisher` and `ObservabilityLevel`

When running the graph directly (without `agentflow api`), pass `OtelPublisher` to `StateGraph` at init time:

```python
from tenxgraph.core.graph import StateGraph
from tenxgraph.runtime.publisher import OtelPublisher
from tenxgraph.runtime.publisher.otel_publisher import ObservabilityLevel

graph = StateGraph(publisher=OtelPublisher(level=ObservabilityLevel.STANDARD))
# ... add nodes and edges ...
compiled = graph.compile()
```

`ObservabilityLevel` controls how much data is emitted as span attributes:

| Level | What it includes |
|---|---|
| `STANDARD` | Token counts, model name, request params, finish reason *(default)* |
| `FULL` | All of STANDARD + prompt messages, completions, tool I/O — may contain PII |

With an explicit `TracerProvider` (e.g. to export to Jaeger or Honeycomb):

```python
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry import trace

from tenxgraph.core.graph import StateGraph
from tenxgraph.runtime.publisher import OtelPublisher
from tenxgraph.runtime.publisher.otel_publisher import ObservabilityLevel

provider = TracerProvider()
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter(endpoint="http://collector:4317")))
trace.set_tracer_provider(provider)

graph = StateGraph(publisher=OtelPublisher(level=ObservabilityLevel.FULL))
# ... add nodes and edges ...
compiled = graph.compile()
```

### Span hierarchy

Every graph run produces a consistent span tree:

```
tenxgraph.graph          ← one per ainvoke / astream call
  tenxgraph.node         ← one per node execution (e.g. "MAIN", "TOOL")
    tenxgraph.llm        ← one per LLM call (tokens, model, finish reason)
    tenxgraph.tool       ← one per tool call (name, type: local | mcp)
```

The `tenxgraph.graph` span carries `thread_id` as `session.id` so tools like Langfuse automatically group multi-turn conversations.

**Install:**

```bash
pip install "10xgraph[otel]"           # graph-level spans (OtelPublisher)
pip install "10xscale-agentflow-cli[otel]"       # API layer (FastAPIInstrumentor + OTLP exporter)
```

---

## What's next

| Page | What it covers |
|---|---|
| [Connecting Clients](/docs/concepts/connecting-clients) | TypeScript SDK, streaming, remote tools |
| [Memory](/docs/concepts/memory) | `PgCheckpointer`, Redis cache, long-term vector store |
| [Extensibility](/docs/concepts/extensibility) | `BaseAuth`, `AuthorizationBackend`, `BasePublisher` and all other ABCs |
| [Quality & Observability](/docs/qa) | `GraphLifecycleHook` with OpenTelemetry, evaluation, testing |
