WebSocket streaming and realtime
In shortBidirectional WebSocket endpoints for turn-based agent streaming and realtime audio.
- 9 min read
- 11 sections
- Updated
- v0.10.0
- Markdown
10xGraph exposes two WebSocket endpoints: /v1/graph/ws for turn-based agent interactions and /v1/graph/live for realtime audio conversations. WebSockets keep a persistent connection open so the server can stream responses token-by-token, the client can resume from checkpoints, and (for realtime agents) both sides can exchange audio in near-real-time. Both endpoints support the same authentication, rate limiting, and concurrency controls as REST APIs.
When to use WebSocket
Use WebSocket when the client needs true bidirectional streaming or low-latency two-way communication. For simple invoke-and-wait workflows, REST’s /v1/graph/invoke is simpler. For streaming where the client reads until done, REST’s /v1/graph/stream (NDJSON) is also sufficient and works in more environments (proxies, browsers without WebSocket support).
WebSocket shines for:
- Resume workflows: the client detects a remote tool call in the stream, executes it, and resumes the same run without losing context.
- Interactive agents: the client sends new instructions mid-run without waiting for the current response.
- Realtime audio: the client streams audio input while the model speaks output, with near-zero latency.
- Long-lived connections: applications that keep one socket open for many sequential runs over hours.
Authentication
Both endpoints use the same credential methods as REST APIs. The bearer token can come from:
AuthorizationHTTP header (standard):Authorization: Bearer <token>10xgraph-bearerSec-WebSocket-Protocol (recommended for browsers): browsers cannot set anAuthorizationheader on a WebSocket, and this keeps the token out of the URL?token=query parameter (fallback): the token is visible in logs and browser history, so avoid in production
At the WebSocket handshake, the server verifies the token. Before every graph run on /v1/graph/ws (the client can send several messages on one socket), the token is verified again: if it expires or is revoked between runs, the server sends an error chunk and closes the socket with code 1008.
Setting the subprotocol (if using 10xgraph-bearer): offer two subprotocols, the sentinel first and the raw token second.
const socket = new WebSocket(
"wss://api.example.com/v1/graph/ws",
["10xgraph-bearer", TOKEN]
);The server accepts both 10xgraph-bearer and the legacy agentflow-bearer (deprecated until 2.0), and echoes the sentinel back as the accepted subprotocol.
Turn-based streaming: /v1/graph/ws
This endpoint streams graph execution as the agent runs step-by-step. The client sends one message, the server streams back chunks, and when done, the client can send another message or close the connection. It is suitable for turn-based agents (ReactAgent, PlanActReflect, etc.) where each user message is a complete turn.
Protocol
Fresh run (start a new thread or run on an existing one):
Client → Server:
{
"invoke_type": "fresh",
"messages": [
{"role": "user", "content": [{"type": "text", "text": "What is 2 + 2?"}]}
],
"config": {
"thread_id": "optional-id-or-omit-for-new-thread"
}
}
Server → Client (StreamChunk JSON, the same format as POST /v1/graph/stream):
{"event": "updates", "data": {"status": "invoking_graph", ...}}
{"event": "message", "message": {...}}
...
{"event": "updates", "data": {"status": "done"}}The final {"event": "updates", "data": {"status": "done"}} chunk is sent by the server after every run on the socket. The socket stays open for the next frame.
Resume after remote tool call:
When a run ends on a remote (client-executed) tool call, the client executes the tool (for example in the browser) and sends the result back. The server does not inspect chunks for tool calls; detecting them is the client library’s job. The resume frame needs tool_result and config.thread_id:
Client → Server:
{
"invoke_type": "resume",
"tool_result": [
{"role": "tool", "content": [{"type": "text", "text": "result here"}]}
],
"config": {
"thread_id": "same-thread-from-checkpoint"
}
}
Server → Client:
(same format: StreamChunk lines, ending with updates/status:done)Thread IDs
- Omit
thread_idor send"new"to create a fresh thread each time. - Send an existing
thread_idto resume that thread’s conversation. - A blank or missing
thread_idalso means “create a new one”, so the server generates an ID. resumeruns requireconfig.thread_id.- When auth is enabled, the thread is owner-checked before each run. An unauthorized thread gets an
errorchunk withNot authorized to stream thread ...and the socket stays open.
Rate limiting and concurrency
Every WebSocket handshake and every graph run on the socket counts against the global rate limit: the same bucket as REST requests, when rate_limit is configured.
- A handshake over the limit is refused before
accept()with close code 1013. - A run over the limit gets an
errorchunk whosedatahasreasonandretry_after_seconds. The run is skipped and the socket stays open; the check runs again on the next frame.
Connection caps apply at the handshake only: websocket.max_connections (per process) and websocket.max_connections_per_user. Over either cap, the handshake is refused with close code 1013.
Realtime audio: /v1/graph/live
This endpoint bridges a WebSocket to a realtime agent (a graph rooted at a live agent, driven through CompiledGraph.arealtime()), enabling true two-way audio conversation. The client and server exchange audio frames and control messages in real-time, with minimal latency.
Realtime is only for agents built for it: connecting to /v1/graph/ws with a realtime agent closes immediately with code 1008 (use /v1/graph/live instead). The reverse also holds: a turn-based graph on /v1/graph/live is closed with 1008.
Protocol
Init frame (JSON, the first frame sent): a flat object. All keys are optional; the keys the server reads are thread_id, model, voice, modalities, vad, system_prompt and tools_tags.
Client → Server:
{
"model": "gemini-2.5-flash-live",
"voice": "Puck",
"modalities": ["AUDIO"],
"vad": {
"enabled": true,
"prefix_padding_ms": 300,
"silence_duration_ms": 1000
},
"thread_id": "optional-thread-id"
}The model override is honored only if it is listed in websocket.realtime_models; otherwise it is ignored and the agent’s own model is used. A bad value (for example an invalid modalities entry) returns a fatal error event with code invalid_config.
Downstream (Server → Client):
- Binary frames: PCM16 audio deltas from the model
- JSON text frames: every other event (transcripts, turn complete, interrupted, tool calls, errors)
Upstream (Client → Server):
- Binary frames: PCM16 audio input from the user/environment
- JSON text frames:
{"type": "activity_start"},{"type": "activity_end"},{"type": "text", "text": "..."},{"type": "close"}
Refer to the realtime audio guide for details and examples.
Configuration
WebSocket behavior is controlled by keys in the websocket object within 10xgraph.json:
| Key | Type | Default | Purpose |
|---|---|---|---|
max_connections |
integer, null | 1000 | Hard cap on concurrent WebSocket connections (both /v1/graph/ws and /v1/graph/live) across this process. Set to 0 or null for unlimited. |
max_connections_per_user |
integer, null | 10 | Cap on how many WebSocket connections one verified user can hold (by user_id). Set to 0 or null for unlimited. |
realtime_models |
array of strings | [] |
Models a /v1/graph/live client is allowed to request (for example ["gemini-2.5-flash-live"]). Empty means clients cannot override the agent’s model. |
Example:
{
"agent": "graph:app",
"websocket": {
"max_connections": 5000,
"max_connections_per_user": 50,
"realtime_models": ["gemini-2.5-flash-live"]
}
}In a multi-worker deployment (Gunicorn + Uvicorn), these limits are per process, not global. Size them accordingly: if you run 4 workers with max_connections: 1000, the server can hold 4000 concurrent sockets total.
Close codes
WebSocket close codes follow RFC 6455. The server sends these in specific situations:
| Code | Meaning | Retry? |
|---|---|---|
| 1000 | Normal closure (client initiated or clean shutdown) | No |
| 1003 | Unsupported data: the /v1/graph/live init frame is not a JSON object or could not be parsed |
No |
| 1008 | Policy violation: wrong endpoint for the agent type, token expired/revoked between runs, or auth/authorization failure at the handshake | No |
| 1011 | Unexpected server error during run | May retry after brief delay |
| 1013 | Try again later: rate limit or connection cap exceeded at the handshake | Yes, after a delay |
On /v1/graph/ws, an error chunk is sent before closing for the wrong-agent-type and expired-token cases. Handshake rejections (auth failure, rate limit, connection cap) close before the socket is accepted, so no chunk is sent.
Proxy settings
If the API server is behind a reverse proxy (Nginx, HAProxy, CloudFlare, etc.), the proxy must be configured to preserve the WebSocket connection and pass through the upgrade headers. Additionally, when authentication is enabled and the proxy forwards client IP via X-Forwarded-For, the server must be told to trust that header.
Proxy compatibility
- The proxy must support HTTP/1.1 upgrade requests (WebSocket handshake).
- Connection and Upgrade headers must be forwarded as-is.
- Most reverse proxies support this by default; confirm in their documentation.
Nginx example (snippet):
location /v1/graph/ws {
proxy_pass http://backend:8000;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_set_header Host $host;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_read_timeout 86400s; # Allow long-lived sockets
}Client IP detection
If you use the trusted_proxy_headers setting in rate_limit, the server reads the client IP from X-Forwarded-For to bucket rate limits correctly:
{
"agent": "graph:app",
"rate_limit": {
"enabled": true,
"by": "ip",
"trusted_proxy_headers": true,
"trusted_proxy_hops": 1
}
}trusted_proxy_headers: whether to honorX-Forwarded-Forat all. Settrueonly if there is a proxy you control.trusted_proxy_hops: how many proxy layers to count back from the right of theX-Forwarded-Forlist. Set to 1 if you have one proxy, 2 for two proxies in series, etc.
Without this, clients behind a proxy appear to have the proxy’s IP and all land in the same rate-limit bucket.
Example: TypeScript client
Here’s a minimal example using the 10xGraph TypeScript client:
import { TenxGraphClient, Message } from "10xgraph-client";
const client = new TenxGraphClient({
baseUrl: "http://localhost:8000",
authToken: "your-jwt-token",
});
// wsStream() has the same signature as stream(), but runs over /v1/graph/ws
const stream = client.wsStream([Message.text_message("What is 2 + 2?")], {
config: { thread_id: "my-thread" },
});
for await (const chunk of stream) {
if (chunk.event === "message") {
console.log("Message:", chunk.message);
}
}wsStream() handles remote tool calls for you: it runs registered client-side tools and sends the resume frame on the same socket. Use client.stream() for the HTTP equivalent. See Client streaming.
Example: Raw WebSocket in JavaScript
If you prefer to manage the WebSocket directly:
const token = "your-jwt-token";
const socket = new WebSocket("wss://api.example.com/v1/graph/ws", ["10xgraph-bearer", token]);
socket.onopen = () => {
console.log("Connected");
socket.send(JSON.stringify({
invoke_type: "fresh",
messages: [
{ role: "user", content: [{ type: "text", text: "Hello" }] }
]
}));
};
socket.onmessage = (event) => {
const chunk = JSON.parse(event.data);
console.log("Chunk:", chunk.event, chunk.data);
if (chunk.event === "updates" && chunk.data.status === "done") {
// Run finished
socket.close(1000);
}
};
socket.onerror = (event) => {
console.error("WebSocket error:", event);
};
socket.onclose = (event) => {
console.log(`Closed: code=${event.code}, reason=${event.reason}`);
};Troubleshooting
Connection refused or 403 forbidden
- Check that your token is valid and not expired. Verify it is passed via the
Authorizationheader, the10xgraph-bearersubprotocol, or?token=query. A rejected handshake closes with code 1008. - Confirm authentication is enabled and configured in
10xgraph.json("auth": "jwt"or custom).
Socket closes with code 1008
- If the message is “This graph is a live agent…”, you’re using
/v1/graph/wson a realtime agent. Switch to/v1/graph/live. - If the message is “Session expired…”, your token has expired. Reconnect with a fresh token.
Socket closes with code 1013
- Rate limit or connection cap exceeded at the handshake. Wait and reconnect.
- Check
websocket.max_connectionsandwebsocket.max_connections_per_userin your config.
Proxy forwarding issues
- Confirm the proxy forwards the
UpgradeandConnectionheaders. Most proxies do by default. - If behind a proxy, set
trusted_proxy_headers: truein rate_limit config so the server reads the real client IP fromX-Forwarded-For. - Increase proxy timeouts: WebSocket connections can be long-lived, so set read/write timeouts to several minutes or higher.
Audio glitches or delays on /v1/graph/live
- Increase
vad.prefix_padding_msto keep more audio from before speech is detected. - Tune
start_sensitivityandend_sensitivityin thevadobject if speech is missed or background noise triggers it. - Check network latency; realtime is sensitive to round-trip time.
Related pages
- Invoke and stream over REST covers HTTP POST endpoints for the same operations.
- Realtime audio guide provides complete examples and configuration for audio agents.
- Client streaming documents the TypeScript client’s streaming API.
- Authentication explains JWT setup and token validation.