# Rate Limiting

> Complete reference for the rate_limit block in 10xgraph.json: the backends, limits, and every option for throttling requests to the 10xGraph API server.

Source: https://10xgraph.com/docs/reference/api-cli/rate-limiting
Last updated: 2026-09-29

The `rate_limit` block in `10xgraph.json` activates 10xGraph's built-in sliding-window
rate limiter. The limiter is disabled by default — remove the block or set it to `null`
to turn it off.

## Configuration fields

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `enabled` | boolean | `true` | Enables the middleware when the `rate_limit` block exists. Set to `false` to temporarily disable without removing the block. |
| `backend` | string | `"memory"` | Counter storage. `"memory"`, `"redis"`, or `"custom"`. |
| `requests` | integer | `100` | Maximum requests allowed within each window. |
| `window` | integer | `60` | Window size in seconds. |
| `by` | string | `"ip"` | Bucket key. `"ip"`, `"user"`, or `"global"`. See [Bucket keys](#bucket-keys). |
| `exclude_paths` | string array | `[]` | Request paths that bypass rate limiting entirely. The `/ping` health check is always exempt, so a probe never gets `429`, even when the backend is down and `fail_open` is `false`. |
| `trusted_proxy_headers` | boolean | `false` | Use `X-Forwarded-For` to resolve the client IP. Only enable behind a proxy you control. |
| `trusted_proxy_hops` | integer | `1` | How many proxies of your own sit in front of the app. See [Proxy hops](#proxy-hops). Must be `>= 1`. |
| `trusted_proxies` | string array | `[]` | IPs or CIDR ranges your proxies connect from. When set, `X-Forwarded-For` is honoured only for requests whose peer address is in one of them. See [Proxy hops](#proxy-hops). An invalid network raises a `ValueError`. |
| `redis` | object or string | `null` | Redis connection for the `"redis"` backend, as `{"url": ..., "prefix": ...}` or the URL as a bare string (which keeps the default prefix). |
| `redis.url` | string | `null` | Redis connection URL. Required for the `"redis"` backend unless a Redis client is already bound in InjectQ. Supports `$ENV_VAR` and `${ENV_VAR}` expansion; an unset variable stops the server from starting. |
| `redis.prefix` | string | `"agentflow:rate-limit"` | Key prefix used for all Redis entries. |
| `fail_open` | boolean | `true` | When `true`, requests are allowed if the Redis backend is unreachable. When `false`, they are denied. Only applies to the `"redis"` backend. |

Invalid values are rejected at config load: `by` outside `ip`/`user`/`global`, `backend` outside `memory`/`redis`/`custom`, a non-positive `requests` or `window`, `trusted_proxy_hops` below `1`, or an entry in `trusted_proxies` that is not an IP or CIDR range all raise a `ValueError` and stop the server from starting.

## Bucket keys

`by` decides which bucket a request is counted against.

| `by` | Key | Notes |
| --- | --- | --- |
| `"ip"` | the resolved client address | The default. One bucket per client address. |
| `"user"` | `user:<user_id>` | One bucket per authenticated user. Falls back to `ip:<address>` when there is no authenticated user, so anonymous traffic is still limited per caller rather than sharing one bucket everybody can exhaust. |
| `"global"` | `__global__` | One bucket for the whole service. |

`"user"` is what you want once auth is enabled. Limiting purely by IP gives a single user roaming between addresses an effectively unlimited budget, while a NAT'd office sharing one address gets throttled as though it were one caller.

## Proxy hops

`X-Forwarded-For` is a list that each proxy **appends** to. Whatever the caller sent arrives at the left of the list; only the entries your own proxies appended, on the right, are trustworthy. Reading the leftmost entry would let a caller send a different value on every request, land in a fresh bucket each time, and never be limited at all.

`trusted_proxy_hops` is how many entries, counted from the **right**, your own infrastructure appended. With the default of `1` (one proxy in front of the app) the last entry is the address that proxy actually observed. If the header carries fewer entries than the configured hop count, the header is ignored entirely and the peer address is used, with a warning.

`trusted_proxy_hops` only has an effect when `trusted_proxy_headers` is `true`.

The hop count assumes every request comes through your proxy. A client that can reach the app
directly would otherwise be free to send its own `X-Forwarded-For`. Set `trusted_proxies` to the
networks your proxies connect from, and the header is honoured only for requests whose peer
address is in one of them; anyone else is keyed by their peer address.

```json
"rate_limit": {
  "trusted_proxy_headers": true,
  "trusted_proxy_hops": 1,
  "trusted_proxies": ["10.0.0.0/8"]
}
```

## WebSocket handshakes

Rate limiting is HTTP middleware, and Starlette runs middleware for HTTP scopes only, so WebSocket handshakes would otherwise bypass it. `WS /v1/graph/ws` and `WS /v1/graph/live` therefore re-apply the check at the handshake, using the **same backend and the same bucket** as REST requests. Opening a socket counts exactly like any other request; exceeding the limit refuses the handshake with WebSocket close code `1013` (Try Again Later) before `accept()`.

The separate [`websocket.max_connections`](/docs/reference/api-cli/configuration#websocket-ag_ui-and-observability) cap is enforced at the same point and uses the same close code.

## Minimal example

```json
{
  "agent": "graph.react:app",
  "rate_limit": {
    "enabled": true,
    "backend": "memory",
    "requests": 100,
    "window": 60,
    "by": "ip",
    "exclude_paths": ["/ping", "/docs", "/redoc", "/openapi.json"]
  }
}
```

## Full Redis example

```json
{
  "agent": "graph.react:app",
  "rate_limit": {
    "enabled": true,
    "backend": "redis",
    "requests": 1000,
    "window": 60,
    "by": "ip",
    "trusted_proxy_headers": true,
    "exclude_paths": ["/ping", "/metrics", "/docs", "/redoc", "/openapi.json"],
    "redis": {
      "url": "${RATE_LIMIT_REDIS_URL}",
      "prefix": "agentflow:rate-limit"
    },
    "fail_open": true
  }
}
```

```bash
# .env
RATE_LIMIT_REDIS_URL=redis://localhost:6379/0
```

Install the Redis extra before using `"backend": "redis"`:

```bash
pip install "10xscale-agentflow-cli[redis]"
```

## Backend comparison

| Backend | When to use |
| --- | --- |
| `memory` | Local development, tests, demos, single-process services |
| `redis` | Production: Gunicorn/Uvicorn with multiple workers, Docker/Kubernetes |
| `custom` | Custom storage, external quota services, non-standard enforcement |

> **The memory backend counts per process**
>
> Enabling the `memory` backend logs a startup warning. It keeps counters in process memory, so with N workers the effective limit is `requests x N`, and every counter resets when a worker restarts. Use the `redis` backend for any multi-worker deployment.

## Response headers

Every response includes rate-limit headers:

| Header | Description |
| --- | --- |
| `X-RateLimit-Limit` | Configured request limit |
| `X-RateLimit-Remaining` | Requests remaining in the current window |
| `X-RateLimit-Reset` | Unix timestamp for the window reset estimate |
| `X-RateLimit-Reset-After` | Seconds until the window resets |
| `Retry-After` | Present on `429` responses only |

## 429 response body

```json
{
  "error": {
    "code": "RATE_LIMIT_EXCEEDED",
    "message": "Too many requests. Limit: 100 per 60s. Retry after 12s.",
    "limit": 100,
    "window_seconds": 60,
    "retry_after_seconds": 12
  },
  "metadata": {
    "request_id": "request-id",
    "status": "error"
  }
}
```

## Custom backend interface

```python
from agentflow_cli.src.app.core.middleware.rate_limit import (
    BaseRateLimitBackend,
    RateLimitDecision,
)

class MyRateLimitBackend(BaseRateLimitBackend):
    async def check(self, key: str, *, limit: int, window: int) -> RateLimitDecision:
        allowed = True
        remaining = limit - 1
        reset_after = window
        return RateLimitDecision(
            allowed=allowed,
            remaining=remaining,
            reset_after=reset_after,
        )

    async def close(self) -> None:
        return None
```

Set `"backend": "custom"` in `10xgraph.json` and bind the instance through InjectQ.

## See also

- [Configure Rate Limiting](/docs/how-to/api-cli/configure-rate-limiting) — step-by-step setup guide
- [10xgraph.json configuration](/docs/reference/api-cli/configuration)
- [Environment variables](/docs/reference/api-cli/environment)
