Configure Rate Limiting

In shortHow to enable and configure the built-in sliding-window rate limiter in the 10xGraph API.

  • 9 min read
  • 13 sections
  • Updated
  • v0.10.0
  • Markdown

Rate limiting protects your API from overuse and abuse by enforcing a ceiling on how many requests a client can make within a time window. Without it, a single user or malicious actor can consume all your infrastructure resources and degrade service for everyone. The 10xGraph API server includes a built-in sliding-window rate limiter that you enable with a few lines of configuration.

By default, the 10xGraph API accepts unlimited requests. This guide shows how to enable the rate limiter, choose a backend suitable for your deployment, and handle edge cases like proxy headers and WebSocket connections.

How it works

The rate limiter is HTTP middleware that executes before your graph logic runs on every request. It maintains a count of requests per client (or globally) inside a rolling time window and returns 429 Too Many Requests when the limit is exceeded. The limit is never active until you add a rate_limit block to 10xgraph.json.

The limiter uses a sliding-window algorithm, not fixed time buckets. This means the window rolls with each request: a client with a limit of 100 requests per 60 seconds can send 100 requests in the first 30 seconds, then nothing until 60 seconds have passed, then 100 more. This prevents bursts that would succeed with a simple bucket-based counter.

Three backends are available: memory for development, Redis for production deployments, and custom if you need to store counters somewhere else. Your choice depends on how many workers you run and whether you prioritize simplicity or distributed enforcement.

Method 1: In-memory backend (development / single process)

The simplest setup stores counters in the process memory. It works with a single Uvicorn worker and requires no extra dependencies.

Update 10xgraph.json:

JSON
{
  "agent": "graph.react:app",
  "rate_limit": {
    "enabled": true,
    "backend": "memory",
    "requests": 100,
    "window": 60,
    "by": "ip",
    "exclude_paths": ["/ping", "/docs", "/redoc", "/openapi.json"]
  }
}

Each client IP can make up to 100 requests every 60 seconds. Health-check and documentation paths are excluded so monitoring and browser access are not affected.

Verify it works: Start the server and make requests from the same IP until you exceed the limit.

Terminal
10xgraph api

In another terminal, test the limiter:

Terminal
# First 100 requests succeed
for i in {1..100}; do
  curl -s http://localhost:8000/v1/graph/invoke \
    -H "Content-Type: application/json" \
    -d '{"messages": [{"role": "user", "content": [{"type": "text", "text": "test"}]}]}' > /dev/null
done

# Next request is rate-limited
curl -i http://localhost:8000/v1/graph/invoke \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": [{"type": "text", "text": "test"}]}]}'
# Returns: 429 Too Many Requests

Method 2: Redis backend (production / multi-process)

When you run multiple Uvicorn workers, containers, or servers, each process would have its own in-memory counter. Use Redis to store counters centrally so the limit is enforced across the whole deployment.

Install the Redis extra

Terminal
pip install "10xgraph-api[redis]"

Update 10xgraph.json

JSON
{
  "agent": "graph.react:app",
  "rate_limit": {
    "enabled": true,
    "backend": "redis",
    "requests": 1000,
    "window": 60,
    "by": "ip",
    "trusted_proxy_headers": true,
    "exclude_paths": ["/ping", "/metrics", "/docs", "/redoc", "/openapi.json"],
    "redis": {
      "url": "${RATE_LIMIT_REDIS_URL}",
      "prefix": "10xgraph:rate-limit"
    },
    "fail_open": true
  }
}

Set the environment variable

Terminal
# .env
RATE_LIMIT_REDIS_URL=redis://localhost:6379/0

The ${RATE_LIMIT_REDIS_URL} placeholder is expanded from the environment at startup. Never commit Redis credentials into 10xgraph.json.

How it works: The Redis backend uses a Lua script with sorted sets to atomically check and record a request in one operation. This prevents concurrent requests from slipping past the configured limit. Each request adds a timestamp to a sorted set; older entries are automatically cleaned up by the script, and Redis evicts idle keys using TTL.

Verify it works: Set up Redis, update your config, restart the server, and load-test from multiple processes:

Terminal
# Start Redis if needed
redis-server

# Run the server
10xgraph api

# In another terminal, hammer it with concurrent requests
ab -n 2000 -c 50 http://localhost:8000/v1/graph/invoke
# Some requests will return 429

Method 3: Custom backend

Implement BaseRateLimitBackend when you want to store counters somewhere else: a SQL database, a distributed cache, or an external rate-limit service.

Python
# graph/rate_limit.py
from tenxgraph_api.src.app.core.middleware.rate_limit import (
    BaseRateLimitBackend,
    RateLimitDecision,
)

class MyRateLimitBackend(BaseRateLimitBackend):
    async def check(self, key: str, *, limit: int, window: int) -> RateLimitDecision:
        # Replace with your real logic
        # key: the client identifier (IP, user, or "global")
        # limit: the configured request limit
        # window: the time window in seconds
        allowed = True
        remaining = limit - 1
        reset_after = window
        return RateLimitDecision(
            allowed=allowed,
            remaining=remaining,
            reset_after=reset_after,
        )

    async def close(self) -> None:
        # Release any resources (connections, file handles)
        return None

Register it in your InjectQ container and set backend to "custom":

Python
# graph/__init__.py
from injectq import InjectQ
from tenxgraph_api.src.app.core.middleware.rate_limit import BaseRateLimitBackend

from graph.rate_limit import MyRateLimitBackend

container = InjectQ()
container.bind_instance(BaseRateLimitBackend, MyRateLimitBackend())

Bind the instance under BaseRateLimitBackend: that is the key the server looks up. If backend is "custom" and nothing is bound there, startup fails.

Then in 10xgraph.json:

JSON
{
  "injectq": "graph:container",
  "rate_limit": {
    "enabled": true,
    "backend": "custom",
    "requests": 200,
    "window": 60,
    "by": "ip"
  }
}

Identity modes

JSON
{
  "rate_limit": { "requests": 100, "window": 60, "by": "ip" }
}

Limits requests by client IP address. This is the default and works well when you do not have authentication. Anonymous users, scripts, and browser clients are each limited individually.

JSON
{
  "rate_limit": { "requests": 100, "window": 60, "by": "user" }
}

Limits requests by authenticated user_id. This is the right choice as soon as you have auth: limiting purely by IP gives a single user roaming between addresses (mobile, corporate proxy) an effectively unlimited budget, while a NAT’d office sharing one address gets throttled as though it were one caller.

Requests with no authenticated user fall back to an ip:<address> bucket, so anonymous traffic is still limited per caller rather than sharing one bucket any single client could exhaust for everyone.

Global (one shared quota for all clients)

JSON
{
  "rate_limit": { "requests": 5000, "window": 60, "by": "global" }
}

Enforces a single limit for the entire service. Use this when you want to control total throughput rather than per-client fairness. This is rare; per-IP or per-user is almost always better.

WebSocket handshakes count too

Rate limiting is HTTP middleware, and Starlette runs middleware only for HTTP scopes. WebSocket handshakes would otherwise bypass the limiter entirely. The API server therefore re-applies the rate-limit check at WebSocket handshake time, against the same backend and the same bucket as REST requests.

Opening a socket costs one request from the client’s quota. When the quota is exhausted the handshake is refused before accept() with WebSocket close code 1013 (Try Again Later), not with an HTTP 429.

Each graph run started over /v1/graph/ws also costs one request from the same bucket. A run over the limit is not executed: the server sends an error chunk with retry_after_seconds and keeps the socket open.

Budget for this when sizing limits for a streaming client. A browser that reconnects on every network blip spends a request each time, and every run on a long-lived socket spends one more.

Concurrent socket count is capped separately by websocket.max_connections (and websocket.max_connections_per_user), which use the same close code. See 10xgraph.json configuration.

Behind a reverse proxy

If your API runs behind nginx, a load balancer, or a cloud gateway, the real client IP is forwarded in the X-Forwarded-For header. Set trusted_proxy_headers: true to use that header instead of the direct connection IP.

JSON
{
  "rate_limit": {
    "backend": "redis",
    "requests": 1000,
    "window": 60,
    "by": "ip",
    "trusted_proxy_headers": true,
    "trusted_proxy_hops": 1
  }
}

Set trusted_proxy_hops to match your topology

X-Forwarded-For is a list that each proxy appends to. Whatever the caller sent arrives at the left; only the entries your own proxies appended, on the right, are trustworthy.

trusted_proxy_hops is how many entries, counted from the right, your own infrastructure appended. It defaults to 1, which is correct for a single proxy in front of the app.

Topology trusted_proxy_hops
One nginx or one load balancer 1
CDN in front of a load balancer, both yours 2
No proxy at all leave trusted_proxy_headers off

To honor X-Forwarded-For only for connections that come from your proxies, also set trusted_proxies to a list of IPs or CIDR ranges. A client that reaches the app directly is then keyed by its own address.

Excluding paths

Add monitoring, health-check, and documentation paths to exclude_paths so they never count against the rate limit:

JSON
{
  "rate_limit": {
    "exclude_paths": ["/ping", "/metrics", "/docs", "/redoc", "/openapi.json"]
  }
}

Health checks (/ping) are excluded automatically and do not need to be listed. Other status endpoints, readiness checks, and documentation paths should be added here so your monitoring and developers do not trigger limits.

Disabling rate limiting

Remove the rate_limit block (or set it to null) to disable the middleware entirely:

JSON
{
  "agent": "graph.react:app",
  "rate_limit": null
}

What clients see when limited

When the limit is exceeded, the API returns 429 Too Many Requests:

JSON
{
  "error": {
    "code": "RATE_LIMIT_EXCEEDED",
    "message": "Too many requests. Limit: 100 per 60s. Retry after 12s.",
    "limit": 100,
    "window_seconds": 60,
    "retry_after_seconds": 12
  },
  "metadata": {
    "request_id": "abc123",
    "status": "error"
  }
}

Every response also includes these headers so clients can track their quota:

Header Description
X-RateLimit-Limit Configured request limit
X-RateLimit-Remaining Requests remaining in the current window
X-RateLimit-Reset Unix timestamp for the window reset
X-RateLimit-Reset-After Seconds until the window resets
Retry-After Seconds to wait before retrying (only on 429)

Clients can check X-RateLimit-Remaining before sending the next request to avoid hitting the limit, and use Retry-After to back off intelligently.

Troubleshooting

I still get unlimited requests even though I configured a limit.

Verify that 10xgraph.json has a rate_limit block with "enabled" not set to false and that you restarted the server after editing the file. Check the startup logs for rate-limit configuration messages. If the backend is Redis, ensure the connection is working: a backend error with fail_open: true allows all requests.

Different workers seem to have different limits.

You are using the memory backend with multiple workers. Each process maintains its own counters, so the real limit is requests × workers. Switch to Redis backend for consistent enforcement across workers.

The proxy’s real IP is not being recognized.

Verify trusted_proxy_headers: true is set and trusted_proxy_hops matches your topology (1 for single proxy, 2 for CDN + load balancer). Check the server logs for warnings about mismatched hop counts. If the header has fewer entries than the hop count, it is ignored and the direct peer address is used instead.

Rate-limit headers are not appearing in responses.

They should appear on every response, including successful ones. If missing, check that rate limiting is enabled and the response is not being cached by a proxy. Middleware rate-limit headers are always fresh.

See also

Frequently asked questions

Do I need rate limiting in development?
In development, you can run without limits. In any production environment, rate limiting protects your infrastructure from both accidental overuse and abuse.
How much does Redis add to the setup?
The limiter reuses a Redis client already bound in the InjectQ container (under `redis`, `redis_client` or `Redis`). Otherwise it needs `rate_limit.redis.url` and opens its own connection.
What happens if Redis is unreachable?
With `fail_open: true` (the default), requests pass through; with `fail_open: false`, requests are denied. Choose based on whether uptime matters more than protection.
Last updated for v0.10.0Edit this page on GitHubReport an issue