Handle files and multimodal input
In shortEnable clients to upload images, audio, and documents to your agent via the file API.
- 7 min read
- 10 sections
- Updated
- v0.10.0
- Markdown
The 10xGraph API server accepts file uploads from clients and injects the media into your graph’s execution. Upload once, reference by ID, and rely on the server to resolve the reference before model calls. This works the same way across REST (/v1/graph/invoke, POST /v1/graph/stream) and WebSocket (WS /v1/graph/ws), with built-in ownership enforcement so users cannot reference each other’s files.
How it works
The file pipeline flows in four steps:
- Upload: Client calls
POST /v1/files/uploadwith binary data; the server stores it and returns afile_id. - Reference: Client includes the
file_idin a content block (image, audio, or document) on a message sent to the graph. - Rewrite: Before execution, the server resolves the reference:
- Images and audio → converted to
graph://media/{file_id}URLs, resolved by model adapters at call time. - Documents → replaced with extracted plain text if extraction was cached at upload; otherwise the block is rewritten to a
graph://media/{file_id}reference like images and audio.
- Images and audio → converted to
- Execute: Your graph receives the rewritten messages with fully resolved media.
This rewrite happens in the same input-preparation step for all three endpoints, so behavior is identical whether the client sends over REST, WebSocket, or the server-to-server AG-UI protocol. Ownership is checked at every step: when user B downloads or inspects a file uploaded by user A, the API returns 404, never 403. When user B references it in a message, the request fails with a “File not found” error (422 on invoke). Either way the API never confirms that a foreign file_id exists.
The rewrite is a no-op when no media service is configured, so development without file uploads requires no extra setup.
Note: WebSocket live streaming (WS /v1/graph/live) is audio-only and does not use the file upload API; see its documentation.
Upload a file
To upload a file, send a multipart POST to /v1/files/upload with the binary data:
curl -X POST http://127.0.0.1:8000/v1/files/upload \
-H "Authorization: Bearer $TOKEN" \
-F "[email protected]"The server responds with metadata and a unique file_id:
{
"data": {
"file_id": "a8f2k9x1m5p3",
"mime_type": "image/png",
"size_bytes": 184320,
"filename": "invoice.png",
"extracted_text": null,
"url": "/v1/files/a8f2k9x1m5p3",
"direct_url": null,
"direct_url_expires_at": null
},
"metadata": {"request_id": "...", "timestamp": "...", "message": "OK"}
}Permissions and ownership
The uploading user is automatically recorded as the file’s owner. Every read operation, including retrieval, metadata queries, and message rewriting, checks ownership and rejects requests from other users (404 on the file routes, a “File not found” error on message references). This owner recording is immutable and survives as long as the file does.
The server reads uploads in 1 MiB chunks, so oversized files are rejected before the entire body is buffered in memory. Two HTTP status codes signal problems:
| Status | Meaning |
|---|---|
415 Unsupported Media Type |
The file’s MIME type is not in MEDIA_ALLOWED_CONTENT_TYPES. |
413 Payload Too Large |
The file exceeds MEDIA_MAX_SIZE_MB. The limit is checked during streaming, so rejection occurs before the full body is loaded. |
Content-type restrictions
By default, all MIME types are allowed. Before exposing the upload endpoint to untrusted callers, set MEDIA_ALLOWED_CONTENT_TYPES to restrict what can be uploaded:
# Exact types
MEDIA_ALLOWED_CONTENT_TYPES=image/png,image/jpeg,application/pdf
# Wildcard subtypes
MEDIA_ALLOWED_CONTENT_TYPES=image/*,application/pdf
# Any image type
MEDIA_ALLOWED_CONTENT_TYPES=image/*Matching is case-insensitive and ignores charset parameters in the Content-Type header. A rejected upload returns 415 Content type not allowed: <mime>.
Reference a file in a message
Once uploaded, use the file_id in a content block. Send the message to the graph with /v1/graph/invoke, /v1/graph/stream, or over WS /v1/graph/ws:
{
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "What is the total on this invoice?"
},
{
"type": "image",
"media": {
"kind": "file_id",
"file_id": "a8f2k9x1m5p3"
}
}
]
}
],
"config": {
"thread_id": "thread-123"
}
}The server rewrites this block to graph://media/a8f2k9x1m5p3 before execution. Your graph receives the rewritten reference, and the model adapter resolves it when needed.
If you resend a message that already contains a graph://media/ URL, it is left untouched, so retries are safe.
Block types
You can reference files in image, audio, or document blocks:
{
"type": "image",
"media": {"kind": "file_id", "file_id": "..."}
}{
"type": "audio",
"media": {"kind": "file_id", "file_id": "..."}
}{
"type": "document",
"media": {"kind": "file_id", "file_id": "..."}
}Document extraction and caching
Documents (PDFs, Word files, HTML, etc.) are handled differently from images and audio because most models expect text input, not binary files.
Extraction at upload time
When you upload a document and DOCUMENT_HANDLING=extract_text (the default), the server extracts plain text immediately and returns it in the extracted_text field of the upload response. This extraction is cached:
- In-process cache (per worker): immediate lookups for the same file.
- Checkpointer cache (if configured): shared across workers with a 24-hour TTL in the
media:extractionnamespace.
Extractable MIME types are:
application/pdf(PDF documents)application/msword(legacy Word .doc)application/vnd.openxmlformats-officedocument.wordprocessingml.document(Word .docx)text/html,text/xml,application/xml(markup)text/markdown,text/csv,application/json,text/plain(text formats)
Document extraction requires the media extra:
pip install "10xgraph-api[media]"Document handling modes
Three modes control what happens to document blocks when referenced:
| Mode | Behavior |
|---|---|
extract_text |
Extract text at upload and replace the document block with the cached text on reference. Default. |
pass_raw |
No extraction at upload; the document block stays a media reference for your graph or model adapter to handle. |
skip |
No extraction at upload; documents are not turned into text. |
Set the mode via the DOCUMENT_HANDLING environment variable.
Multi-worker extraction
If you run multiple workers without a checkpointer, each worker maintains its own in-process extraction cache. A document uploaded through worker A may be re-extracted on worker B, wasting CPU. Configure a checkpointer (PostgreSQL + Redis recommended for production) to share the extraction cache across workers.
Query the server’s media configuration
Clients can call GET /v1/config/multimodal to discover what the server accepts without hard-coding assumptions:
curl -H "Authorization: Bearer $TOKEN" \
http://127.0.0.1:8000/v1/config/multimodal{
"data": {
"media_storage_type": "local",
"media_max_size_mb": 25.0,
"document_handling": "extract_text"
},
"metadata": {"request_id": "...", "timestamp": "...", "message": "OK"}
}This requires the config:read permission. The response intentionally omits the content-type allowlist; treat HTTP 415 as the signal that a type is refused.
Storage backends
Configure where and how uploaded files are stored via environment variables. Three backends are available: in-memory (development only), local filesystem (single-server), and cloud (S3 or GCS, recommended for production).
Memory storage
MEDIA_STORAGE_TYPE=memoryAll files are held in RAM and lost on restart. Use only for local development and testing. A worker restart or crash loses all uploads.
Local filesystem storage
MEDIA_STORAGE_TYPE=local
MEDIA_STORAGE_PATH=./uploadsFiles are written to the configured directory. The default is ./uploads relative to the working directory. This backend assumes a shared filesystem across workers (e.g., NFS or attached storage), so each worker can read files uploaded by any other worker. Without shared storage, files uploaded to one worker are not visible to the others.
Cloud storage (AWS S3 or Google Cloud Storage)
MEDIA_STORAGE_TYPE=cloud
MEDIA_CLOUD_PROVIDER=aws
MEDIA_CLOUD_BUCKET=my-bucket
MEDIA_CLOUD_REGION=us-east-1
MEDIA_CLOUD_PREFIX=10xgraph-mediaFiles are uploaded to S3 or GCS, accessible to all workers instantly. This is the recommended setup for multi-worker deployments. The cloud backend is stateless and horizontally scalable.
AWS S3 credentials
Provide credentials one of three ways:
Via environment variables:
MEDIA_CLOUD_ACCESS_KEY_ID=AKIA...
MEDIA_CLOUD_SECRET_ACCESS_KEY=...Via temporary AWS session token:
MEDIA_CLOUD_ACCESS_KEY_ID=ASIA...
MEDIA_CLOUD_SECRET_ACCESS_KEY=...
MEDIA_CLOUD_SESSION_TOKEN=...Via IAM role (recommended): Omit credentials and rely on the instance’s attached IAM role.
Google Cloud credentials
Provide a service account JSON key:
MEDIA_CLOUD_PROVIDER=gcp
MEDIA_CLOUD_PROJECT_ID=my-project
MEDIA_CLOUD_CREDENTIALS_JSON='{"type": "service_account", ...}'Or set GOOGLE_APPLICATION_CREDENTIALS and omit MEDIA_CLOUD_CREDENTIALS_JSON.
Signed URLs and direct access
With a cloud storage backend, clients can request signed direct URLs to bypass the API server. With local or memory storage there is no direct URL, and url is the API path /v1/files/{file_id}. expires_at is a Unix timestamp in seconds.
curl -H "Authorization: Bearer $TOKEN" \
http://127.0.0.1:8000/v1/files/{file_id}/url{
"data": {
"url": "https://s3.amazonaws.com/bucket/10xgraph-media/abc123?X-Amz-Algorithm=...",
"file_id": "a8f2k9x1m5p3",
"expires_at": 1791470000,
"mime_type": "image/png"
},
"metadata": {"request_id": "...", "timestamp": "...", "message": "OK"}
}Signed URLs are useful for large files or direct browser downloads. Ownership is checked before issuing a URL, but the signed URL itself is a bearer link: anyone who holds it can download the file until it expires.
Signed URL configuration
| Variable | Default | Purpose |
|---|---|---|
MEDIA_SIGNED_URL_TTL_SECONDS |
3600 |
Lifetime of a signed URL (1 hour). |
MEDIA_SIGNED_URL_REFRESH_BUFFER_SECONDS |
60 |
Re-sign this many seconds before expiry instead of handing out a URL about to die. |
The server caches the signed URL and reuses it until fewer than MEDIA_SIGNED_URL_REFRESH_BUFFER_SECONDS remain, then signs a new one. With the defaults, a client always receives a URL with at least 60 seconds of validity left.
Production deployment checklist
Before shipping to production, configure the following:
- Content types: Set
MEDIA_ALLOWED_CONTENT_TYPESto only the types your agent handles. Do not leave it empty. - File size: Set
MEDIA_MAX_SIZE_MBto the smallest value that works for your use case. This is independent ofMAX_REQUEST_SIZE(10 MB by default), which applies to entire HTTP request bodies. - Storage backend: Use
cloudstorage (S3 or GCS) for any deployment with multiple workers. Thelocalbackend requires a shared filesystem and is not suitable for containerized deployments. Memory storage is development-only. - Extraction cache: Configure a checkpointer (PostgreSQL + Redis recommended) to share document extraction cache across workers. Without it, the same document is extracted per worker per process, wasting CPU.
- Ownership enforcement: Set
MEDIA_REQUIRE_OWNER=trueif you need absolute certainty that no file without a recorded owner can be accessed. Files uploaded before ownership tracking was enabled have no owner and are otherwise allowed with a warning. - Signed URL TTL: Adjust
MEDIA_SIGNED_URL_TTL_SECONDSbased on how long clients hold URLs (longer for downloads from slow connections, shorter for security).
See also
- REST API: Files reference
- REST API: Graph endpoints (invoke, stream, stop, fix)
- Environment variables
- Media concepts (sender side)
- Production checklist (full hardening guide)
Frequently asked questions
- How do I restrict which file types my agent accepts?
- Set MEDIA_ALLOWED_CONTENT_TYPES to a comma-separated list of MIME types or wildcards (e.g., image/*,application/pdf). Empty (the default) accepts all types.
- What happens to documents when they are uploaded?
- If DOCUMENT_HANDLING=extract_text (the default), the server extracts plain text at upload time and caches it. When the client later references that file, the document block is replaced with the cached text before your graph runs.
- Should I use local or cloud storage?
- Use cloud storage (S3 or GCS) for any multi-worker deployment. Local storage assumes a shared filesystem and loses data on restart; memory storage loses everything on every restart.