Send media to models

In shortSend images, audio, and documents to your models from Python using content blocks, MediaRef, and storage backends.

  • 6 min read
  • 12 sections
  • Updated
  • v0.10.0
  • Markdown

10xGraph supports multimodal input natively: send images, audio, video, and documents directly alongside text in messages. Each content type is a typed block, and media is referenced via MediaRef, which decouples message structure from where bytes actually live. You choose the transport: external URLs, inline base64, or deduplicated references through a storage backend.


Content blocks and MediaRef

Any message is a list of content blocks. 10xGraph provides typed blocks for each medium:

Python
from tenxgraph.core.state import (
    TextBlock,
    ImageBlock,
    AudioBlock,
    VideoBlock,
    DocumentBlock,
    DataBlock,
    Message,
    MediaRef,
)

Each block except TextBlock holds a MediaRef that tells 10xGraph where the bytes are. MediaRef has three kind values:

Kind Use case
url External URL, or graph://media/{key} from a media store (use media_store.to_media_ref(key, mime_type) to build it)
data Inline base64 (data_base64) for small payloads only
file_id A provider-managed file ID, for example one returned by the OpenAI or Gemini file APIs

An Agent only keeps media blocks when you pass it a multimodal_config. Without one, image, audio, video and document blocks are stripped before the model call. The examples below all pass MultimodalConfig().

Install the provider extra for the model you use, for example pip install "10xgraph[google-genai]" for Gemini or pip install "10xgraph[openai]" for OpenAI.


Example 1: Image from a public URL

Send an image directly from the web:

Python
from tenxgraph.core.graph import Agent, StateGraph
from tenxgraph.core.state import Message, TextBlock, ImageBlock, MediaRef
from tenxgraph.storage.media import MultimodalConfig
from tenxgraph.utils import END

agent = Agent(
    model="gemini-2.5-flash",
    provider="google",
    multimodal_config=MultimodalConfig(),
)

graph = StateGraph()
graph.add_node("MAIN", agent)
graph.set_entry_point("MAIN")
graph.add_edge("MAIN", END)
app = graph.compile()

messages = [
    Message(
        role="user",
        content=[
            TextBlock(text="What is in this image?"),
            ImageBlock(
                media=MediaRef(
                    kind="url",
                    url="https://example.com/photo.png",
                    mime_type="image/png",
                )
            ),
        ],
    )
]

result = app.invoke({"messages": messages})
print(result["messages"][-1].text())

This is the simplest approach: you give the model a URL and 10xGraph resolves it for the provider. Agent is a graph node, so run it through a compiled graph (app), which later examples reuse.


Example 2: Inline base64 for small media

Embed bytes directly in the message for small payloads:

Python
import base64
from tenxgraph.core.state import Message, TextBlock, ImageBlock, MediaRef

# Read a local image
with open("photo.jpg", "rb") as f:
    image_bytes = f.read()

b64 = base64.b64encode(image_bytes).decode()

messages = [
    Message(
        role="user",
        content=[
            TextBlock(text="Describe this photo in one sentence."),
            ImageBlock(
                media=MediaRef(
                    kind="data",
                    data_base64=b64,
                    mime_type="image/jpeg",
                )
            ),
        ],
    )
]

result = app.invoke({"messages": messages})

This works well for development and testing. For production or repeated references, use a media store instead.


Example 3: Media store for reusable references

Store media once, reference it many times:

Python
import asyncio
from tenxgraph.core.state import Message, TextBlock, ImageBlock, MediaRef
from tenxgraph.storage.media import InMemoryMediaStore

media_store = InMemoryMediaStore()

# Bind the store to the graph so graph://media/ references can be resolved
app = graph.compile(media_store=media_store)

# Upload and get a storage key
async def setup():
    with open("chart.png", "rb") as f:
        key = await media_store.store(data=f.read(), mime_type="image/png")
    return key

key = asyncio.run(setup())

# Reference by key in a message (builds a graph://media/{key} URL)
messages = [
    Message(
        role="user",
        content=[
            TextBlock(text="Analyze this chart and extract the numbers."),
            ImageBlock(media=media_store.to_media_ref(key, "image/png")),
        ],
    )
]

result = app.invoke({"messages": messages})

The store returns an opaque key. You can reference that key in unlimited messages without re-uploading the file. This is recommended for production.


Audio and documents

Audio blocks work the same way as images:

Python
import base64
from tenxgraph.core.state import AudioBlock

with open("clip.wav", "rb") as f:
    audio_bytes = f.read()

messages = [
    Message(
        role="user",
        content=[
            TextBlock(text="Transcribe this audio."),
            AudioBlock(
                media=MediaRef(
                    kind="data",
                    data_base64=base64.b64encode(audio_bytes).decode(),
                    mime_type="audio/wav",
                ),
                sample_rate=16000,
                channels=1,
            ),
        ],
    )
]

Document blocks can include both extracted text and the raw file. Most models take text, so provide the text when you have it:

Python
from tenxgraph.core.state import DocumentBlock

messages = [
    Message(
        role="user",
        content=[
            TextBlock(text="Summarize this contract."),
            DocumentBlock(
                text="Optional: pre-extracted text if you have it already.",
                media=MediaRef(
                    kind="url",
                    url="https://example.com/contract.pdf",
                    mime_type="application/pdf",
                ),
            ),
        ],
    )
]

Configure media with MultimodalConfig

Pass MultimodalConfig to Agent to enable media for that agent. Agents without it strip media blocks, which keeps text-only agents in a multi-agent graph from receiving images added by an earlier agent.

Python
from tenxgraph.core.graph import Agent
from tenxgraph.storage.media import (
    ImageHandling,
    DocumentHandling,
    MultimodalConfig,
)

agent = Agent(
    model="gpt-4o",
    provider="openai",
    multimodal_config=MultimodalConfig(
        image_handling=ImageHandling.BASE64,
        document_handling=DocumentHandling.EXTRACT_TEXT,
        max_image_size_mb=10.0,
        max_image_dimension=2048,
        supported_image_types={"image/jpeg", "image/png", "image/webp"},
        supported_doc_types={"application/pdf"},
    ),
)

The fields and defaults are:

Field Default
image_handling ImageHandling.BASE64 (also URL, FILE_ID)
document_handling DocumentHandling.EXTRACT_TEXT (also FORWARD_RAW, SKIP)
max_image_size_mb 10.0
max_image_dimension 2048
supported_image_types JPEG, PNG, WebP, GIF
supported_doc_types PDF, DOCX

The size, dimension and type limits are enforced by MediaProcessor (tenxgraph.storage.media), which you call yourself to validate or resize images before sending. The Agent does not apply them automatically. Document text extraction is not part of the core library; it lives in the API server, so pass pre-extracted text in DocumentBlock(text=...) when you need it.


Complete graph with media store

Here is an example that combines a graph, checkpointer, and media store:

Python
import asyncio
from tenxgraph.core.graph import Agent, StateGraph
from tenxgraph.core.state import (
    ImageBlock,
    Message,
    TextBlock,
)
from tenxgraph.storage.checkpointer import InMemoryCheckpointer
from tenxgraph.storage.media import (
    ImageHandling,
    InMemoryMediaStore,
    MultimodalConfig,
)
from tenxgraph.utils import END

async def main():
    checkpointer = InMemoryCheckpointer()
    media_store = InMemoryMediaStore()

    agent = Agent(
        model="gemini-2.5-flash",
        provider="google",
        system_prompt=[{"role": "system", "content": "You are a helpful image analyst."}],
        multimodal_config=MultimodalConfig(
            image_handling=ImageHandling.BASE64,
        ),
    )

    graph = StateGraph()
    graph.add_node("agent", agent)
    graph.set_entry_point("agent")
    graph.add_edge("agent", END)

    app = graph.compile(checkpointer=checkpointer, media_store=media_store)

    # Upload an image
    with open("sample.jpg", "rb") as f:
        image_key = await media_store.store(
            data=f.read(),
            mime_type="image/jpeg",
        )

    # Create a message referencing the stored image
    messages = [
        Message(
            role="user",
            content=[
                TextBlock(text="What objects do you see?"),
                ImageBlock(media=media_store.to_media_ref(image_key, "image/jpeg")),
            ],
        )
    ]

    # Invoke
    result = await app.ainvoke(
        {"messages": messages},
        config={"thread_id": "image-analysis-1"},
    )

    # Print the response
    print(result["messages"][-1].text())

asyncio.run(main())

Media store backends

10xGraph provides three storage backends. Choose based on your deployment model:

InMemoryMediaStore: Development and testing. Data is lost on process restart.

Python
from tenxgraph.storage.media import InMemoryMediaStore

store = InMemoryMediaStore()
key = await store.store(data=image_bytes, mime_type="image/png")
bytes_back, mime = await store.retrieve(key)

LocalFileMediaStore: Single-server deployments. Stores files on disk with sharding.

Python
from tenxgraph.storage.media import LocalFileMediaStore

store = LocalFileMediaStore(base_dir="./media")
key = await store.store(data=pdf_bytes, mime_type="application/pdf")

Files are sharded as {base_dir}/{key[:2]}/{key[2:4]}/{key}.{ext} with a .meta.json sidecar for metadata.

CloudMediaStore: Multi-worker and cloud deployments. Stores in S3 or GCS.

Terminal
pip install "10xgraph[cloud-storage]"
Python
from cloud_storage_manager import CloudStorageFactory, StorageProvider, StorageConfig, AwsConfig
from tenxgraph.storage.media import CloudMediaStore

config = StorageConfig(
    aws=AwsConfig(bucket_name="my-bucket", access_key_id="...", secret_access_key="...")
)
storage = CloudStorageFactory.get_storage(StorageProvider.AWS, config)
store = CloudMediaStore(storage, prefix="10xgraph-media")

All stores expose the same async interface: store(), retrieve(), delete(), exists(), and get_metadata(), plus to_media_ref() to build a message reference.


Common patterns and pitfalls

Don’t re-encode for every message. If you send the same image multiple times, store it once and reference the key. This saves bandwidth and encoding overhead.

Provide MIME types. Always include mime_type in MediaRef. The provider needs it to interpret the bytes correctly.

Size limits are configurable. When running behind the API server, check MEDIA_MAX_SIZE_MB and MEDIA_ALLOWED_CONTENT_TYPES environment variables. For Python graphs, MultimodalConfig.max_image_size_mb is enforced only when you use MediaProcessor.

Pre-extract text when possible. Document blocks can carry both raw bytes and pre-extracted text. If you have text extracted already (e.g. from a PDF parser), include it via the text parameter so the model has text to work with even if it cannot process the raw file.

Video is supported. You can send video via VideoBlock the same way you send images and audio. Whether the model processes it depends on provider capabilities.


What you learned

  • Message content is a list of typed blocks: TextBlock, ImageBlock, AudioBlock, DocumentBlock, etc.
  • MediaRef decouples message structure from media transport: kind="url" for public URLs, kind="data" for inline base64, kind="file_id" for provider file IDs; use media_store.to_media_ref() for storage references.
  • Use InMemoryMediaStore for development, LocalFileMediaStore for single servers, CloudMediaStore for multi-worker setups.
  • MultimodalConfig on the Agent enables media for that agent; without it, media blocks are stripped.
  • Always provide MIME types and respect size limits.

Frequently asked questions

When should I use each MediaRef kind?
Use `url` for public or CDN-hosted media, and for `graph://media/{key}` references built with `media_store.to_media_ref()`. Use `data` for small payloads embedded directly in the message. Use `file_id` only for provider-managed file IDs.
Do I need a media store for production?
For single-server setups, LocalFileMediaStore works. For multi-worker or cloud deployments, use CloudMediaStore (S3/GCS) so all workers can access the same files.
Can I send media over the REST API?
Yes. Upload via POST /v1/files/upload, get a file_id, and reference it in your message. That flow is separate from Python; see /docs/server/files-and-multimodal.
Last updated for v0.10.0Edit this page on GitHubReport an issue