Build a realtime audio conversation
In shortCapture microphone input, stream it to a live agent with client.realtime(), receive PCM16 audio output, and play it back in the browser.
- 11 min read
- 13 sections
- Updated
- v0.10.0
- Markdown
client.realtime() opens a WebSocket connection that carries bidirectional PCM16 audio frames and JSON control events between your browser and a live agent. The API is transport only: it handles the audio bytes and event delivery, but leaves microphone capture and speaker playback to you. This guide covers both sides: building a mic capturer to send PCM16 at 16 kHz, a PCM16 player to receive audio at 24 kHz, managing the session lifecycle, and handling interrupts.
For the complete RealtimeSession API (every event type, reconnect behavior, init parameters) see reference/client/realtime.
Prerequisites
- A graph rooted at a
LiveAgent. A turn-based graph rejects the connection: the server sends a fatalerrorevent withcode: 'not_live'and closes the socket with code1008. Checkinfo.is_realtimefromclient.graph()before rendering an audio UI. - A configured
TenxGraphClientinstance. On Node 18 or 20, you must passwebSocketImpl(thewspackage) in the client config; see create-client for details. Browsers ship WebSocket natively. - A secure context.
getUserMediarequires HTTPS orlocalhost. Onhttp://,getUserMediarejects immediately.
The audio contract
| Direction | Format |
|---|---|
Up, via sendAudio() |
PCM16, mono, 16 kHz |
Down, on the 'audio' channel |
PCM16, mono, 24 kHz |
The two rates differ, and getting them backwards produces audio that plays at the wrong speed rather than an error. The package exports REALTIME_INPUT_SAMPLE_RATE (16000) and REALTIME_OUTPUT_SAMPLE_RATE (24000), use them rather than literals.
Step 1: Build a PCM16 player
The agent sends audio as a stream of PCM16 frames (typically 50-200ms each). If you play each frame independently starting at the current time, you get silence between them. Instead, schedule each frame to start exactly when the previous one ends. This keeps playback continuous and eliminates the stuttering that makes the agent’s voice sound broken.
export const createPcmPlayer = (defaultSampleRate = 24000) => {
let context = null;
let nextTime = 0;
const ensureContext = () => {
if (!context) {
const Ctor = window.AudioContext || window.webkitAudioContext;
context = new Ctor();
}
if (context.state === 'suspended') context.resume();
return context;
};
return {
play: (pcmBytes, sampleRate = defaultSampleRate) => {
if (!pcmBytes || pcmBytes.byteLength < 2) return;
const c = ensureContext();
const frames = Math.floor(pcmBytes.byteLength / 2);
const view = new DataView(pcmBytes.buffer, pcmBytes.byteOffset, pcmBytes.byteLength);
const buffer = c.createBuffer(1, frames, sampleRate);
const channel = buffer.getChannelData(0);
for (let i = 0; i < frames; i++) {
channel[i] = view.getInt16(i * 2, true) / 32768; // little-endian
}
const source = c.createBufferSource();
source.buffer = buffer;
source.connect(c.destination);
const startAt = Math.max(c.currentTime, nextTime);
source.start(startAt);
nextTime = startAt + buffer.duration;
},
close: () => {
if (context) {
try { context.close(); } catch { /* already closed */ }
context = null;
nextTime = 0;
}
},
};
};Two details are critical. First, getInt16(offset, true) reads little-endian because that is what the server sends. Second, nextTime tracks the end of the previous frame so you know where to start the next one. Without it, each frame starts at currentTime, which may be before the previous frame finishes, causing overlap or gaps.
Always construct the player inside a user gesture (click or tap). Browsers start an AudioContext in a suspended state until user interaction, so building it outside the event handler results in silence. The resume() call wakes a suspended context on demand.
Step 2: Capture the microphone as PCM16 at 16 kHz
Request a 16 kHz AudioContext and the browser automatically resamples the microphone input to that rate. This saves you from writing a resampler. The sample rate must be exactly 16 kHz because the agent’s speech-to-text model expects that bitrate. The code also enables echo cancellation and noise suppression to improve audio quality on the server side.
export const createMicCapture = async (onFrame, sampleRate = 16000) => {
const stream = await navigator.mediaDevices.getUserMedia({
audio: { channelCount: 1, echoCancellation: true, noiseSuppression: true },
});
const Ctor = window.AudioContext || window.webkitAudioContext;
const context = new Ctor({ sampleRate });
if (context.state === 'suspended') await context.resume();
const source = context.createMediaStreamSource(stream);
const processor = context.createScriptProcessor(4096, 1, 1);
// Route through a muted gain node so onaudioprocess keeps firing without
// echoing the microphone back out of the speakers.
const sink = context.createGain();
sink.gain.value = 0;
processor.onaudioprocess = (event) => {
const input = event.inputBuffer.getChannelData(0); // Float32, -1..1
const pcm = new Int16Array(input.length);
for (let i = 0; i < input.length; i++) {
const s = Math.max(-1, Math.min(1, input[i]));
pcm[i] = s < 0 ? s * 32768 : s * 32767;
}
onFrame(new Uint8Array(pcm.buffer));
};
source.connect(processor);
processor.connect(sink);
sink.connect(context.destination);
return {
stop: () => {
try {
processor.disconnect();
source.disconnect();
sink.disconnect();
} catch { /* already torn down */ }
stream.getTracks().forEach((t) => t.stop());
try { context.close(); } catch { /* already closed */ }
},
};
};Echo cancellation is not optional. Without it, the agent hears its voice playing from your speakers through the microphone, interprets it as new user input, and interrupts itself in an endless loop. If a user turns off echo cancellation or is using a speaker without proper shielding, recommend headphones.
createScriptProcessor is deprecated but universally supported. For production apps with high-traffic requirements, use AudioWorklet instead (it runs in a separate thread and reduces main-thread blocking). The frame-to-PCM conversion code is identical.
Always keep the returned handle and call stop() when the session ends. Forgetting to call stop() leaves the microphone indicator on in the browser, confusing users and potentially creating a privacy issue.
Step 3: Open the realtime session
Call client.realtime() with the model you want and audio modality. The method returns a RealtimeSession object immediately and begins connecting in the background.
import { TenxGraphClient } from '10xgraph-client';
const client = new TenxGraphClient({
baseUrl: 'http://localhost:8000',
authToken: token,
});
let player = createPcmPlayer(); // built inside the click handler
const session = client.realtime(
{ model: 'gemini-2.5-flash-live', modalities: 'AUDIO' },
{ reconnect: { enabled: false } },
);
session.on('open', () => setStatus('live'));
session.on('audio', (pcm, rate) => player.play(pcm, rate));
session.on('error', (e) => setError(e.message));
session.on('close', () => {
player.close();
setStatus('ended');
});
await session.ready;Wait for session.ready to resolve before proceeding, which confirms the WebSocket is open and the init frame has been sent. reconnect: { enabled: false } makes a close permanent, suitable for a “Start / End” button UI. For long-lived assistants that should survive network hiccups, leave reconnect enabled (the default) so the session reopen with the same thread_id and resume from the checkpoint.
The init model parameter is a per-session override, honoured only when the model is listed in the server’s websocket.realtime_models setting. Any other value is ignored and the agent’s own model is used. Live model availability varies by API key and region, so omit model unless you need to pick from that list.
Step 4: Implement push-to-talk (manual turn-taking)
In push-to-talk mode, the user holds down a button to speak and releases to stop. Call activityStart() before opening the mic and activityEnd() when closing it. The server uses these signals to detect turn boundaries: it stops listening and responds.
let mic = null;
const toggleMic = async () => {
if (status !== 'live') return;
if (mic) {
mic.stop();
mic = null;
session.activityEnd(); // ends the turn; the agent now responds
return;
}
try {
session.activityStart();
mic = await createMicCapture((frame) => session.sendAudio(frame));
} catch (e) {
// getUserMedia rejects when the user denies the permission prompt,
// or when the page is not on HTTPS/localhost.
setError(e?.message || 'Microphone unavailable');
session.activityEnd(); // do not leave the turn hanging open
}
};For hands-free (continuous) operation, leave voice activation detection (VAD) enabled on the server (the default) and omit activityStart / activityEnd. Stream mic frames continuously and let the server detect silence and turn boundaries automatically. Only set vad: { enabled: false } in the init frame if you want manual turn control as shown above.
Step 5: Display live transcripts
As the user speaks and the agent responds, both sides stream in as text deltas (small chunks). Each delta event includes a finished flag that signals the end of a sentence or turn. Buffer consecutive deltas from the same speaker into a single message bubble to avoid a cluttered UI.
const open = { user: null, agent: null }; // id of the in-progress bubble per role
const append = (role, text, finished) => {
setMessages((previous) => {
const openId = open[role];
if (openId == null) {
const id = nextId();
open[role] = finished ? null : id;
return [...previous, { id, role, text }];
}
const next = previous.map((m) => (m.id === openId ? { ...m, text: m.text + text } : m));
if (finished) open[role] = null;
return next;
});
};
session.on('input_transcript', (e) => append('user', e.text || '', e.finished));
session.on('output_transcript', (e) => append('agent', e.text || '', e.finished));Step 6: Handle interruption and cleanup
When a user starts talking while the agent is still speaking, the server detects this and sends an interrupted event. If you don’t clear the audio queue, the agent’s previous response will keep playing while the new one starts, creating a chaotic overlap. Always discard queued playback and create a fresh player on interrupt.
session.on('interrupted', () => {
player.close(); // discard everything queued
player = createPcmPlayer(); // fresh player for the next reply
});Always clean up when the session ends or the component unmounts. If a user navigates away without calling stop() on the mic and close() on the session, the browser shows a microphone indicator, the socket stays open, and resources leak.
useEffect(() => () => {
mic?.stop();
session?.close();
player?.close();
}, []);session.close() sends a close control frame to the server, closes the WebSocket, and prevents automatic reconnection.
Step 7: Check if the agent supports realtime audio
Not all agents are configured for realtime conversations. Before rendering an audio UI, fetch the graph info once at startup and check the is_realtime flag. If it is false, show a notice and use the standard chat interface instead.
const { info } = (await client.graph()).data;
const liveCapable = Boolean((info as { is_realtime?: boolean }).is_realtime);
if (!liveCapable) {
showNotice('This agent is not a realtime agent. Use the chat interface instead.');
}If you skip this check and try to open a realtime session against a turn-based agent, the server sends a fatal not_live error event and closes the connection. The playground’s Live page implements this gate (through its connection capabilities) to prevent users from seeing a broken interface.
Reference implementation
The 10xGraph playground (agentflow-playground/) includes a complete, production-ready implementation of realtime audio:
| File | Contains |
|---|---|
src/lib/realtime-audio.js |
createPcmPlayer(24000) and createMicCapture(onFrame, 16000) (the code from Steps 1 and 2). |
src/pages/live/components/live-session.jsx |
Session lifecycle, push-to-talk button handler, transcript coalescing, interruption handling, and cleanup. |
src/pages/live/live-page.jsx |
The live-capable UI gate (Step 7), based on the is_realtime flag. |
To run it: 10xgraph play (starts the API server and playground together), then open the Live page to test the realtime audio session.
Troubleshooting common issues
| Symptom | Cause | Fix |
|---|---|---|
| Playback is too fast or too slow. | Playing output frames intended for 24 kHz at 16 kHz playback rate, or vice versa. The audio itself is correct; only the playback speed is wrong. | Pass the sampleRate argument of the 'audio' listener to the player, as in Step 3. |
| No audio at all, no errors in the console. | The AudioContext was created before the first user interaction and is in suspended state. Browsers enforce this for security and battery reasons. |
Create the AudioContext and player inside the click handler, and call resume() on a suspended context. |
| Choppy playback with audible gaps between frames. | Each frame starts at the current time instead of when the previous frame ends. With typical frame durations of 50-200ms, this creates noticeable silence. | Schedule frames using max(currentTime, nextTime) and update nextTime by the buffer duration. See Step 1. |
| The agent constantly interrupts itself mid-sentence. | Echo cancellation is disabled. The agent hears its own voice playing from the speaker, mistakes it for new user input, and stops to respond. | Enable echoCancellation: true in the getUserMedia call (Step 2). If echo is still a problem, recommend headphones. |
getUserMedia rejects immediately with a permission error. |
The user denied the permission prompt, or the page is not on HTTPS or localhost. | Handle the rejection in a catch block, show the user a clear error message, and call activityEnd() to clean up the pending turn. |
Socket closes immediately after opening, error code not_live. |
The graph is turn-based, not a live agent. | Check info.is_realtime before offering the audio UI. |
Live session fails even though the graph is a LiveAgent. |
The live model is not available for your API key, region, or quota. | Omit the model field so the agent’s own model is used, or pick one listed in websocket.realtime_models. |
No WebSocket implementation available error in Node.js. |
Node 18 and 20 do not have a global WebSocket object. |
Pass webSocketImpl in the TenxGraphClient config: { webSocketImpl: require('ws') }. |
| Microphone indicator remains on after ending the session. | The mic.stop() method was not called or an exception prevented it from running. |
Always call stop() in the “End” button handler and in a React useEffect cleanup. Wrap calls in try/catch or use optional chaining (mic?.stop()). |
Summary
You now understand the full lifecycle of a realtime audio conversation:
- Sample rates are intentional. Input at 16 kHz optimizes speech-to-text quality; output at 24 kHz is what the model produces. Using the browser’s automatic resampling avoids decoder complexity.
- Playback scheduling prevents gaps. Tracking the end time of the previous frame and scheduling the next one there keeps the audio stream seamless.
- Echo cancellation is essential. Without it, the agent hears its own voice and loops. Always enable it in the
getUserMediaconstraints. - Turn-taking has two modes. Push-to-talk (manual
activityStart/activityEnd) suits voice-assistant UIs. Hands-free (server VAD) suits always-listening assistants. - Interruption and cleanup are critical. The
interruptedevent requires clearing the audio queue. Cleanup on unmount and on error prevents resource leaks and privacy issues. - Runtime checks prevent broken experiences. Always verify
is_realtimebefore offering the audio UI, and omit themodelfield unless you need one fromwebsocket.realtime_models.
Next steps
- See the reference for the complete
RealtimeSessionAPI: every event type, reconnect configuration, and resumption behavior. - Explore the playground’s Live page implementation for a production example.
- For Python server-side realtime audio, see use-realtime-audio (the guide for building a
LiveAgent).