Voice
Real-time speech-to-text in the chat composer. The user speaks, the runtime transcribes, the agent runs the resulting prompt.
This feature (voice) hasn't been tagged in any Deep Agents cell yet. Try CopilotKit's Built-in Agent, LangGraph (Python), LangGraph (TypeScript).
You have a working chat surface and you want users to be able to speak instead of type. By the end of this guide, the chat composer will sprout a mic button, recorded audio will be transcribed by the runtime, and the transcript will auto-send to the agent like any other message.
When to use this#
- Hands-free or accessibility flows where typing isn't the right input modality.
- Mobile or kiosk surfaces where a long voice query is faster than thumb-typing.
- Demo and test loops where you want canned audio to drive the chat without a microphone.
If you only need file uploads (audio, images, video, documents), use Multimodal Attachments instead. Voice is specifically about live transcription of recorded speech into chat input.
Frontend#
<CopilotChat /> from @copilotkit/react-core/v2 renders the mic button automatically when the runtime advertises audioFileTranscriptionEnabled: true on its /info endpoint. There's nothing to wire up on the chat surface itself:
deepagents::voice. Known demos are bundled from manifest demos[i]; check the cell id and framework slug.The runtimeUrl="/api/copilotkit-voice" points the browser to your Next.js API route. When the user clicks the mic, the chat captures audio, POSTs it to that runtime route's /transcribe endpoint, drops the resulting transcript into the composer, and submits.
Driving the demo without a mic#
For Playwright runs, screenshots, or any flow where prompting for mic permissions is awkward, ship a button that emits a canned sample phrase through an onTranscribed callback, bypassing the transcription endpoint entirely:
deepagents::voice. Known demos are bundled from manifest demos[i]; check the cell id and framework slug.The parent chat component can then drop that text into the composer's textarea (matched via data-testid="copilot-chat-textarea") using the native value setter and a synthetic input event so React's managed state updates correctly.
Backend#
Next.js API route#
Create a dedicated API route at app/api/copilotkit-voice/[[...slug]]/route.ts. The [[...slug]] catch-all pattern lets the V2 runtime handle its internal URL routing (/info, /agent/:id/run, /transcribe, etc.) under the /api/copilotkit-voice base path.
Wire up the V2 runtime with a TranscriptionService. The V1 wrapper drops the transcriptionService option, so use createCopilotRuntimeHandler from @copilotkit/runtime/v2 directly:
deepagents::voice. Known demos are bundled from manifest demos[i]; check the cell id and framework slug.The basePath: "/api/copilotkit-voice" in createCopilotRuntimeHandler must match the API route's directory path. With transcriptionService set, the runtime advertises audioFileTranscriptionEnabled: true on /info (which is what tells the chat to render the mic button) and routes POST /transcribe to the service.
Without a service, `/transcribe` answers 503
A runtime with no transcriptionService still serves the route, and answers every request
503 with { "error": "service_not_configured" }. The mic button never appears, so the
symptom is a chat with no voice input rather than a visible server error — check /info for
audioFileTranscriptionEnabled when voice silently doesn't show up.
Calling `/transcribe` yourself
The chat handles this for you; these are the rules if you post to the route directly. As
multipart, the audio field must be named audio — any other name reads as absent and the
route answers invalid_request. As JSON, mimeType is required alongside the base64
audio, and a payload without it is rejected the same way.
Custom transcription backends#
TranscriptionService from @copilotkit/runtime/v2 is an abstract class. Subclass it to plug in any transcription provider — Whisper, AssemblyAI, Deepgram, your own model. The library ships TranscriptionServiceOpenAI as the canonical reference implementation.
Return a string, and let provider errors through
transcribe returns the transcript as a string — the handler wraps it into
{ transcription } itself, so returning a richer object is a type error.
Let the provider's own errors propagate unchanged. The runtime classifies failures by reading
the error text for markers like rate, 429, auth and too long, so a provider message
such as OpenAI returned 429 rate limited maps to the right error code on its own. Replacing
it with your own wording bypasses that and everything lands as a generic provider error.
A useful pattern is wrapping your service in a guard that returns a clean 4xx when credentials aren't configured, instead of an opaque 5xx from the underlying SDK:
deepagents::voice. Known demos are bundled from manifest demos[i]; check the cell id and framework slug.