Voice
Real-time speech-to-text in the chat composer. The user speaks, the runtime transcribes, the agent runs the resulting prompt.
You have a working chat surface and you want users to be able to speak instead of type. By the end of this guide, the chat composer will sprout a mic button, recorded audio will be transcribed by the runtime, and the transcript will auto-send to the agent like any other message.
When to use this#
- Hands-free or accessibility flows where typing isn't the right input modality.
- Mobile or kiosk surfaces where a long voice query is faster than thumb-typing.
- Demo and test loops where you want canned audio to drive the chat without a microphone.
If you only need file uploads (audio, images, video, documents), use Multimodal Attachments instead. Voice is specifically about live transcription of recorded speech into chat input.
Frontend#
<CopilotChat /> from @copilotkit/react-core/v2 renders the mic button automatically when the runtime advertises audioFileTranscriptionEnabled: true on its /info endpoint. There's nothing to wire up on the chat surface itself:
agent-spec::voice. Known demos are bundled from manifest demos[i]; check the cell id and framework slug.The runtimeUrl="/api/copilotkit-voice" points the browser to your Next.js API route. When the user clicks the mic, the chat captures audio, POSTs it to that runtime route's /transcribe endpoint, drops the resulting transcript into the composer, and submits.
Driving the demo without a mic#
For Playwright runs, screenshots, or any flow where prompting for mic permissions is awkward, ship a button that emits a canned sample phrase through an onTranscribed callback, bypassing the transcription endpoint entirely:
agent-spec::voice. Known demos are bundled from manifest demos[i]; check the cell id and framework slug.The parent chat component can then drop that text into the composer's textarea (matched via data-testid="copilot-chat-textarea") using the native value setter and a synthetic input event so React's managed state updates correctly.
Backend#
Next.js API route#
Create a dedicated API route at app/api/copilotkit-voice/[[...slug]]/route.ts. The [[...slug]] catch-all pattern lets the V2 runtime handle its internal URL routing (/info, /agent/:id/run, /transcribe, etc.) under the /api/copilotkit-voice base path.
Wire up the V2 runtime with a TranscriptionService. The V1 wrapper drops the transcriptionService option, so use createCopilotRuntimeHandler from @copilotkit/runtime/v2 directly:
agent-spec::voice. Known demos are bundled from manifest demos[i]; check the cell id and framework slug.The basePath: "/api/copilotkit-voice" in createCopilotRuntimeHandler must match the API route's directory path. With transcriptionService set, the runtime advertises audioFileTranscriptionEnabled: true on /info (which is what tells the chat to render the mic button) and routes POST /transcribe to the service.
Without a service, `/transcribe` answers 503
A runtime with no transcriptionService still serves the route, and answers every request
503 with { "error": "service_not_configured" }. The mic button never appears, so the
symptom is a chat with no voice input rather than a visible server error — check /info for
audioFileTranscriptionEnabled when voice silently doesn't show up.
Calling `/transcribe` yourself
The chat handles this for you; these are the rules if you post to the route directly. As
multipart, the audio field must be named audio — any other name reads as absent and the
route answers invalid_request. As JSON, mimeType is required alongside the base64
audio, and a payload without it is rejected the same way.
Custom transcription backends#
TranscriptionService from @copilotkit/runtime/v2 is an abstract class. Subclass it to plug in any transcription provider — Whisper, AssemblyAI, Deepgram, your own model. The library ships TranscriptionServiceOpenAI as the canonical reference implementation.
Return a string, and let provider errors through
transcribe returns the transcript as a string — the handler wraps it into
{ transcription } itself, so returning a richer object is a type error.
Let the provider's own errors propagate unchanged. The runtime classifies failures by reading
the error text for markers like rate, 429, auth and too long, so a provider message
such as OpenAI returned 429 rate limited maps to the right error code on its own. Replacing
it with your own wording bypasses that and everything lands as a generic provider error.
A useful pattern is wrapping your service in a guard that returns a clean 4xx when credentials aren't configured, instead of an opaque 5xx from the underlying SDK:
agent-spec::voice. Known demos are bundled from manifest demos[i]; check the cell id and framework slug.