Voice Mode
Voice Mode turns a chat into a spoken conversation. You talk, Open WebUI transcribes what you said and sends it as your message, and the model's reply is read back to you a sentence at a time as it is written. The conversation stays an ordinary chat, so every turn is saved as text you can read and continue later.
Voice Mode uses the instance's speech-to-text and text-to-speech settings, so it works with whichever engines your admin has set up.
Starting Voice Mode
- Open a chat with one model selected. Voice Mode talks to a single model at a time.
- Leave the message box empty. While it is empty and nothing is attached, the Voice mode button takes the place of the send button.
- Click Voice mode and allow microphone access when the browser asks.
On a wide screen, Voice Mode opens in the side panel next to the chat, so the transcript stays in view. On a phone or a narrow window, it fills the screen. Your screen is kept awake for as long as Voice Mode is open, where the browser supports it.
To end it, click End call. That also turns off the microphone and any camera or screen share.
You can also open a chat with Voice Mode already running by adding ?call=true to its URL. See URL Parameters.
Talking With the Model
- Speak, then pause. Voice Mode starts recording when it hears you and sends what you said after about two seconds of silence.
- The reply is spoken as it streams. Each sentence is read aloud as soon as it is complete, so a long answer starts playing before the model has finished writing it. Reasoning and tool-call details are not read aloud.
- Tap to interrupt. While the model is speaking, tap its picture or the status line to stop it. If the reply is still being written, that stops the reply as well.
- Or talk over it. With Allow Voice Interruption in Call turned on (see Your Voice Mode Settings), starting to speak cuts the model off the same way. With it off, the microphone ignores everything while the model is speaking.
- Mute. Click the microphone button, or press M on a keyboard. Muting in the middle of a sentence discards that sentence instead of sending part of it. If you mute while the model is speaking, the microphone turns back on when it finishes.
The status line under the model's picture shows what Voice Mode is doing:
| Status | What it means |
|---|---|
| Listening... | Waiting for you to speak, or recording you |
| Thinking... | Transcribing what you said and sending it |
| Tap to interrupt | The model is speaking |
| Muted | Your microphone is muted |
Show the Model What You See
Click Camera to add video to the conversation. While video is on, each time you finish speaking, a still frame from the camera is attached to your message, so a model that can read images can answer questions about what it sees. The frames are saved with the chat like any other attached image.
Switch camera lists your cameras and, where the browser supports it, Screen Share, which sends frames of a screen or window instead. Click Stop camera in the corner of the video to turn it off.
Voices and Playback
The reply is spoken in:
- the model's own TTS Voice, if one is set when editing the model in Workspace > Models
- otherwise, the voice you picked under Settings > Audio > Set Voice
- otherwise, the default voice your admin set
Speech Playback Speed in Settings > Audio sets how fast replies are read.
The speech itself comes from the text-to-speech engine your admin chose in Admin Panel > Settings > Audio. If the admin leaves it on Web API, your browser's built-in voices are used. You can also pick Kokoro.js (Browser) as your Text-to-Speech Engine in Settings > Audio to generate speech in your browser.
The admin's Response Splitting setting decides how much of the reply is spoken at once:
| Response Splitting | What Voice Mode does |
|---|---|
| Punctuation (default) | Speaks one sentence at a time as the reply streams |
| Paragraphs | Speaks one paragraph at a time |
| None | Waits for the whole reply, then speaks it |
What you say is always transcribed on the server with the admin's speech-to-text engine, even if you chose Web API for dictation in your own settings. The Language under Settings > Audio > STT Settings is passed along with it. If the admin's speech-to-text engine is set to Web API, Voice Mode is not available.
Your Voice Mode Settings
Two options in Settings > Interface, under Voice, apply only to Voice Mode. Both are off by default.
| Setting | What it does |
|---|---|
| Allow Voice Interruption in Call | Lets you interrupt the model by speaking while it talks |
| Display Emoji in Call | Shows an emoji matching each spoken sentence in place of the model's picture. Each emoji is picked by the task model, one request per sentence |
The Voice Mode Prompt
While Voice Mode is on, Open WebUI adds voice instructions to the start of the system prompt. The built-in instructions ask the model for short, warm, conversational answers that are easy to follow when heard rather than read.
Admins can turn this off, or replace the instructions with their own, under Admin Panel > Settings > Interface > Voice Mode Prompt. The same settings are available as ENABLE_VOICE_MODE_PROMPT and VOICE_MODE_PROMPT_TEMPLATE. An empty template uses the built-in instructions.
Permissions
Admins can control access to Voice Mode on a per-role or per-group basis.
- Location: Admin Panel > Users > Groups > Default permissions > Permissions > Allow Call. To set a group-specific value, edit the group and open its Permissions tab.
- Environment Variable:
USER_PERMISSIONS_CHAT_CALL(Default:True)
If disabled, users do not see the Voice mode button. Admins always do.
Troubleshooting
- "Permission denied when accessing media devices": the browser blocked the microphone. Browsers only allow microphone access on HTTPS or
localhost, so opening Open WebUI by its IP address from another device usually causes this. See Microphone Access Issues. - "Select only one model to call": more than one model is selected. Remove the others and try again.
- "Call feature is not supported when using Web STT engine": the admin's speech-to-text engine is set to Web API. An admin needs to choose a server-side engine in Admin Panel > Settings > Audio.
- It keeps listening and never sends: Voice Mode sends after about two seconds of silence, so steady background noise can keep it recording. Mute while it is noisy, or move somewhere quieter.
- No sound from the model: see Text-to-Speech (TTS) Issues.