Using Azure AI Speech for Text-to-Speech
This guide covers how to use Azure AI Speech for Text-to-Speech with Open WebUI. Open WebUI sends text to Azure's text to speech REST API and plays back the audio it returns.
See the companion guide: Using Azure AI Speech for Speech-to-Text
Requirements
- An Azure Speech resource (Microsoft's documentation also calls it a Foundry resource for Speech). See Microsoft's Speech service regions table for text to speech availability by region.
- The resource's key, from Resource Management > Keys and Endpoint in the Azure portal. Either of the two keys works.
- Open WebUI installed and running
Quick Setup (UI)
- Click your profile icon (bottom-left corner)
- Select Admin Panel, then Settings. Settings opens in a window.
- In the sidebar, choose Audio under Admin, in the Experience group. The other Audio entry, in the Preferences group, holds your personal audio settings. On a narrow screen these labels are hidden, so open the admin one directly at
/?settings=admin:audioinstead. - Configure the following:
| Setting | Value |
|---|---|
| Text-to-Speech Engine | Azure AI Speech |
| API Key | Your Speech resource key |
| Azure Region | The region of your Speech resource, as an identifier such as westus or westeurope |
| Endpoint URL | Leave blank, unless you use a sovereign cloud (see Endpoint URL) |
- Click Save. Choosing the engine already saves on its own and clears TTS Voice, which is why the voice comes after this step.
- Switch to another settings tab and back to Audio. The voice list is fetched when the Audio tab opens, so it only loads with your key after this.
- Configure the voice and format:
| Setting | Value |
|---|---|
| TTS Voice | Type part of a voice name or language to search, then pick a voice, for example en-US-JennyNeural |
| Output format | Leave the default audio-24khz-160kbitrate-mono-mp3 (see Output Format) |
- Click Save
The key only works with the region the resource was created in. Microsoft notes that keys are region-scoped, so a key used with a different region fails to authenticate.
If Azure Region and Endpoint URL are both blank, speech is still requested from eastus, but the voice list is not loaded at all. Set the region your resource is in.
Choosing Voices
Open WebUI loads the voice list for your region from Azure. Each row in TTS Voice shows the voice's full name, such as en-US-JennyNeural, with its display label next to it in grey. Only a few rows show at a time, so type part of a name or a language, such as en-GB, to narrow the list.
The full name is what Open WebUI sends to Azure, and it takes the speech language from the first two parts of that name (en-US from en-US-JennyNeural). If you type a voice instead of picking one, use the full name.
Which voice is used for a reply:
- The model's own TTS Voice, if one is set when editing the model in Workspace → Models
- Otherwise, the voice the user picked under Set Voice in their personal Audio settings, as long as the admin's TTS Voice is still the one that was set when they picked it. If the admin changes it, users hear the new default until they pick again.
- Otherwise, the admin's TTS Voice
For the full list of voices and languages, see Microsoft's Language and voice support.
Endpoint URL
By default, Open WebUI sends requests to https://<region>.tts.speech.microsoft.com, built from Azure Region. Endpoint URL replaces that address. Open WebUI adds /cognitiveservices/v1 to it to request speech and /cognitiveservices/voices/list to load the voice list, so enter only the scheme and host, with no path.
This is not your resource's endpoint from Keys and Endpoint, the one the STT guide can use. Microsoft documents the voice list at /tts/cognitiveservices/voices/list on that address, while Open WebUI requests /cognitiveservices/voices/list, so the voice list would not load.
For the sovereign clouds, Microsoft documents these addresses, with <region> replaced by your region identifier:
| Cloud | Endpoint URL |
|---|---|
| Azure Government | https://<region>.tts.speech.azure.us |
| Azure operated by 21Vianet | https://<region>.tts.speech.azure.cn |
See Speech service in sovereign clouds for the region identifiers.
Output Format
Output format is sent to Azure as the audio format to return. Keep one of the MP3 formats, such as the default audio-24khz-160kbitrate-mono-mp3. Open WebUI stores the audio it receives as an .mp3 file and, except in the slim image, serves it as MP3 whatever format Azure returned. The slim image refuses audio types outside its list of browser-playable formats, such as MP3, WAV and Ogg. Microsoft's audio outputs list has every value.
Environment Variables Setup
If you prefer to configure via environment variables:
services:
open-webui:
image: ghcr.io/open-webui/open-webui:main
environment:
- AUDIO_TTS_ENGINE=azure
- AUDIO_TTS_API_KEY=your-speech-resource-key
- AUDIO_TTS_AZURE_SPEECH_REGION=westeurope
- AUDIO_TTS_VOICE=en-US-JennyNeural
# ... other configurationSet AUDIO_TTS_VOICE to an Azure voice name. Its default, alloy, is an OpenAI voice.
All Azure TTS Environment Variables
| Variable | Description | Default |
|---|---|---|
AUDIO_TTS_ENGINE | Set to azure | empty (uses browser-only TTS) |
AUDIO_TTS_API_KEY | Your Speech resource key (shared with the ElevenLabs engine) | empty |
AUDIO_TTS_AZURE_SPEECH_REGION | Region identifier of your Speech resource | empty (speech uses eastus; without a base URL, the voice list does not load) |
AUDIO_TTS_AZURE_SPEECH_BASE_URL | Replaces https://<region>.tts.speech.microsoft.com | empty |
AUDIO_TTS_AZURE_SPEECH_OUTPUT_FORMAT | Audio output format | audio-24khz-160kbitrate-mono-mp3 |
AUDIO_TTS_VOICE | Full Azure voice name | alloy (not an Azure voice) |
Saving the admin Audio settings stores every field on that tab, empty ones included, and choosing an engine saves too. With the default ENABLE_PERSISTENT_CONFIG=true, those stored values take precedence over the AUDIO_* environment variables from then on. See ENABLE_PERSISTENT_CONFIG.
Testing TTS
- Start a new chat
- Send a message to any model
- Click the speaker icon on the AI response to hear it read aloud
Troubleshooting
No voices shown in the list
- Confirm API Key and Azure Region are set and saved
- Switch to another settings tab and back to Audio to reload the list
- Check the Open WebUI logs for
Error fetching Azure voices
Speech fails with a 401 error
Azure rejected the key. Check that the key is copied correctly from Keys and Endpoint and that Azure Region is the region your Speech resource is in.
Speech fails with another error
- Confirm TTS Voice is set to a full Azure voice name from the list
- Confirm Output format is one of the MP3 values from Microsoft's audio outputs list
- Check the Open WebUI logs for the full error
For broader audio debugging, see the Audio Troubleshooting Guide.
Cost Considerations
Azure bills Speech usage under your Azure subscription. Check Azure's Speech pricing page for current rates.