# Using Azure AI Speech for Text-to-Speech Source: https://docs.openwebui.com/features/chat-conversations/audio/text-to-speech/azure-tts-integration This guide covers how to use Azure AI Speech for Text-to-Speech with Open WebUI. Open WebUI sends text to Azure's [text to speech REST API](https://learn.microsoft.com/azure/ai-services/speech-service/rest-text-to-speech) and plays back the audio it returns. > **Looking for STT?** > > See the companion guide: [Using Azure AI Speech for Speech-to-Text](/features/chat-conversations/audio/speech-to-text/azure-stt-integration) ## Requirements - An Azure Speech resource (Microsoft's documentation also calls it a Foundry resource for Speech). See Microsoft's [Speech service regions table](https://learn.microsoft.com/azure/ai-services/speech-service/regions?tabs=tts) for text to speech availability by region. - The resource's key, from **Resource Management** > **Keys and Endpoint** in the Azure portal. Either of the two keys works. - Open WebUI installed and running ## Quick Setup (UI) 1. Click your **profile icon** (bottom-left corner) 2. Select **Admin Panel**, then **Settings**. Settings opens in a window. 3. In the sidebar, choose **Audio** under **Admin**, in the _Experience_ group. The other **Audio** entry, in the _Preferences_ group, holds your personal audio settings. On a narrow screen these labels are hidden, so open the admin one directly at `/?settings=admin:audio` instead. 4. Configure the following: | Setting | Value | | --- | --- | | Text-to-Speech Engine | Azure AI Speech | | API Key | Your Speech resource key | | Azure Region | The region of your Speech resource, as an identifier such as westus or westeurope | | Endpoint URL | Leave blank, unless you use a sovereign cloud (see Endpoint URL) | 1. Click **Save**. Choosing the engine already saves on its own and clears **TTS Voice**, which is why the voice comes after this step. 2. Switch to another settings tab and back to **Audio**. The voice list is fetched when the Audio tab opens, so it only loads with your key after this. 3. Configure the voice and format: | Setting | Value | | --- | --- | | TTS Voice | Type part of a voice name or language to search, then pick a voice, for example en-US-JennyNeural | | Output format | Leave the default audio-24khz-160kbitrate-mono-mp3 (see Output Format) | 1. Click **Save** The key only works with the region the resource was created in. Microsoft notes that keys are region-scoped, so a key used with a different region fails to authenticate. > **Always set the region** > > If **Azure Region** and **Endpoint URL** are both blank, speech is still requested from `eastus`, but the voice list is not loaded at all. Set the region your resource is in. ## Choosing Voices Open WebUI loads the voice list for your region from Azure. Each row in **TTS Voice** shows the voice's full name, such as `en-US-JennyNeural`, with its display label next to it in grey. Only a few rows show at a time, so type part of a name or a language, such as `en-GB`, to narrow the list. The full name is what Open WebUI sends to Azure, and it takes the speech language from the first two parts of that name (`en-US` from `en-US-JennyNeural`). If you type a voice instead of picking one, use the full name. Which voice is used for a reply: 1. The model's own **TTS Voice**, if one is set when editing the model in **Workspace** → **Models** 2. Otherwise, the voice the user picked under **Set Voice** in their personal **Audio** settings, as long as the admin's **TTS Voice** is still the one that was set when they picked it. If the admin changes it, users hear the new default until they pick again. 3. Otherwise, the admin's **TTS Voice** For the full list of voices and languages, see Microsoft's [Language and voice support](https://learn.microsoft.com/azure/ai-services/speech-service/language-support?tabs=tts). ## Endpoint URL By default, Open WebUI sends requests to `https://.tts.speech.microsoft.com`, built from **Azure Region**. **Endpoint URL** replaces that address. Open WebUI adds `/cognitiveservices/v1` to it to request speech and `/cognitiveservices/voices/list` to load the voice list, so enter only the scheme and host, with no path. This is not your resource's endpoint from **Keys and Endpoint**, the one the [STT guide](/features/chat-conversations/audio/speech-to-text/azure-stt-integration#endpoint-url) can use. Microsoft documents the voice list at `/tts/cognitiveservices/voices/list` on that address, while Open WebUI requests `/cognitiveservices/voices/list`, so the voice list would not load. For the sovereign clouds, Microsoft documents these addresses, with `` replaced by your region identifier: | Cloud | Endpoint URL | | --- | --- | | Azure Government | https://.tts.speech.azure.us | | Azure operated by 21Vianet | https://.tts.speech.azure.cn | See [Speech service in sovereign clouds](https://learn.microsoft.com/azure/ai-services/speech-service/sovereign-clouds) for the region identifiers. ## Output Format **Output format** is sent to Azure as the audio format to return. Keep one of the MP3 formats, such as the default `audio-24khz-160kbitrate-mono-mp3`. Open WebUI stores the audio it receives as an `.mp3` file and, except in the slim image, serves it as MP3 whatever format Azure returned. The slim image refuses audio types outside its list of browser-playable formats, such as MP3, WAV and Ogg. Microsoft's [audio outputs list](https://learn.microsoft.com/azure/ai-services/speech-service/rest-text-to-speech#audio-outputs) has every value. ## Environment Variables Setup If you prefer to configure via environment variables: ``` services: open-webui: image: ghcr.io/open-webui/open-webui:main environment: - AUDIO_TTS_ENGINE=azure - AUDIO_TTS_API_KEY=your-speech-resource-key - AUDIO_TTS_AZURE_SPEECH_REGION=westeurope - AUDIO_TTS_VOICE=en-US-JennyNeural # ... other configuration ``` Set `AUDIO_TTS_VOICE` to an Azure voice name. Its default, `alloy`, is an OpenAI voice. ### All Azure TTS Environment Variables | Variable | Description | Default | | --- | --- | --- | | AUDIO_TTS_ENGINE | Set to azure | empty (uses browser-only TTS) | | AUDIO_TTS_API_KEY | Your Speech resource key (shared with the ElevenLabs engine) | empty | | AUDIO_TTS_AZURE_SPEECH_REGION | Region identifier of your Speech resource | empty (speech uses eastus; without a base URL, the voice list does not load) | | AUDIO_TTS_AZURE_SPEECH_BASE_URL | Replaces https://.tts.speech.microsoft.com | empty | | AUDIO_TTS_AZURE_SPEECH_OUTPUT_FORMAT | Audio output format | audio-24khz-160kbitrate-mono-mp3 | | AUDIO_TTS_VOICE | Full Azure voice name | alloy (not an Azure voice) | > **Settings saved in the UI take precedence** > > Saving the admin **Audio** settings stores every field on that tab, empty ones included, and choosing an engine saves too. With the default `ENABLE_PERSISTENT_CONFIG=true`, those stored values take precedence over the `AUDIO_*` environment variables from then on. See [`ENABLE_PERSISTENT_CONFIG`](/reference/env-configuration#enable_persistent_config). ## Testing TTS 1. Start a new chat 2. Send a message to any model 3. Click the **speaker icon** on the AI response to hear it read aloud ## Troubleshooting ### No voices shown in the list 1. Confirm **API Key** and **Azure Region** are set and saved 2. Switch to another settings tab and back to **Audio** to reload the list 3. Check the Open WebUI logs for `Error fetching Azure voices` ### Speech fails with a 401 error Azure rejected the key. Check that the key is copied correctly from **Keys and Endpoint** and that **Azure Region** is the region your Speech resource is in. ### Speech fails with another error 1. Confirm **TTS Voice** is set to a full Azure voice name from the list 2. Confirm **Output format** is one of the MP3 values from Microsoft's [audio outputs list](https://learn.microsoft.com/azure/ai-services/speech-service/rest-text-to-speech#audio-outputs) 3. Check the Open WebUI logs for the full error For broader audio debugging, see the [Audio Troubleshooting Guide](/troubleshooting/audio). ## Cost Considerations Azure bills Speech usage under your Azure subscription. Check [Azure's Speech pricing page](https://azure.microsoft.com/pricing/details/cognitive-services/speech-services/) for current rates.