Task Models
Open WebUI makes small model calls in the background, separately from the answer a user is waiting for. They write chat titles, generate tags, suggest follow-up questions, produce autocomplete ghost text, rewrite a message into a retrieval or web search query, rewrite a message into an image prompt and summarize older messages when context compaction is on.
By default all of that runs on the model the user is chatting with. On an expensive flagship model it is a bill nobody planned for, and on a slow local model it makes the whole interface feel sluggish. Pointing background work at a small, fast model is the highest-value change an administrator can make here.
The settings below are in Settings > Admin > Interface.
Choosing the task model
The Tasks section holds two pickers. Open WebUI picks between them based on the connection the chat model came from:
| Field | Environment variable | Used when the chat model comes from |
|---|---|---|
| Local Task Model | TASK_MODEL | a connection typed Local (Ollama and anything else you marked local) |
| External Task Model | TASK_MODEL_EXTERNAL | every chat model not typed Local, including pipe and function models and models with no connection type |
Both default to Current Model, which means background work runs on the chat's own model. If you name a model that is no longer available, background work falls back to the chat's model rather than failing.
Small, fast and non-reasoning. A reasoning model spends seconds thinking before producing a three-word title, and you pay for those thinking tokens.
- External Task Model:
gpt-5.4-nano,gemini-3.5-flash-lite,claude-haiku-4-5-20251001. - Local Task Model:
qwen3.5:2b,gemma4:e2b,llama3.2:3b.
The main chat experience does not change. The background work just stops dragging.
Task model parameters
Task Model Parameters > Configure, in the same Tasks section, opens the advanced parameter controls you already know from a model and applies them to every background request. The equivalent environment variable is TASK_MODEL_PARAMS, which takes a JSON object such as {"max_tokens":4000,"temperature":0.3}.
Leave it empty and background requests behave as they always have: task calls use the task model's own max_tokens from its model editor when one is set, and 1000 output tokens when none is, and every other parameter is left out. Context compaction has its own Context Compaction Model picker under Chat (CONTEXT_COMPACTION_MODEL, default Current Model) and does not use the Local or External task model.
Reach for it when:
- A title or a compaction summary comes back cut off. This is the usual symptom of a reasoning task model. Thinking counts against the same 1000 tokens, so the budget can be gone before any answer is written. Raise
max_tokensor turnreasoning_effortdown. - Background calls cost more than you expected. Tag, follow-up, query and autocomplete requests carry no output limit of their own. Setting
max_tokensputs a ceiling on all of them at once. - Background output is too loose. Lower
temperaturefor steadier titles and tags.
As soon as you set a single parameter, the built-in 1000-token limit on titles and compaction summaries is no longer applied. If you still want a limit, include max_tokens in what you set. Emoji generation is the one exception: it always asks for 4 tokens and ignores this setting.
Five of the controls in the panel have no effect on background requests: Stream Chat Response, Stream Delta Chunk Size, Function Calling, Reasoning Tags and Context Compaction Threshold. Setting them here changes nothing.
Background requests also receive the task model's own per-model parameters from its model editor, with TASK_MODEL_PARAMS winning on conflicting keys. A user's per-chat and per-account chat parameters never reach them, and neither do the global model defaults.
Turning individual tasks off
Each background task has a switch under Generation in the same panel, and an environment variable:
| Task | Admin toggle | Environment variable |
|---|---|---|
| Chat title generation | Title Generation | ENABLE_TITLE_GENERATION |
| Tag generation | Tags Generation | ENABLE_TAGS_GENERATION |
| Follow-up suggestions | Follow Up Generation | ENABLE_FOLLOW_UP_GENERATION |
| Autocomplete (fires as users type) | Autocomplete Generation | ENABLE_AUTOCOMPLETE_GENERATION |
| Query rewriting for knowledge retrieval | Retrieval Query Generation | ENABLE_RETRIEVAL_QUERY_GENERATION |
| Query rewriting for web search | Web Search Query Generation | ENABLE_SEARCH_QUERY_GENERATION |
Image prompt generation is switched off in Settings > Admin > Images instead (ENABLE_IMAGE_PROMPT_GENERATION).
Autocomplete is the first thing to turn off. It fires as users type, so a slow task model turns the prompt box to molasses.
Changing the prompts
Most tasks' prompts sit under Generation, each empty by default, meaning the built-in prompt is used. Retrieval and web search share one Query Generation Prompt, the Image Prompt Generation Prompt sits under Generation while its toggle is on the Images tab, and the Context Compaction Prompt is under Chat. The Tasks section of the environment configuration reference has the full text of each built-in prompt and the placeholders it accepts.
Prompt modifiers are worth knowing about here, since task prompts often have to cope with pasted documents and long code blocks.