Skip to main content

Task Models

Open WebUI makes small model calls in the background, separately from the answer a user is waiting for. They write chat titles, generate tags, suggest follow-up questions, produce autocomplete ghost text, rewrite a message into a retrieval or web search query, rewrite a message into an image prompt and summarize older messages when context compaction is on.

By default all of that runs on the model the user is chatting with. On an expensive flagship model it is a bill nobody planned for, and on a slow local model it makes the whole interface feel sluggish. Pointing background work at a small, fast model is the highest-value change an administrator can make here.

The settings below are in Settings > Admin > Interface.


Choosing the task model

The Tasks section holds two pickers. Open WebUI picks between them based on the connection the chat model came from:

FieldEnvironment variableUsed when the chat model comes from
Local Task ModelTASK_MODELa connection typed Local (Ollama and anything else you marked local)
External Task ModelTASK_MODEL_EXTERNALa connection typed External (OpenAI-compatible providers)

Both default to Current Model, which means background work runs on the chat's own model. If you name a model that is no longer available, background work falls back to the chat's model rather than failing.

What to pick

Small, fast and non-reasoning. A reasoning model spends seconds thinking before producing a three-word title, and you pay for those thinking tokens.

  • External Task Model: gpt-5.4-nano, gemini-3.5-flash-lite, claude-haiku-4-5-20251001.
  • Local Task Model: qwen3.5:2b, gemma4:e2b, llama3.2:3b.

The main chat experience does not change. The background work just stops dragging.


Task model parameters

Task Model Parameters > Configure, in the same Tasks section, opens the advanced parameter controls you already know from a model and applies them to every background request. The equivalent environment variable is TASK_MODEL_PARAMS, which takes a JSON object such as {"max_tokens":4000,"temperature":0.3}.

Leave it empty and background requests behave as they always have: title generation and context compaction summaries stop at 1000 output tokens, and every other background task sends no parameters at all.

Reach for it when:

  • A title or a compaction summary comes back cut off. This is the usual symptom of a reasoning task model. Thinking counts against the same 1000 tokens, so the budget can be gone before any answer is written. Raise max_tokens or turn reasoning_effort down.
  • Background calls cost more than you expected. Tag, follow-up, query and autocomplete requests carry no output limit of their own. Setting max_tokens puts a ceiling on all of them at once.
  • Background output is too loose. Lower temperature for steadier titles and tags.
Setting anything replaces the built-in limit

As soon as you set a single parameter, the built-in 1000-token limit on titles and compaction summaries is no longer applied. If you still want a limit, include max_tokens in what you set. Emoji generation is the one exception: it always asks for 4 tokens and ignores this setting.

Five of the controls in the panel have no effect on background requests: Stream Chat Response, Stream Delta Chunk Size, Function Calling, Reasoning Tags and Context Compaction Threshold. Setting them here changes nothing.

This is the only place background requests get parameters from. A user's per-chat and per-account chat parameters never reach them, and neither do the global model defaults.


Turning individual tasks off

Each background task has a switch under Generation in the same panel, and an environment variable:

TaskAdmin toggleEnvironment variable
Chat title generationTitle GenerationENABLE_TITLE_GENERATION
Tag generationTags GenerationENABLE_TAGS_GENERATION
Follow-up suggestionsFollow Up GenerationENABLE_FOLLOW_UP_GENERATION
Autocomplete (fires as users type)Autocomplete GenerationENABLE_AUTOCOMPLETE_GENERATION
Query rewriting for knowledge retrievalRetrieval Query GenerationENABLE_RETRIEVAL_QUERY_GENERATION
Query rewriting for web searchWeb Search Query GenerationENABLE_SEARCH_QUERY_GENERATION

Image prompt generation is switched off in Settings > Admin > Images instead (ENABLE_IMAGE_PROMPT_GENERATION).

Weak hardware

Autocomplete is the first thing to turn off. It fires as users type, so a slow task model turns the prompt box to molasses.


Changing the prompts

Every task's prompt sits next to its toggle under Generation, and each one is empty by default, meaning the built-in prompt is used. The Tasks section of the environment configuration reference has the full text of each built-in prompt and the placeholders it accepts.

Prompt modifiers are worth knowing about here, since task prompts often have to cope with pasted documents and long code blocks.

This content is for informational purposes only and does not constitute a warranty, guarantee, or contractual commitment. Open WebUI is provided "as is." See your license for applicable terms.