CHAPTER 00
Roadmap & Architecture
A working mental model for the system you are about to build.
On this page
Faithful English web edition · Original chapter, structure and illustrations from the learning guide.
LEARNING OBJECTIVES
Build the map before touching the controls
By the end of this chapter, you should be able to explain what Open WebUI does, where the model actually runs, how a base model differs from a configured assistant, how Knowledge differs from a Tool, and why MCP or OpenAPI services can live outside Open WebUI.
01. What Open WebUI is — and what it is not
Open WebUI is an interface and orchestration platform for working with AI models. It centralises conversations, connected models, knowledge, tools, permissions and extensions in one experience. It can talk to Ollama, commercial providers and services that expose OpenAI-compatible APIs.
The most important fact is also the simplest: Open WebUI is not the model. It does not ship with language models of its own. Installing the interface and obtaining an LLM are separate responsibilities.
Open WebUI is the cockpit. The LLM is the engine. Ollama can be one of the hangars where those engines are stored and run.
This separation means one chat can use a local model served by Ollama, another can call a remote API, and a third can use a purpose-built assistant. They can coexist behind the same interface.
A web-development analogy
Think of a Django application that coordinates PostgreSQL, Redis, an email service and a payment provider. Django does not become those services; it integrates them. Open WebUI plays a comparable coordinating role for inference, document retrieval and executable tools.
02. The platform layers
Separating the layers prevents terms such as model, Knowledge, Function, Tool, MCP, Pipe and RAG from collapsing into one vague idea.
Enlarge illustration ↗- Experience
- The browser, chats, model picker, mobile interface, users and presentation controls.
- Orchestration
- Open WebUI keeps conversation state and composes instructions, history, retrieved context, tool definitions and generation parameters.
- Inference
- The provider loads an LLM and generates tokens. Locally this may be Ollama using your CPU or GPU; remotely it may be a commercial API.
- Knowledge
- Documents and collections are extracted, indexed and retrieved so relevant evidence can enter the model's context.
- Actions
- Tools query APIs, search services, access data or trigger automations. They extend the system beyond text generation.
03. What happens when you send a message
Imagine asking: “Review these requirements and tell me which endpoints my API needs.” The visible chat hides a sequence of distinct operations.
- The browser sends the message to Open WebUI.
- Open WebUI identifies the selected base model or Workspace Model and loads its instructions, parameters and attached capabilities.
- If Knowledge is connected, relevant passages may be retrieved from your documents and added to the context.
- If compatible tools are available, their schemas are supplied so the model can request a structured action.
- Open WebUI sends the prepared context to the model provider.
- The provider performs inference. Ollama uses local hardware; a remote provider performs the work on its own infrastructure.
- The provider streams tokens back or asks for a tool call. Open WebUI coordinates any additional step and displays the result.
04. The core building blocks
Base model
The base model is the LLM that produces the response. It has its own context window, speed, reasoning quality, tool-calling behaviour and memory requirements. Open WebUI cannot add a capability the underlying model fundamentally lacks.
Workspace Models: presets and configured agents
The Models area in Workspace creates reusable configurations over a base model. A configuration can combine a system prompt, generation parameters, Knowledge, Tools and Skills. It behaves like a specialised assistant without changing the model's weights.
The same base model could power a Python Tutor, a Requirements Analyst and a Code Reviewer. Their instructions and connected resources differ; the underlying weights may be identical.
Knowledge and RAG
Knowledge organises documents and collections. Retrieval-Augmented Generation does not retrain the LLM or permanently insert a document into its weights. It finds relevant information at question time and places that information in the working context:
question → search or vector representation → relevant passages
→ model context → grounded responseLater chapters examine chunking, embeddings, hybrid search, vector databases and reranking. For now, remember that RAG is a retrieval layer.
Tools, MCP and OpenAPI
A Tool exposes an operation the model may request during a conversation. Open WebUI supports native capabilities, Python-based Workspace Tools and external tool servers. MCP standardises how AI clients discover tools and resources. OpenAPI describes HTTP operations through a standard contract.
The architectural consequence matters: the code does not have to execute inside Open WebUI. Specialised services can run in separate processes or machines, with narrower credentials and clearer ownership.
Functions
Functions extend the platform more deeply than a normal tool. Pipes can appear as models, Filters can intercept messages, Actions can add interface controls and Events can react to activity. Because these components may execute server-side code, they belong to a high-trust boundary.
05. Ollama and model providers
Ollama is the starting point for local inference in this guide. Open WebUI calls its API and displays the available models in the picker. In a Docker installation, Open WebUI and Ollama may run in different containers, processes or machines; they only need a valid network route.
Browser → Open WebUI → Ollama API → loaded model → GPU / CPUThe GPU therefore does not “belong” to Open WebUI. The runtime performing inference uses it. This distinction allows the interface to remain stable while Ollama moves to a more powerful machine or additional providers are added.
06. From chatbot to agent
A basic chatbot receives text and returns text. An agentic system can decide what information it needs, use tools, observe results and continue for several steps. In Open WebUI this behaviour can emerge from a model with reliable native tool calling, clear instructions, Knowledge, state and carefully scoped Tools.
LLM + instructions + state + knowledge + tools + decision loop
= agentic behaviourNot every model that writes well is a reliable agent. A model may produce beautiful prose but choose the wrong tool, generate invalid arguments or abandon a multi-step plan. Evaluation must include operational behaviour, not only answer quality.
07. Security boundaries
Capability increases impact. A model without tools can only generate output. A model with filesystem, terminal, email or API access can affect real systems. Local execution does not make code trustworthy.
Security is part of the architecture: identity, exposed ports, application permissions, provider credentials, container isolation, data volumes, tool servers and external APIs are separate boundaries. Later chapters harden each of them.
08. The Zero to Hero path
The goal is not merely to learn an interface. It is to design a small AI platform consciously: select the right model, decide where data lives, control which actions are permitted and know how to inspect failures.
PRACTICAL CHECKPOINT
Can you draw the system from memory?
Draw User → Open WebUI → provider/model. Add Knowledge and Tools beside Open WebUI and explain what each contributes.
- Who executes the model when Open WebUI connects to Ollama?
- Is a Workspace Model necessarily a newly trained LLM?
- Does RAG change the model's weights?
- How does retrieving information differ from executing a Tool?
- Why can an MCP server run outside Open WebUI?
- What new risk appears when Python or terminal access is enabled?
Check your answers
The provider performs inference. A Workspace Model is a configuration over a base model. RAG retrieves context without changing weights. Tools perform operations. MCP separates the tool server from the client. Executable code widens the trust boundary to files, secrets, networks and the host.
09. Essential vocabulary
- LLM
- A Large Language Model that processes and generates language.
- Inference
- Running trained weights to produce an output from an input.
- Provider
- A runtime or service exposing one or more models to Open WebUI.
- Context window
- The token budget the model can consider in one request.
- Embedding
- A vector representation used for semantic comparison and retrieval.
- RAG
- Retrieval-Augmented Generation: retrieve external evidence before generating.
- Tool calling
- The model requests a structured, executable operation.
- MCP
- A protocol for exposing tools and resources to AI clients.
- Agent
- A system combining a model, state, instructions and actions across multiple steps.
Source and further reading
This edition preserves the chapter's teaching sequence and examples. Screenshots reflect the source edition; controls may move between releases.
Open the original chapter ↗Open WebUI features ↗Workspace documentation ↗