← All chapters

CHAPTER 17 / Architecture

Build the
whole platform

The final platform is not one giant assistant. It is a set of specialised components with clear jobs, data ownership and boundaries that can grow without becoming mysterious.

Enter the chapter
Original cover of chapter 17: Building Our Open WebUI Platform

CHAPTER 17

Building Our Open WebUI Platform

Assemble the system from deliberate, observable parts.

On this page

Faithful English web edition · Original chapter, structure and illustrations from the learning guide.

LEARNING OBJECTIVES

Assemble a platform from clear responsibilities

This chapter turns the previous building blocks into one architecture: a model portfolio, curated assistants, governed Knowledge, bounded Tools, coding workspaces, Automations, APIs, security and operating discipline.

01. From installation to platform

An installation answers “can I chat with a model?” A platform also answers which model is appropriate, which sources are authoritative, which actions are allowed, how identity flows, where evidence is kept, what happens when a service fails and who owns the decision to expand capability.

02. A realistic 12 GB starting architecture

Current homelab architecture with Ollama Open WebUI Knowledge tools and ComputerEnlarge illustration ↗
Figure 1 · A compact system remains useful when responsibilities are visible.
Open WebUI
The front door for chat, assistants, Knowledge and access policy.
Ollama
The local inference runtime serving a deliberately small model portfolio.
Knowledge
Focused, owned collections close to the user experience.
Tools / MCP
External services with separate lifecycle and credentials when appropriate.
Computer
A bounded workspace for files, terminal, Git and coding agents.

Keep the first version simple enough to understand and capable enough to be useful. Do not add a new service merely because it is available.

03. Target architecture: separate responsibilities

Target architecture separating experience orchestration inference knowledge and action layersEnlarge illustration ↗
Figure 2 · Every layer has a different job, scaling path and trust boundary.
LayerResponsibilityExamples
ExperienceHuman interaction and product workflowOpen WebUI, custom Django UI
OrchestrationPrompts, models, Tools, agents and automationWorkspace Models, Functions
InferenceGenerate and embedOllama, remote providers
KnowledgeSources, extraction, retrieval and citationsCollections, vector search, storage
ActionRead or modify external systemsMCP, OpenAPI, Computer, APIs
Open WebUI model and connection settings supporting the orchestration and inference boundaryEnlarge illustration ↗
Interface · Provider administration belongs behind the assistants people select.

04. Assign models by job

Model portfolio with daily coding reasoning task embedding and vision modelsEnlarge illustration ↗
Figure 3 · Specialise by workload instead of forcing one model to cover every trade-off.
Daily assistant
Balanced quality and speed for normal conversation.
Coding model
Repository and tool-oriented work.
Reasoning model
Hard analysis where extra latency is acceptable.
Task Model
Small fast model for titles, tags and background tasks.
Embedding model
Stable vector representation for retrieval.
Vision model
Image and screenshot understanding when required.

A 12 GB server may not host all roles simultaneously. A portfolio is a logical assignment; it can mix local and remote providers and use scheduled loading rather than permanent residency.

05. Workspace Models make the portfolio usable

Open WebUI Workspace Model editor with prompts knowledge tools and capabilitiesEnlarge illustration ↗
Interface · A Workspace Model turns a raw provider ID into a purposeful assistant.

Each preset should have a clear name, short job description, focused system prompt, chosen base model, deliberate Knowledge and Tools, sensible parameters and explicit audience.

  • Daily Assistant: concise, no privileged Tools, general model.
  • Document Analyst: citation rules and one approved Knowledge collection.
  • Business Analyst: analysis Skill, project templates and read-only work Tools.
  • Coding Agent: coding model plus an isolated Computer workspace.

Opinionated presets reduce accidental complexity. A “do everything” assistant hides authority and makes failures difficult to attribute.

06. Design Knowledge as a source-of-truth system

Open WebUI Knowledge workspace with documents and collectionsEnlarge illustration ↗
Interface · Collections need owners, update paths and a known authoritative source.

Recommended collections are bounded by purpose: platform operations, active project documentation, policy/reference material and personal notes should not become one anonymous corpus.

Focused Retrieval selects passages and scales to larger corpora. Full Context is appropriate only when the complete document fits and every part matters. The index is a retrieval aid, not the authoritative document. Record where originals live, who updates them and when re-indexing occurs.

Every collection needs a small golden set of questions with known answers and expected sources. RAG belongs to evaluation, not faith.

07. Memory, Prompts and Skills solve different problems

ComponentUse it forDo not use it for
MemoryStable user preferences and durable personal factsLarge documents or frequently changing operational data
PromptReusable invocation text and structured inputHidden authority or secret storage
SkillReusable instruction, method and supporting resourcesUnreviewed code execution

Choose the smallest component that matches the persistence and behaviour required.

08. Tools and connectors reach live systems

Open WebUI Tools workspace and configured integrationsEnlarge illustration ↗
Interface · Every connected service needs an owner, credential boundary and test path.

Use OpenAPI first when an existing REST service already has a clear contract. Use MCP for a tool ecosystem designed for agents. Use Workspace Tools for small trusted extensions that genuinely belong inside the application process.

Begin read-only. Prove authentication, schema quality, observability and model selection before adding write operations. Grant one narrow credential per integration where possible and make every connector easy to disable.

For every operation, define the business owner, data classification, expected latency, retry safety and audit evidence. A technically healthy connector without ownership becomes an invisible production dependency.

09. General agents and coding workspaces

A general agent reasons over a goal and calls bounded Tools. A coding agent additionally needs a persistent repository, filesystem, terminal, tests and Git. Do not confuse chat capability with safe operating authority.

Open WebUI Computer workspace with repository editor terminal and agent conversationEnlarge illustration ↗
Interface · Coding work belongs in a real isolated workspace with objective verification.

For parallel work, give each task its own worktree or checkout and branch. Decompose by independent files and contracts; shared mutable state is not useful parallelism.

10. Automate only stable work

Good candidates have a predictable trigger, bounded inputs, repeatable decisions, narrow side effects and inspectable outputs: daily briefs, health summaries, CI triage and document freshness checks.

If the manual process changes every time, automation will encode ambiguity and repeat it faster. Stabilise the human workflow, then schedule or expose it to a webhook.

Promote autonomy in stages: manual execution, a human-triggered Action, scheduled read-only work, notification-only response and finally a narrowly defined remediation. Each stage keeps the evidence and rollback of the previous one.

11. Treat Open WebUI as an internal AI platform API

A custom Django service can call Open WebUI for model resolution, RAG and Tools while owning its own interface and business data. Keep shared API keys in the backend and map product users to explicit permissions.

browser → Django product boundary
        → Open WebUI platform API
        → model / Knowledge / Tools

This makes the web app replaceable without duplicating every AI integration, and makes providers replaceable without rewriting the product UI.

The service account should expose only the model, Knowledge and Tool set the product requires. Django still authenticates its own users and applies business rules; one shared upstream key must not silently turn every product user into the same administrator.

12. Define trust tiers

  1. Conversation: model can answer but not act.
  2. Read-only context: approved Knowledge and read-only APIs.
  3. Human-approved actions: Actions with an explicit confirmation.
  4. Bounded agent: isolated workspace and narrow write authority.
  5. Unattended automation: pre-approved environment, idempotency and monitoring.

Movement upward should require evidence, an owner and a rollback path. Prompt wording alone is not a security boundary.

13. Minimum operational discipline

  • Service, provider and deep functional health checks.
  • TTFT, total latency, errors and queue depth.
  • GPU memory, offload, disk and container restarts.
  • RAG golden tests and source verification.
  • Tool call failures and external dependency health.
  • Backup age and last successful restore drill.

Keep an operations notebook containing versions, model digests, configuration decisions, change dates, incidents and rollback instructions. It turns future troubleshooting from archaeology into engineering.

Change one layer at a time where possible. Record the expected effect, validation and rollback before an upgrade, embedding change or permission redesign. The notebook is the bridge between architecture and daily operation.

14. Scale the bottleneck you can name

BottleneckLikely response
InferenceSmaller models, shorter context, more suitable GPU or additional provider node.
RAG ingestion/searchDedicated embedding capacity, external vector service, bounded workers.
ApplicationMore workers plus externalised database and compatible shared state.
AgentsQueue, workspace isolation, concurrency limits and stronger supervision.

Scaling one layer can move the bottleneck elsewhere. Re-run end-to-end measurements after each change.

15. Build in a stable sequence

Seven implementation stages from core through knowledge actions agents automation hardening and scaleEnlarge illustration ↗
Figure 4 · Each layer earns the next by meeting a clear success test.
  1. Core: reliable authenticated chat with a known model.
  2. Knowledge: one owned collection and a passing golden set.
  3. Actions: one read-only service with clear logs.
  4. Agents: one bounded workspace and objective verification.
  5. Automation: one proven workflow with selective notification.
  6. Hardening: public path, secrets, networks, backup and restore.
  7. Scale: capacity added only for a measured bottleneck.

16. Finished-platform assistants

Personal Daily Assistant
General conversation, safe preferences and no privileged Tools by default.
Business Analyst
Project Knowledge, analysis Skill and read-only ticket or reporting Tools.
Coding Agent
Repository workspace, coding model, tests, Git and supervised approvals.
RAG Researcher
Focused sources, citation requirement and an evaluation set.
Operations Assistant
Health and logs with read-only diagnosis before remediation.

17. Final architecture checklist

  • Every component has one job and owner.
  • Every model has a workload and resource profile.
  • Every Knowledge collection has a source and test set.
  • Every Tool has a credential, network boundary and disable path.
  • Every agent has an isolated workspace and verification method.
  • Every automation has safe failure and idempotency.
  • Every public route uses the intended identity boundary.
  • Every critical state has a tested restore.
  • Every expansion begins from measured evidence.

A useful platform card also records version, last validation, next review, dependency owner and decommission condition. Architecture remains clear when obsolete components can be identified and removed.

18. Practical checkpoint

Create one platform card per component with its owner, inputs, outputs, secrets, permissions, monitoring signal and disable condition. Then choose the next unfinished layer and write the success test it must pass before anything else is added.

19. Essential vocabulary

Platform
A governed collection of reusable capabilities, not merely one application.
Experience layer
The interface and product workflow seen by people.
Inference layer
The runtimes that generate text, images or embeddings.
Action layer
Tools and agents able to affect live systems.
Trust boundary
A point where identity, authority or confidence changes.
Task Model
A smaller model assigned to background platform work.
Golden test set
A stable set of known cases used to compare behaviour.

Source and further reading

This edition preserves the chapter's teaching sequence and examples. Screenshots reflect the source edition; controls may move between releases.

Open the original chapter ↗Workspace ↗Models ↗Knowledge ↗MCP ↗Computer ↗Hardening ↗

Find your next step

Search chapter titles and section headings