← All chapters

CHAPTER 10 / Agents

Make a plan.
Then verify it

An agent is not a model with a dramatic prompt. It is a loop that interprets a goal, plans, acts through tools, observes the result and stops under defined conditions.

Enter the chapter
Original cover of chapter 10: Agents

CHAPTER 10

Agents

Build multi-step work that remains inspectable and bounded.

On this page

Faithful English web edition · Original chapter, structure and illustrations from the learning guide.

LEARNING OBJECTIVES

Build an agent as a system, not a dramatic prompt

By the end, you should be able to describe the agent loop, distinguish chat history from memory and external state, use structured tasks, decide when delegation helps, design approval and stop conditions, and assemble a bounded Open WebUI agent from Models, Skills, Knowledge and Tools.

01. What is an agent — really?

A language model receives context and generates a continuation. By itself it does not open GitHub, call an API, run tests or remember yesterday. An agent appears when a harness surrounds the model with a goal, instructions, Tools, state, permissions, an action environment and a loop.

Agent =
model + goal + instructions + tools
+ state + loop + guardrails

The model proposes a next step. The system executes it, returns evidence and lets the model decide again. That is closer to a worker following feedback than to a single response.

Agent does not mean autonomous forever

Two searches followed by an answer are already agentic. Editing twenty files and repairing failed tests is more autonomous. A scheduled workflow that begins without a waiting user adds another dimension. Ask four concrete questions: can it choose actions, observe results, preserve state and continue for multiple steps?

02. The spectrum from chatbot to autonomous workflow

Spectrum from chatbot through tool user and agent to autonomous workflowEnlarge illustration ↗
Figure 1 · Agentic behaviour is a spectrum; each additional freedom creates new failure modes as well as capability.
SystemBehaviour
ChatbotAnswers from supplied context.
Tool-enabled assistantCan call an external capability.
Agentic assistantCan chain calls and adapt to their results.
AgentCan plan, maintain state and pursue a multi-step goal.
AutomationCan start or continue work without a human prompting every step.

03. The agent loop: goal, plan, act, observe, repeat

Agent loop with goal, plan, action, observation, state update and stop conditionsEnlarge illustration ↗
Figure 2 · Observation changes the next decision; the plan is not executed blindly.

“Investigate the failed deployment, fix it and verify the tests” becomes a sequence: inspect the repository, reproduce the failure, observe the error, locate its cause, edit, rerun checks, revise the hypothesis if necessary and stop only after verification.

state = understand_goal(request)
while not state.done:
    action = model_decides_next_action(state)
    if action is a tool_call:
        result = execute_tool(action)
        state = update_state(state, result)
    else:
        state.final_answer = action.response
        state.done = True

Real systems also enforce schemas, permissions, approvals, retries, budgets and streaming. The key idea remains: an action produces new evidence, and evidence determines the next action.

04. Where the plan lives

Implicit
The model selects actions without keeping a formal list. Efficient for a short search-and-answer task.
Visible
The model writes a plan in the conversation. Humans can review it, but its states are still prose.
Structured
A system tool stores tasks with explicit pending, in-progress, completed or cancelled states.

Long work becomes less reliable when the plan exists only inside transient reasoning. A structured checklist does not make the model smarter; it makes omissions visible and preserves the operational state outside a paragraph.

05. State: how an agent remembers the work

Diagram showing current-turn context, chat state, long-term memory and external system stateEnlarge illustration ↗
Figure 3 · Different state belongs in different sources of truth.

State is information that must survive between steps: inspected tables, a backup created, a failed command and remaining tasks. Some belongs in the chat, some in structured tasks and some in the real system.

State is not context window

State is everything retained by the system. The context window is only what the model can consider during one inference. A long-running agent may have extensive state but must retrieve, summarise and select what enters each prompt.

Choose the source of truth

If the agent created ticket INC123, ServiceNow owns its actual status. “I think it was created” in chat is not authoritative; the agent should query the system again.

06. Chat history, Memory and external state

StoreBest usePoor use
Chat historyImmediate continuity and recent tool results.An unlimited permanent database.
Long-term MemoryDurable preferences and facts across conversations.Temporary transactional state such as “test 17 is failing now.”
External systemGit branches, tickets, database records, files and calendars.Facts copied into chat and never checked again.

Open WebUI Memory can add, search, update and delete durable facts through native tool calling. Use it for stable preferences such as a project language or documentation location, while leaving live state in the system that owns it.

07. Native tool calling powers agentic behaviour

Open WebUI chat controls showing Native function callingEnlarge illustration ↗
Chat Controls determine the current conversation’s model and enabled capability surface.

Native mode passes structured Tool schemas to the provider and receives structured calls. Modern built-ins depend on this path; Legacy prompt-parsed calls are not a reliable foundation for agent loops.

Depending on configuration and access, built-in categories include Memory, Notes, Chat History, Channels, Task Management, Automations, Knowledge, Web Search, Image Generation, Code Interpreter and Sub-agents, alongside Workspace, MCP, OpenAPI and Terminal Tools.

Open WebUI Workspace Model capabilities configurationEnlarge illustration ↗
Workspace Models package a base model with selected built-in capabilities instead of exposing everything.

A model that can call one Tool is not necessarily capable of a ten-step workflow. Planning, result interpretation, error recovery and state maintenance must also hold across the loop.

08. Task Management: visible multi-step progress

create_tasks(tasks=[...])
update_task(id="task-id", status="completed")

Task states are pending, in_progress, completed and cancelled. They are stored with the conversation, so a temporary chat cannot provide the same persistent list.

  • Create tasks only for genuinely multi-step work.
  • Keep one primary task in progress unless parallelism is intentional.
  • Mark completion after verification, not after intention.
  • Cancel obsolete work instead of leaving it pending.
  • Update the plan when evidence changes it.

The checklist is an operational window into the agent’s state, not decorative status text.

09. Skills, Knowledge and Tools

Skill     = how to work
Knowledge = what to know
Tool      = what it can do

A Skill can teach a code-review method. Knowledge can supply architecture and project conventions. Tools can read GitHub and execute tests. Keeping these roles separate produces a system that is easier to maintain than one enormous prompt that tries to encode method, facts and actions together.

10. Sub-agents and delegation

Parent agent delegating focused tasks to sub-agents and combining resultsEnlarge illustration ↗
Figure 4 · Delegation isolates focused investigations and lets the parent synthesise final output.

The built-in delegate_task Tool lets a parent give a bounded task to another model execution. The sub-agent works in a separate conversation, using the same model, Tools, Skills and Filters, then returns its final response.

Open WebUI Sub-agents configuration interfaceEnlarge illustration ↗
Sub-agents are explicitly enabled because every delegation adds cost, concurrency and operational complexity.

Delegation helps when

  • A smaller, clearer problem improves focus.
  • Independent investigations can proceed in parallel.
  • Intermediate detail should not pollute the parent context.
  • The parent needs to compare specialist results.

It is not free. Every delegation is another full model run and may trigger several calls. If one agent can solve the task clearly, that is usually easier to operate.

11. Open Terminal: a real action workspace

Open WebUI connected to Open Terminal with chat, files and execution workspaceEnlarge illustration ↗
Open Terminal gives an agent a filesystem, shell, processes, packages and observable outputs.

A model without execution can suggest a patch. An agent with Terminal can inspect a repository, edit, run pytest or npm test, read the failure and correct it. Verification closes the feedback loop.

Docker isolation and direct host access are different security decisions. Mount only necessary paths, use development credentials and decide whether persistence is useful enough to justify its larger blast radius.

12. Automations: agentic work without a waiting user

An automation answers “When does work start?” An agent answers “How does the system choose and execute steps?” A schedule may trigger one simple prompt or wake a tool-enabled multi-step workflow.

schedule triggers work + agent loop performs work

Automations belong to their creator and administrators can limit count or frequency. A scheduled task must be safe to run without an interactive person available for clarification or approval.

13. Human approval, guardrails and stop conditions

Experimental Tool Permissions can pause calls for Allow or Deny. Denial returns to the model as an error so it can adapt. This is valuable for sending mail, deleting files, deploying, closing tickets or modifying records.

LayerControl
PromptDescribe intent and operating rules.
Tool schemaExpose only valid operations and arguments.
API/serverValidate identity, values and permissions.
EnvironmentRestrict files, network and credentials.
Human approvalGate costly or irreversible actions.
Audit logsMake actions inspectable afterwards.

Every loop also needs a success condition and operational limits: maximum retries, timeout, token or Tool-call budget, and user cancellation. “Keep trying until it works” is not a safe policy.

14. Build an agent in Open WebUI

Workspace → Models can wrap a base model with a system prompt, parameters, Knowledge, Tools and Skills. That preset becomes a purpose-built agent without requiring a new framework.

Technical Support Agent

  1. Choose a model with reliable native tool calling.
  2. Define scope and escalation in a concise system policy.
  3. Attach runbooks and architecture as Knowledge.
  4. Attach a troubleshooting Skill.
  5. Add read-only ticket, health and log Tools.
  6. Enable Task Management for longer investigations.
  7. Test one-Tool cases before multi-step incidents.
  8. Add narrow writes only after read-only behaviour is reliable.
You are a Technical Support Agent.
Use evidence, not guesses.
Inspect documented and live state.
Plan multi-step investigations visibly.
Prefer read-only tools.
Verify every proposed fix.
Ask before destructive actions.
State what is missing when evidence is insufficient.

15. External agents behind Open WebUI

Open WebUI can also be the frontend for an agent framework exposed through an OpenAI-compatible endpoint. That external system brings its own loop, terminal, memory or session logic while Open WebUI provides users, chat and a shared interface.

A provider returns a completion. An agent endpoint may take many actions before returning. They can look similar in the model picker but have different operational ownership. This pattern fits specialised runtimes, session/resume semantics or capabilities that should not live inside the Open WebUI backend.

16. Evaluate the process, not only the final prose

Goal completion
Did the requested outcome actually occur?
Tool selection
Did the agent choose the correct capability?
Arguments
Were parameters valid, bounded and safe?
State tracking
Did it retain completed and pending work?
Recovery
Did new evidence change the next step?
Verification
Did it prove success rather than assume it?
Efficiency
How many model and Tool calls were needed?
Safety
Did it remain inside approval and permission boundaries?

Build a benchmark of 10–20 representative tasks and retain Tool calls, unnecessary steps, recovered errors, latency and required human interventions — not only pass/fail.

17. Troubleshooting agent failures

FailureResponse
Never calls ToolsConfirm Native mode; test one trivial Tool with a capable model.
Chooses the wrong ToolReduce the catalogue and remove overlapping names/descriptions.
Bad argumentsUse types, enums and server-side validation.
Forgets progressUse a saved chat, Task Management and less noisy context.
LoopsAdd success signals, structured errors and retry budgets.
Claims success without proofAdd an explicit verification Tool or read-back step.
Local model fails long workflowsReduce complexity or use a stronger orchestration model.

18. A future local agent stack

  • Open WebUI: shared control plane, users, chats and permissions.
  • Ollama / remote providers: models chosen by task.
  • Workspace Models: specialised assistants.
  • Knowledge: long-lived reference material.
  • Skills: reusable working methods.
  • MCP / OpenAPI: live services and business systems.
  • Open Terminal: isolated execution and verification.
  • Task Management: visible plans.
  • Memory: selected durable preferences.
  • Sub-agents: focused delegation when it truly helps.
  • Automations: scheduled work with safe unattended behaviour.
The goal is not one magical model. It is a platform where every layer has a clear responsibility.

PRACTICAL CHECKPOINT

Can you trace a complete, bounded agent loop?

  • What does the harness add around the LLM?
  • Why is observe → decide as important as plan → act?
  • Where should task state, durable Memory and real-world state live?
  • When does delegation help?
  • Which stop and verification conditions prove the work is complete?
Check your answers

The harness supplies Tools, state, execution, permissions and control. Observations update the plan. Tasks belong in chat state, durable preferences in Memory and transactions in their external source. Delegate focused independent work. Stop on verified success or explicit budgets, not on model confidence.

19. Essential vocabulary

Agent
A system that pursues a goal through decisions, Tools, state and execution loops.
Agent harness
The context, permissions, runtime, feedback and controls around a model.
Agent loop
Decide → act → observe → update → repeat.
State
Working information preserved between steps.
Memory
Durable, retrievable information across conversations.
Sub-agent
A focused delegated model execution returning a result to a parent.
Tool approval
A human gate before an operation executes.
Stop condition
A success or limit criterion that ends the loop.
Automation
Scheduled work initiated without a user waiting in the chat.

Source and further reading

This edition preserves the chapter's teaching sequence and examples. Screenshots reflect the source edition; controls may move between releases.

Open the original chapter ↗Workspace Models ↗Task Management ↗Sub-agents ↗Open Terminal ↗

Find your next step

Search chapter titles and section headings