← All chapters

CHAPTER 15 / Operations

Observe first.
Then change

Troubleshooting is a method: reproduce, isolate, observe, change one thing and verify. A layered system becomes manageable when you can locate the first boundary that fails.

Enter the chapter
Original cover of chapter 15: Observability & Troubleshooting

CHAPTER 15

Observability & Troubleshooting

Diagnose the stack without collecting random fixes.

On this page

Faithful English web edition · Original chapter, structure and illustrations from the learning guide.

LEARNING OBJECTIVES

Diagnose the first broken boundary

You will turn a vague failure into a reproducible symptom, move through a layered diagnostic ladder, collect useful logs, separate GPU and context problems, trace RAG as a pipeline and write an incident record another operator can follow.

01. Troubleshooting is a method

Evidence-driven troubleshooting loop from capture through verify and documentEnlarge illustration ↗
Figure 1 · Capture evidence, isolate a layer, change one thing and verify the result.

Random changes erase evidence. Begin with the exact time, user, URL, model, input, expected result, observed result and error. Determine scope: one browser, one user, one model, one provider, one document or the whole service.

Reproduce before changing anything. A repeatable failure is a controlled experiment; an intermittent one needs timestamps and conditions. Change one variable, run the same reproduction and record the result.

02. The diagnostic ladder

Diagnostic ladder from browser proxy Open WebUI provider GPU and storageEnlarge illustration ↗
Figure 2 · Start at the user-visible boundary and stop at the first failing layer.
  1. Browser/UI: console, request, response and account.
  2. Proxy/network: routing, TLS, streaming and reachability.
  3. Open WebUI: health, logs, permissions and configuration.
  4. Provider/Ollama: direct model list and minimal generation.
  5. GPU/system: VRAM, offload, disk and process pressure.
  6. Storage/data: permissions, database, extraction and vector state.
“Everything is broken” becomes useful only after you identify the first boundary that does not satisfy its contract.

03. Healthy means usable

A process can be running while its model provider or retrieval path is failing. Use three levels:

Liveness
The Open WebUI process responds.
Connectivity
The application can reach and enumerate the configured provider.
Deep check
A controlled prompt completes through the real path, optionally including retrieval or one Tool.

Keep the deep check small and predictable. It should prove service, not consume production-scale inference.

Run probes from the same route as their consumer. A host-level curl can pass while the container network fails; an internal URL can pass while the public HTTPS proxy is broken. Label each probe with the boundary it proves.

04. Logs are evidence, not decoration

The browser console exposes frontend exceptions, blocked requests and CORS. Proxy logs show public status and upstream timing. Container logs show application exceptions; provider logs show model load and inference failures.

docker compose ps
docker compose logs --since=15m open-webui
docker compose logs --since=15m ollama

Use the least verbose log level that captures the event, raising it temporarily for a reproduction. Structured JSON helps central processing. Audit logs answer who changed what; operational logs answer why a request failed. Neither should contain raw secrets.

Align clocks and timezones before correlating sources. A browser timestamp, proxy access line and provider error that appear several minutes apart may describe the same request if one component logs UTC and another local time.

05. Containers and deployed reality

docker compose ps
docker inspect open-webui --format '{{json .State}}'
docker inspect open-webui --format '{{json .Config.Env}}'

Restart loops often follow invalid configuration, a missing mount, failed migration or resource pressure. Verify the actual container environment, image tag, networks, mounts and command — not the Compose file you believe was deployed. Redact credentials before sharing output.

Do not begin with restart. Capture exit code, previous logs and resource state first. A restart may temporarily remove the symptom while destroying the evidence needed to explain it.

06. Ollama and model connectivity

Open WebUI connection settings for Ollama and OpenAI-compatible providersEnlarge illustration ↗
Interface · A configured provider is only useful when it is reachable from the Open WebUI network namespace.
curl http://ollama:11434/api/tags

curl http://ollama:11434/api/generate \
  -d '{"model":"<MODEL>","prompt":"Reply with OK","stream":false}'

Test the provider from the same network context as Open WebUI. localhost inside one container does not mean another container or the host. If models disappear intermittently, correlate provider restarts, network errors, configuration changes and UI cache rather than repeatedly deleting connections.

Use the smallest direct generation that still loads the failing model. If direct generation succeeds but Open WebUI fails, the fault lies above the provider: identity, model mapping, connection configuration or request enrichment.

07. GPU, VRAM, context and slow models

Decision chart explaining why a local model becomes slowEnlarge illustration ↗
Figure 3 · Model size, KV cache, context, concurrency and CPU offload have different signatures.
ollama ps
nvidia-smi

ollama ps identifies loaded models, context and processor split. A model partly offloaded to CPU can fit yet run slowly. Longer context expands KV-cache pressure; parallel requests multiply memory use. Disk size describes stored weights, not the entire runtime footprint.

nvidia-smi is supporting evidence: GPU memory, utilisation and competing processes. Compare it with provider state and the exact request rather than treating one number as the diagnosis.

08. Prompt and context failures

Open WebUI model advanced parameters and capabilitiesEnlarge illustration ↗
Interface · Model defaults can change context and behaviour without an obvious prompt edit.

System prompt, conversation history, retrieved chunks, tool schemas and the planned answer all share the context budget. An oversized request may fail explicitly, become slow, lose older instructions or leave too little room for useful generation.

When follow-ups degrade, test the same question in a new chat. Compare the Workspace Model, per-model defaults and provider context settings. This separates accumulated history from a general model problem.

09. Debug RAG as a pipeline

RAG pipeline from document extraction chunking embeddings retrieval context and answerEnlarge illustration ↗
Figure 4 · Inspect the earliest broken RAG stage before tuning later stages.
  1. Extraction: is the expected text present?
  2. Chunking: are related facts kept together with useful metadata?
  3. Embeddings: did the model load and index complete?
  4. Retrieval: does the known question return the known passage?
  5. Context assembly: did that passage fit into the prompt?
  6. Generation: did the model follow evidence and citation instructions?
Open WebUI Knowledge collection with uploaded filesEnlarge illustration ↗
Interface · Begin with one small document and a question whose supporting sentence you already know.

A fluent answer is not proof of retrieval. Verify sources. If retrieval fails, changing the generator’s temperature cannot repair it.

Keep one golden document and question whose supporting sentence is known. Run it after extractor, embedding-model, chunking or reranker changes. Without a fixed case, each new test introduces new uncertainty.

10. Tools, MCP and authentication

Separate connection from intelligence. First invoke the service directly with known arguments. Then test the Open WebUI connection and authentication. Only after both work should you evaluate whether the model chooses the Tool correctly.

  • Test from Open WebUI’s network namespace.
  • Verify API token, OAuth owner and expiry.
  • Inspect Tool name, description and schema.
  • Use a model capable of reliable tool calling.
  • Make the prompt’s need for the Tool explicit during diagnosis.
Open WebUI error shown after a failed MCP or tool requestEnlarge illustration ↗
Interface · Preserve the exact error and request context before changing credentials or schemas.

11. Browser, proxy, CORS and streaming failures

If the shell request works but the browser does not, inspect the browser’s network panel and console. A UI that loads while generation never appears often points to buffering, SSE/WebSocket forwarding, timeout or mixed-content policy at the proxy.

A CORS error is a browser origin decision, not a generic network failure. Compare scheme, host and port exactly. Test both the canonical proxy URL and the internal direct route to identify which boundary introduces the failure, but do not leave the direct route exposed afterwards.

12. Database, storage and uploads

Terminal and system output used to diagnose disk and storage stateEnlarge illustration ↗
Interface · Disk, inode, permission and volume evidence belongs beside the application error.
df -h
df -i
docker inspect open-webui --format '{{json .Mounts}}'

Check free bytes and inodes, volume ownership and write permission. Multiple application workers require compatible shared state; local vector-store assumptions may not hold across workers. Large document ingestion can spike memory and temporary disk even when steady-state usage looks safe.

13. Read the shape of slowness

Observed shapeLikely investigation
Slow first token, then fastCold model load, prompt processing, retrieval or queue time.
Every token slowCPU offload, model size, quantisation, GPU saturation.
Fast for one user, poor under loadConcurrency, VRAM multiplication and queue policy.
Only long chats degradeContext and KV-cache growth.

Measure time to first token, token rate, total latency, queue time and error rate. One “it feels slow” measurement cannot distinguish them.

14. Monitoring and alerts

Monitor service health, request error rate, latency, restarts, CPU/RAM/GPU, disk, model load failures and retrieval or Tool errors. Alerts need owners and actions. A noisy alert teaches operators to ignore the system; alert on symptoms that require intervention.

Prefer service-level signals over isolated hardware numbers: “deep check failed three times” is more actionable than “GPU reached 95%” during healthy generation. Keep dashboards for investigation and alerts for conditions that demand a response.

15. Repeatable incident checklist

  • Record time, account, request and exact symptom.
  • Determine scope and recent changes.
  • Reproduce with the smallest case.
  • Walk the diagnostic ladder.
  • Capture relevant logs before restarting.
  • State one hypothesis and one change.
  • Run the same reproduction.
  • Roll back if evidence does not improve.
  • Document root cause, fix and prevention.

16. Worked diagnoses

No models appear

Confirm Open WebUI health, inspect the configured provider URL, call /api/tags from the application network and verify permissions. Do not reinstall the UI before testing the provider boundary.

A model suddenly becomes very slow

Record the exact model, prompt and concurrency; inspect ollama ps and GPU state; compare a short fresh chat. Look for context expansion, a second loaded model or CPU offload.

RAG invents an answer that exists in the PDF

Search extracted text, inspect the chunk containing the answer, test retrieval with a known phrase and confirm the chunk entered the final context. Fix the earliest failing stage.

MCP is connected but unused

Call the server directly, verify the user’s auth, inspect schema descriptions, then test a forced simple prompt with one Tool and a capable model.

17. Practical checkpoint

  1. Write one failure as expected versus observed behaviour.
  2. Name its scope and the first three ladder checks.
  3. Capture a correlated browser, Open WebUI and provider timestamp.
  4. Make one reversible change and repeat the same test.
  5. Write a short incident note with cause, evidence and prevention.

18. Essential vocabulary

Liveness
Proof that a process responds.
Deep check
A functional probe that traverses several real layers.
Reproduction
The smallest stable sequence that causes a failure.
Scope
The users, components or conditions affected.
Offload
Running some model work outside the GPU.
Time to first token
Delay before the first generated content reaches the user.
Root cause
The underlying condition that explains the symptom.

Source and further reading

This edition preserves the chapter's teaching sequence and examples. Screenshots reflect the source edition; controls may move between releases.

Open the original chapter ↗Troubleshooting ↗Logging ↗Monitoring ↗RAG troubleshooting ↗Performance ↗

Find your next step

Search chapter titles and section headings