← All chapters

CHAPTER 16 / Operations

Fit the model.
Respect the budget

Performance is a system property. The model, context, hardware, provider runtime, RAG pipeline and concurrent workload all contribute to the time a person experiences.

Enter the chapter
Original cover of chapter 16: Performance & Optimization

CHAPTER 16

Performance & Optimization

Improve speed by measuring the right constraint first.

On this page

Faithful English web edition · Original chapter, structure and illustrations from the learning guide.

LEARNING OBJECTIVES

Tune the system, not one impressive number

You will measure latency and throughput, budget VRAM for weights and context, recognise CPU offload, choose a keep-alive and concurrency policy, tune RAG separately and build a fair benchmark for a 12 GB-class GPU.

01. Performance is a system property

The experience is produced by browser, network, Open WebUI, retrieval, model loading, prompt processing, generation, Tools and concurrent users. Optimisation therefore means choosing a trade-off, not turning a universal “fast” setting on.

A faster response that becomes unreliable, loses evidence or blocks every other user is not an improvement.

02. Measure what a person actually experiences

Performance metrics including time to first token token rate latency throughput VRAM and errorsEnlarge illustration ↗
Figure 1 · Speed is several measurements, not one stopwatch.
Time to first token
Waiting before visible generation begins: queue, cold load, retrieval and prompt processing.
Tokens per second
Generation rate after output begins.
End-to-end latency
Total time from send to complete answer.
Throughput
Total work completed across all users per unit of time.
Error rate
Reliability under the tested operating profile.

Record all five. A high token rate can coexist with a painful cold start; low single-user latency can coexist with collapsing throughput.

03. Build a VRAM budget

VRAM budget composed of model weights KV cache runtime buffers parallel requests and headroomEnlarge illustration ↗
Figure 2 · Weights are only the first claim on GPU memory.
VRAM ≈ model weights
     + KV cache
     + runtime buffers
     + parallel request state
     + safety headroom

Downloaded file size is not runtime use. Quantisation controls weight footprint, while context, batch behaviour and concurrency alter the rest. Leave headroom for spikes and other GPU processes.

ollama ps
nvidia-smi

Verify the actual processor split. A model that “fits” by offloading many layers to CPU can still miss its latency target.

04. Context length and KV cache

Every input token must be processed and its attention state may consume KV-cache memory. System instructions, history, RAG passages, Tool schemas and planned output share the same budget.

Agents are context-hungry because observation and tool results accumulate across rounds. A large advertised context window is a ceiling, not a recommended default. Trim old history, retrieve only relevant passages and reserve enough space for the answer and Tool loop.

When a tool schema or retrieved passage is injected, the user may not see its size. Inspect the resolved request or use controlled comparisons: fresh chat versus long history, Tool disabled versus enabled, and small versus large retrieval Top K.

05. CPU offload: fits does not mean fast

When the GPU cannot hold the full workload, Ollama may execute some layers with CPU and system RAM. This can prevent an out-of-memory failure but move inference across a much slower path.

EvidenceMeaning
Partial CPU/GPU split in ollama psThe runtime is offloading model work.
High RAM and CPU, modest GPU utilisationThe CPU path may dominate.
Short context works, long context slows or failsKV-cache growth is exhausting headroom.

Possible responses are a smaller model, stronger quantisation, shorter context, fewer parallel requests or more suitable hardware.

06. Cold starts and keep-alive

Comparison of model cold start and keep-alive request pathsEnlarge illustration ↗
Figure 3 · Keeping a model warm trades lower first-token time for occupied memory.

A cold request loads model weights before inference. Keep-alive leaves the model resident for later requests. Longer is not automatically better: a rarely used warm model may block a more important model from loading.

Open WebUI or Ollama keep-alive configuration exampleEnlarge illustration ↗
Configuration · Choose residence time from usage patterns and memory pressure.

Use a modest keep-alive for frequently used models, monitor evictions and cold starts, and unload deliberately before a heavy workload when required.

Ollama keep-alive and runtime state exampleEnlarge illustration ↗
Evidence · Runtime state confirms whether a model is still resident.

07. Concurrency, queues and throughput

Concurrency sweet spot comparing one two four and too many parallel requestsEnlarge illustration ↗
Figure 4 · Parallelism helps until memory and compute contention dominate.

One request may have excellent latency while four simultaneous requests increase total throughput. Beyond the hardware’s useful parallelism, each request slows, memory pressure rises and failures appear.

A queue is not inherently bad. A bounded queue protects the runtime and makes overload visible. Start with one or two parallel generations on a 12 GB-class GPU, measure, then increase by one while observing latency, throughput, VRAM and errors.

Choose the target from the product need. An interactive assistant values low tail latency; overnight batch work may accept slower individual requests to complete more total work. Publish queue limits and return a controlled overload response instead of letting requests fail unpredictably.

08. Multiple loaded models compete

A general model, coding model and embedding model may each fit independently but not together. Unplanned model switching can create repeated load/evict cycles that feel like random slowness.

Prefer one primary generation model per constrained GPU. If possible, move embeddings to CPU, another GPU or an external service. Separate latency-sensitive chat from background jobs when their resource profiles conflict.

09. Know which setting is authoritative

Open WebUI model parameters and advanced runtime settingsEnlarge illustration ↗
Interface · A Workspace Model or connection default may override what you expected from Ollama.

Provider environment variables, API request parameters, connection settings and per-model presets can all influence context and behaviour. Choose one source of truth for each setting and document overrides.

Per-model presets are valuable when deliberate: a fast daily assistant can use conservative context while a research profile accepts more latency for a larger evidence window.

Record the effective model ID, digest, context and parallelism with each benchmark. Otherwise a silent UI override can make two apparently identical tests exercise different runtimes.

10. RAG and embedding performance

Open WebUI Knowledge and embedding configurationEnlarge illustration ↗
Interface · Retrieval has its own model, batching, storage and context costs.

Document ingestion performs extraction, chunking, embedding and persistence. Chat-time retrieval performs query embedding, search, optional reranking and context injection. Measure these separately from generation.

  • Move embeddings off the generation GPU when they contend.
  • Limit asynchronous embedding concurrency.
  • Use chunks large enough for meaning but small enough for selective retrieval.
  • Do not retrieve more passages than the answer needs.
  • Cache stable extraction and embeddings, not user-specific secrets blindly.

Optimise for evidence quality as well as milliseconds. Reducing Top K or reranking may improve latency while removing the passage required for a correct answer. Keep retrieval recall in the benchmark.

11. Cache the expensive stable work

Repeated prompts may benefit from provider-level prompt caching, but application caches must include every input that changes the answer: model, system prompt, Knowledge version, Tool state, user permissions and sampling parameters.

RAG system context can be large and stable; retrieval results may not be. Cache document extraction and embeddings aggressively when content is immutable, while keeping permission-aware retrieval fresh.

12. Multiple Ollama servers

Open WebUI connections screen with multiple Ollama serversEnlarge illustration ↗
Interface · Multiple provider connections add capacity only when routing and health are explicit.

Scaling up gives one server more capability. Scaling out adds provider instances and distributes requests. The second approach needs health checks, model availability consistency and a routing policy.

Remove or disable dead connections quickly. A stale backend can turn load balancing into intermittent failure. Decide whether all nodes host the same portfolio or whether requests are routed by job.

13. Scale Open WebUI deliberately

Application workers also consume memory and require shared state. Horizontal scaling may require external database, shared or compatible storage, consistent secret configuration and careful handling of background work. Static assets can be cached at the proxy, but dynamic chat and streaming must not be buffered incorrectly.

14. Sensible 12 GB starting profile

  • Choose one primary quantised model that fully or mostly fits.
  • Use a measured context target instead of the maximum.
  • Keep one generation model warm; unload infrequent alternatives.
  • Start at one parallel generation, then test two.
  • Move embeddings away from the constrained GPU where possible.
  • Reserve VRAM headroom and bound queues.
  • Use a smaller Task Model for background work.

This is a baseline, not a promise. Your exact model architecture, quantisation, context and OS load determine the safe profile.

For Victor’s 12 GB-class target, the important test is not whether a model launches. It is whether the chosen quantisation, normal context and expected concurrency remain mostly GPU-resident with enough headroom for stable daily use.

15. Benchmark correctly

Use a stable test matrix: short chat, long-context chat, one RAG question with known evidence, one Tool call and a small concurrent run. Hold model digest, quantisation, parameters and prompts constant.

RunTTFTToken rateTotalPeak VRAMResult quality
Short / coldRecordRecordRecordRecordPass/fail
Short / warmRecordRecordRecordRecordPass/fail
Long contextRecordRecordRecordRecordPass/fail
Two concurrentRecordRecordRecordRecordError count

Warm up intentionally or label cold runs. Repeat each case. A single anecdotal prompt is not a benchmark.

16. Common anti-patterns

  • Maximising context “just in case.”
  • Keeping every model resident forever.
  • Increasing parallelism without observing memory.
  • Comparing different prompts or model revisions.
  • Calling CPU offload a successful fit without a latency target.
  • Blaming the model for retrieval or queue delay.
  • Optimising speed while ignoring quality and error rate.

17. Practical checkpoint

  1. Record a cold and warm request separately.
  2. Inspect actual GPU/CPU split and peak memory.
  3. Compare one short and one long-context prompt.
  4. Test one and two simultaneous requests.
  5. Change exactly one setting and rerun the same matrix.

18. Essential vocabulary

TTFT
Time from request submission until the first token arrives.
KV cache
Attention state retained for tokens in the active context.
Cold start
The cost of loading a non-resident model.
Keep-alive
How long a model remains resident after use.
Concurrency
The number of active requests at the same time.
Headroom
Capacity intentionally left free for stability and spikes.
Benchmark
A controlled repeatable comparison with relevant variables held constant.

Source and further reading

This edition preserves the chapter's teaching sequence and examples. Screenshots reflect the source edition; controls may move between releases.

Open the original chapter ↗Ollama context ↗Ollama concurrency ↗Open WebUI performance ↗Scaling ↗RAG troubleshooting ↗

Find your next step

Search chapter titles and section headings