SSD Nodes Learn Hosting plans →
指南 Matt Connor作者: Matt Connor · 更新于 2026-08-31

Ollama 如何让模型常驻内存,避免重复加载

Ollama 默认空闲 5 分钟后卸载模型,下一次请求会重新读取权重并等待加载。通过 keep_alive 设置单次请求或服务器默认值,重启后仍可生效。

Ollama 为什么会在几分钟后卸载模型?

Ollama 会在最后一次请求结束后将模型保留在内存中 5 分钟,随后释放模型。下一次请求必须再次从磁盘读取权重,并将其映射到 RAM 或 VRAM,因此第一个 token 到达前会出现停顿。这就是聊天 UI 或代码代理刚开始响应很快,闲置一段时间后,在下一条消息上又变慢的原因。系统没有故障。只是空闲计时器已到期。

该计时器称为 keep_alive。它按模型单独计算,并在每次请求完成后重新计时。正在处理请求的模型不会被卸载,因为服务器只会卸载当前没有活动请求的模型。截至 2026 年 8 月,默认值为 5 分钟,并适用于此服务器加载的所有模型。

有两个位置可以设置 keep_alive:单独设置请求,或设置服务器默认值。systemd drop-in 可使服务器默认值在重启后继续生效。本指南假设 Ollama 已作为服务运行。如果尚未运行,请先阅读在 VPS 上安装 Ollama,然后返回此处。

当前驻留了哪些模型?它们何时过期?

ollama ps
NAME        ID              SIZE      PROCESSOR    CONTEXT    UNTIL
qwen3:8b    500a1f067a9f    6.6 GB    100% GPU     4096       4 minutes from now

输出为空表示当前未加载任何模型,因此下一次请求需要完整加载。PROCESSOR 显示权重所在位置。100% GPU 和 100% CPU 表示明确的情况。类似 25%/75% CPU/GPU 的拆分表示模型无法完全装入 VRAM,因此部分模型在处理器上运行,生成速度会更慢。

UNTIL 是倒计时,会打印类似 4 minutes from now 的相对时间。如果模型加载时使用了负值 keep_alive,则会打印 Forever。服务器正在卸载模型的短暂窗口期间,会打印 Stopping...。

不同版本的列集合可能发生变化,因此应读取表头,不要在脚本中通过字段数量判断。对于自动化场景,请调用 API:

curl -s http://localhost:11434/api/ps

每个条目都包含 expires_at(绝对时间戳,例如 2026-08-09T14:38:31.83753Z)和 size_vram,后者表示该模型驻留在 GPU 内存中的部分。size_vram 为 0 表示模型在 CPU 上运行。

实际重新加载的成本

不要猜测。Ollama 会在每个响应中报告加载时间,字段为 load_duration,单位是纳秒。

sudo apt install -y jq
ollama stop qwen3:8b
curl -s http://localhost:11434/api/generate -d '{"model": "qwen3:8b", "prompt": "hi", "stream": false}' | jq '{load_duration, total_duration}'
curl -s http://localhost:11434/api/generate -d '{"model": "qwen3:8b", "prompt": "hi", "stream": false}' | jq '{load_duration, total_duration}'

第一次调用会加载模型,因此其 load_duration 数值较大。将该数值除以 1000000000,即可换算为秒。第二次调用会在模型驻留内存时执行,报告的数值会小得多。这两个数值之间的差值,就是计时器过期后每个用户都要承担的等待时间,也是修改 keep_alive 的根本原因。这个差值主要来自磁盘读取,因此如果您已将模型目录 迁移到第二个卷,该卷的速度就决定了每次冷加载的最低耗时。至于暂停前后的生成速度,请参阅如何在自己的主机上测量每秒生成的令牌数。

在一次请求中将 Ollama 模型保持加载状态

在请求中发送 keep_alive。从请求完成时起,该设置将应用于此模型。

curl -s http://localhost:11434/api/chat -d '{
  "model": "qwen3:8b",
  "messages": [{"role": "user", "content": "hello"}],
  "keep_alive": "30m"
}'

支持以下 4 种值形式:

  • 时长字符串:"30m"、"24h"、"90s"
  • 普通数字,按秒读取:3600
  • 负值 -1 或 "-1m",表示完全不设置空闲超时
  • 0,表示请求完成后立即卸载

请求中的值会覆盖服务器默认值,双向覆盖均适用。这一点非常重要:客户端发送的 keep_alive 会优先于服务器上的任何配置。

您也可以在不生成任何内容的情况下加载模型。只发送模型名称即可。服务器会加载该模型,并返回包含 "done": true 的空响应。

curl -s http://localhost:11434/api/generate -d '{"model": "qwen3:8b", "keep_alive": "30m"}'

重启后或拉取新模型后运行此命令。这样,首个实际用户请求就无需等待模型加载。CLI 使用一个 flag 执行相同操作:

ollama run --keepalive 30m qwen3:8b "hello"

使用 OLLAMA_KEEP_ALIVE 默认保持模型加载

服务器启动时读取 OLLAMA_KEEP_ALIVE,并将其用于所有未单独设置该值的模型。它支持与请求字段相同的格式,因此 30m、3600 和 -1 均可使用。

关键在于该环境变量必须设置在哪个环境中。在 SSH 会话中运行 export OLLAMA_KEEP_ALIVE=30m 不会产生任何效果,因为通过软件包安装的服务会以独立用户身份作为 systemd 服务运行,并使用自己的环境变量。您的登录 shell 与该服务互不相通。这是该设置看似未生效的最常见原因。

通过 systemd drop-in 配置使其在重启后保持设置

sudo systemctl edit ollama.service

编辑器打开时会显示两个注释标记。请在这两个标记之间输入内容:systemd 会丢弃你写在第二个标记下方的所有内容。

[Service]
Environment="OLLAMA_KEEP_ALIVE=30m"

保存后会写入 /etc/systemd/system/ollama.service.d/override.conf。这是一个 drop-in 文件,不是对已发布 unit 文件的直接修改。因此,Ollama 软件包升级并替换 ollama.service 后,你的设置仍会保留。如果你不熟悉 drop-in 和 unit 文件,systemd 服务和定时器指南介绍了相关机制。

sudo systemctl daemon-reload
sudo systemctl restart ollama
systemctl show ollama --property=Environment

最后一条命令会显示该服务实际运行时使用的环境变量。如果这一行缺少 OLLAMA_KEEP_ALIVE=30m,说明 drop-in 未生效。最常见的原因是缺少 [Service] 标头,或将内容写在了标记下方。重启会卸载所有已加载的模型,因此下一次请求需要冷启动。使用上面的预加载调用预热模型。

What keeping a model resident costs you

The SIZE column in ollama ps is memory held for the whole idle window, not only during a request. An 8B model at 4-bit quantisation sits around 5 to 6 GB. A 27B model is a different conversation, and the memory arithmetic for running one on a CPU-only VPS is worth working through before you decide to hold one resident. Set keep_alive to -1 and you have decided that the model outranks everything else on the box, permanently. On a small VPS that is a direct trade against your database, your web app and your build jobs.

Watch the real numbers rather than trusting an estimate. Run this while a model is loaded, then again after ollama stop:

free -h

The available column is the memory the kernel could still hand to a new process. On an NVIDIA GPU box, nvidia-smi shows the same story in VRAM. If the box runs out, the kernel kills a process to recover:

sudo dmesg -T | grep -i "out of memory"

A line naming ollama means the model server was the victim. A line naming your database means the model won and something you cared about lost. Both outcomes come from the same decision: a long keep-alive window on a box with no headroom.

Two costs here are easy to miss. A longer context length reserves a larger KV cache (key value cache, the per token attention state the model keeps while generating), and that cache is part of the resident size. How big it gets follows from num_ctx, so raising the context window raises the memory a resident model holds for the whole idle period, not only while it is answering. OLLAMA_NUM_PARALLEL above 1 reserves that cache once per parallel slot. If you plan to serve several people from one model, size the memory for the slots, not for the weights alone.

A reasonable default: one model on a box with headroom can use -1. A shared box should use a window that covers the gaps between your requests, such as 30m, so the memory comes back when you stop working.

立即卸载模型

ollama stop qwen3:8b

该命令不返回任何输出,模型会从 ollama ps 中消失。指定未加载的名称时,会返回 couldn't find model "qwen3:8b" to stop。API 形式是不包含提示词的请求,并将 keep_alive 设置为 0:

curl -s http://localhost:11434/api/chat -d '{"model": "qwen3:8b", "messages": [], "keep_alive": 0}'

响应会包含 "done_reason": "unload"。请使用此方法,不要重启服务。systemctl restart ollama 也会释放内存,但会卸载所有其他已加载的模型,并终止正在运行的请求。

在一台服务器上运行多个模型

OLLAMA_MAX_LOADED_MODELS限制同时保持加载状态的模型数量。截至 2026 年 8 月,默认值是每个 GPU 3 个;仅使用 CPU 的服务器默认也是 3 个。该限制按模型数量计算,但真正的限制是内存。因此,第二个大型模型可能在达到 3 个之前很久就因没有足够空间而被拒绝加载。

请求加载新模型时,如果可用内存不足,调度器会卸载一个当前驻留的模型以释放空间。它会优先选择没有活动请求的模型,也会驱逐计时器尚未到期的模型,包括使用 -1 加载的模型。因此,负值 keep_alive 表示不设置空闲超时。它不会将模型权重固定在内存中,以阻止其他模型发起请求。

该决策会以 debug 级别记录。向同一个 drop-in 文件再添加一行 Environment="OLLAMA_DEBUG=1",重启服务,然后监控日志:

sudo journalctl -u ollama -f

如果在触发该操作的请求旁边看到“卸载 runner 以释放空间”的日志,说明这两个模型无法在这台机器上同时运行。解决方法是减少这台服务器上的模型数量,或者为必须快速响应的模型设置较长的窗口,并为很少调用的模型设置 0。

可持续到下一版本的指导

Ollama 经常发布新版本,默认值也会变化。因此,应检查当前使用的构建版本,而不要死记数值:

ollama --version
ollama serve --help

ollama serve --help 会列出该构建实际读取的环境变量,其中包括 OLLAMA_KEEP_ALIVE。有两条规则在各版本中都成立,可以放心据此配置。请求中的值优先于服务器默认值。无论配置文件声明应加载什么,ollama ps 才是实际加载内容的依据。

如果编辑器或代理驱动您的服务器,请先检查客户端发送的内容,再判断是否是服务器的问题。将编码代理指向您自己的 Ollama 服务器介绍了这些请求设置的位置。

FAQ

Ollama 为什么会在 5 分钟后卸载模型?

5 分钟是默认的 keep_alive,即 Ollama 在请求完成后启动的空闲计时器。计时器到期后,服务器会释放模型权重,因此下一次请求必须从磁盘重新加载权重;这段加载过程就是您感受到的暂停。您可以在 JSON 请求体中发送 "keep_alive": "30m",为单个请求延长保留时间;也可以通过 OLLAMA_KEEP_ALIVE 环境变量为整个服务器设置该值。

如何让 Ollama 模型永久保留在内存中?

使用负值:在请求中设置 "keep_alive": -1,或为服务器设置 OLLAMA_KEEP_ALIVE=-1。此时,ollama ps 会在 UNTIL 列中显示 Forever。这只会移除空闲计时器,不会改变其他行为。如果请求加载另一个模型时内存不足,调度器仍会卸载当前模型以腾出空间。

为什么 OLLAMA_KEEP_ALIVE 没有生效?

检查您设置它的位置。运行 systemctl show ollama --property=Environment。如果输出中没有该变量,说明服务器根本没有接收到它,因为在 shell 中导出的变量不会传递给 systemd 服务。使用 sudo systemctl edit ollama.service 设置它,然后运行 sudo systemctl daemon-reload 和 sudo systemctl restart ollama。另一种原因是客户端在请求中发送了自己的 keep_alive,从而覆盖了服务器默认值。

如何在不重启 Ollama 的情况下释放内存?

ollama stop qwen3:8b 会立即卸载指定的一个模型,同时保持服务器和其他已加载模型继续运行。通过 API 发送不带提示词且包含 "keep_alive": 0 的请求,响应会返回 "done_reason": "unload"。使用 ollama ps 确认,输出中不应再列出该模型。