Running Ollama on Proxmox LXC: The Setup That Actually Works (2026)
Proxmox + Ollama — running Ollama on Proxmox LXC, the setup that actually works in 2026
Most of the guides for running Ollama on Proxmox are either outdated, targeted at Docker setups, or gloss over the LXC-specific issues that will bite you. This is the setup I actually run, with the specific configuration choices that took me a few hours of debugging to land on.
Why LXC Over a Full VM?
For a service like Ollama that you're going to run persistently and want to keep lightweight, LXC makes more sense than a full VM. You get:
- Lower overhead: no hypervisor layer, direct kernel access, faster startup
- Easier resource adjustment: CPU and RAM limits can be changed live without shutting down the container
- Cleaner integration with Proxmox storage for model files
- Simpler backup with Proxmox's built-in vzdump
The tradeoff: GPU passthrough in LXC is possible but more complex than in VMs. If you're running a dedicated GPU for inference, a VM is easier to configure. For CPU-only inference on good hardware, LXC is the better choice.
This guide covers CPU-only Ollama on LXC. GPU passthrough for LXC is a separate post.
Hardware
My current Ollama LXC runs on a node with:
- CPU: AMD Ryzen 9 5950X (16 cores / 32 threads)
- RAM: 128GB DDR4
- Storage: NVMe for OS, spinning rust for model files (2TB)
I allocate 8 cores and 32GB RAM to the Ollama LXC. On this hardware I can run mistral-7b for inference in a few seconds and llama3-8b comfortably. For 70B models you need more RAM.
Creating the LXC Container
I use a Debian 12 template. In the Proxmox web UI:
- Download the Debian 12 LXC template from the template library if you don't have it
- Create a new CT with these settings:
- Unprivileged container: No (we need this for some bind mounts)
- CPU: 8 cores
- Memory: 32768 MB
- Swap: 4096 MB
- Disk: 32GB on SSD for OS; we'll mount model storage separately
Or via CLI on the Proxmox host:
pct create 200 /var/lib/vz/template/cache/debian-12-standard_12.2-1_amd64.tar.zst \
--arch amd64 \
--cores 8 \
--memory 32768 \
--swap 4096 \
--hostname ollama-lxc \
--net0 name=eth0,bridge=vmbr0,ip=dhcp \
--storage local-lvm \
--rootfs local-lvm:32 \
--unprivileged 0 \
--features nesting=1
Note --unprivileged 0 — we're running as a privileged container. This has security implications if you're in a multi-tenant environment. For a homelab or dedicated MSP infrastructure, it's fine.
Mounting Model Storage
Model files are large (7B models are 4–8GB, 70B models are 40–80GB). Don't store them on your SSD root. Mount a separate volume.
In /etc/pve/lxc/200.conf, add:
mp0: /mnt/storage/ollama-models,mp=/var/lib/ollama/models
This bind-mounts my NAS path into the container. Proxmox will mount it when the container starts.
Installing Ollama
Start the container, then:
pct start 200
pct enter 200
# Inside the container
apt-get update && apt-get upgrade -y
apt-get install -y curl wget
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
The install script creates a systemd service and a dedicated ollama user. Verify:
systemctl status ollama
● ollama.service - Ollama Service
Loaded: loaded (/etc/systemd/system/ollama.service; enabled)
Active: active (running)
Configuring Ollama for LXC
By default Ollama listens on 127.0.0.1:11434. In an LXC container with a network interface, you want it to listen on all interfaces so you can reach it from your Proxmox host or other containers:
Edit /etc/systemd/system/ollama.service:
[Unit]
Description=Ollama Service
After=network-online.target
[Service]
ExecStart=/usr/bin/ollama serve
User=ollama
Group=ollama
Restart=always
RestartSec=3
Environment="PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
Environment="OLLAMA_HOST=0.0.0.0:11434"
Environment="OLLAMA_MODELS=/var/lib/ollama/models"
[Install]
WantedBy=multi-user.target
The key additions:
OLLAMA_HOST=0.0.0.0:11434— listen on all interfacesOLLAMA_MODELS=/var/lib/ollama/models— explicit model path pointing to our mounted volume
Reload and restart:
systemctl daemon-reload
systemctl restart ollama
Pulling Models
# 7B models — good starting point for CPU inference
ollama pull mistral
# 8B models — good balance of quality and speed
ollama pull llama3.1
# Coding-specific
ollama pull codellama
# Check what's pulled
ollama list
Testing the API
From your Proxmox host or any machine on the same network:
LXC_IP="192.168.1.200" # replace with your container IP
curl http://${LXC_IP}:11434/api/generate -d '{
"model": "mistral",
"prompt": "What is 2+2?",
"stream": false
}'
Response should come back within a few seconds on the hardware spec above.
Automating Health Checks
Add a simple health check that alerts you if Ollama stops responding:
#!/bin/bash
# /usr/local/bin/ollama-healthcheck.sh
OLLAMA_URL="http://localhost:11434"
MAX_RESPONSE_TIME=10 # seconds
RESPONSE=$(curl -s --max-time ${MAX_RESPONSE_TIME} "${OLLAMA_URL}/api/tags" 2>&1)
EXIT_CODE=$?
if [ ${EXIT_CODE} -ne 0 ]; then
logger -t ollama-health "CRITICAL: Ollama not responding (exit code: ${EXIT_CODE})"
# Add your notification method here — email, Zabbix sender, etc.
exit 1
fi
MODEL_COUNT=$(echo "${RESPONSE}" | python3 -c "import sys, json; data=json.load(sys.stdin); print(len(data.get('models', [])))" 2>/dev/null)
logger -t ollama-health "OK: Ollama responding, ${MODEL_COUNT} models loaded"
exit 0
Wire it into cron:
# /etc/cron.d/ollama-health
*/5 * * * * root /usr/local/bin/ollama-healthcheck.sh
Performance Notes
On my hardware, rough inference times for interactive use:
- mistral (7B, Q4): ~8 tokens/sec — comfortable for interactive use
- llama3.1 (8B, Q4): ~6 tokens/sec — comfortable for interactive use
- llama3.1 (70B, Q4): Would need more RAM than I've allocated. Don't try.
For batch processing rather than interactive use, queue requests and let them run at whatever pace the hardware supports.
Common Issues
"cannot allocate memory" errors: You've run out of RAM for the model. Either reduce the model size or increase LXC memory allocation. Allocate the container at least 2x the model file size.
Slow first response: Ollama loads models into RAM on first request. The first response after a model hasn't been used for a while will be slow. This is normal — keep-alive settings help.
Model storage on wrong disk: Check your OLLAMA_MODELS path and ensure the bind mount is actually working: df -h /var/lib/ollama/models should show your storage volume, not your root disk.
Next week I'll cover the MSP "Business in a Box" concept — but the Ollama series will continue with GPU passthrough in a future post.
Matt Fitzgerald runs Fitzgerald Tech Solutions and publishes The Operator's Edge newsletter. Live Life Automated is available now.