The mainstream AI tools all have one thing in common: your data goes through someone else's servers. Every message you send to ChatGPT, Claude, or Gemini is processed on infrastructure you do not control, stored under terms you may not have read, and potentially used to improve models you will pay to use again.
A self-hosted LLM changes that equation completely. You run the model. Your hardware. Your data. No API bills, no usage logs, no third-party terms of service between your thoughts and your AI.
In 2026, self-hosting a capable language model is genuinely accessible. You do not need a data center or a PhD. This guide covers what you need, how to set it up with Ollama, which models to pick, and how to wire it into OpenClaw so your self-hosted LLM has memory, automations, and a full personal assistant interface.
What Is a Self-Hosted LLM?
A large language model (LLM) is the AI engine behind tools like ChatGPT. When you use a cloud service, that model runs on the provider's servers and you access it via API. A self-hosted LLM is the same kind of model, but downloaded and run locally on hardware you control: a personal computer, a home server, a VPS, or a dedicated machine.
The key distinction is inference location. Self-hosted means the model weights live on your disk and run on your CPU or GPU. No request leaves your machine unless you specifically ask it to.
Why Self-Host an LLM?
There are four reasons people move to self-hosted setups:
- Privacy. Sensitive work, client data, personal journals, legal documents. None of it should transit third-party servers. A local LLM guarantees it does not.
- Cost. GPT-4o at $5 per million tokens adds up fast for heavy users. A self-hosted model has zero per-token cost. Hardware is a one-time or sunk expense.
- Offline access. Planes, remote areas, internet outages. A local model works without a connection.
- Control. You choose the model version. You are not subject to sudden price hikes, rate limits, or capability changes made by the provider.
The trade-off is hardware. A capable self-hosted LLM needs RAM and ideally a GPU. But smaller models run surprisingly well on consumer hardware, as you will see below.
Hardware: What You Actually Need
The most common question about self-hosted LLMs is whether your hardware is good enough. The answer depends on which model you want to run.
| Model Size | Example Models | Min RAM (CPU) | GPU VRAM (fast) | Quality |
|---|---|---|---|---|
| 3B params | Phi-3 Mini, Qwen2.5-3B | 4 GB | 4 GB | Good for simple tasks |
| 7B params | Llama 3.1 8B, Mistral 7B | 8 GB | 6 GB | Solid everyday use |
| 14B params | Qwen2.5-14B, Phi-4 | 16 GB | 12 GB | Strong reasoning |
| 32B params | Qwen2.5-32B, DeepSeek-R1 | 32 GB | 24 GB | Near-frontier quality |
| 70B params | Llama 3.3 70B | 64 GB | 48 GB | Frontier-class |
CPU inference works without a GPU, but it is slow. A 7B model running on CPU-only will respond at a conversational pace but not instantly. For comfortable use, 16 GB of system RAM with a modest GPU (RTX 3060 or better) gives you a 14B model at reasonable speed. An Apple Silicon Mac Mini with 24 GB unified memory is a popular self-hosted LLM server because the GPU and CPU share that memory pool.
Setting Up Ollama: The Standard Self-Hosted LLM Server
Ollama is the de facto tool for running local LLMs. It handles model downloading, quantization, serving, and provides an OpenAI-compatible API endpoint that most AI tools (including OpenClaw) can connect to out of the box.
Install Ollama
On Linux or macOS, one command handles the install:
curl -fsSL https://ollama.com/install.sh | sh
On Windows, download the installer from ollama.com. After installation, Ollama runs as a background service on port 11434.
Pull a Model
Download a model with the ollama pull command. For a solid starting point:
# A fast 8B model, good general-purpose
ollama pull llama3.1
# A coding-focused model from Alibaba
ollama pull qwen2.5-coder:7b
# A reasoning model with chain-of-thought
ollama pull deepseek-r1:7b
Ollama automatically downloads the right quantized version for your hardware. Models are stored in ~/.ollama/models and reused across sessions.
Test It Works
ollama run llama3.1
# Type a message and press Enter to chat directly in the terminal
To verify the local LLM server API is running:
curl http://localhost:11434/api/generate \
-d '{"model":"llama3.1","prompt":"Hello","stream":false}'
If you get a JSON response with a text field, your self-hosted LLM is running.
Best Local LLM Models to Run in 2026
The open-source LLM landscape moves fast. These are the top picks across different use cases:
General Use: Llama 3.3 70B
Meta's flagship open model. Best overall quality if your hardware can handle it. Needs 64 GB RAM for CPU inference or a high-VRAM GPU. If you have the hardware, nothing open-source beats it for general tasks.
Best Balance: Qwen2.5-14B
Alibaba's Qwen2.5 series punches well above its weight. The 14B model is outstanding for its size. It handles long context (up to 128K tokens), follows instructions well, and runs on 16 GB RAM. The best self-hosted LLM for most people who cannot run 70B.
Coding: Qwen2.5-Coder 7B or 32B
Specifically trained for code. Outperforms Llama at coding tasks despite smaller size. The 7B version runs on nearly any machine with 8 GB RAM.
Reasoning: DeepSeek-R1
A distilled reasoning model. Produces structured chain-of-thought before answering. Available in sizes from 1.5B to 70B. The 7B version is excellent for problem-solving tasks on modest hardware.
Speed: Phi-3 Mini or Gemma 2 2B
When you need fast responses and have limited hardware (4 GB RAM), these small models are surprisingly capable for simple queries, quick lookups, and low-latency automations.
Connecting Your Self-Hosted LLM to OpenClaw
Running a local model in the terminal is useful. Connecting it to OpenClaw makes it a full personal AI assistant with memory, skills, automations, Telegram integration, and scheduled tasks.
Because Ollama exposes an OpenAI-compatible API, OpenClaw connects to it directly. No adapter needed.
In your OpenClaw gateway config, set the model string to your Ollama model:
model: ollama/llama3.1
# or
model: ollama/qwen2.5:14b
# or
model: ollama/deepseek-r1:7b
Make sure the base URL points to your Ollama instance:
OPENAI_BASE_URL=http://localhost:11434/v1
OPENAI_API_KEY=ollama # Ollama ignores this value but it must be set
Then restart the gateway:
openclaw gateway restart
OpenClaw will now route all inference to your self-hosted LLM. Every conversation, memory read, skill execution, and cron job runs through your local model. No data leaves your machine. See the full guide on running a local AI with OpenClaw and Ollama for step-by-step details.
Running Your Self-Hosted LLM on a VPS
If you want a self-hosted LLM that is always on without leaving a personal computer running, a VPS is the answer. A cloud server running Ollama gives you 24/7 availability, remote access, and no local hardware noise or heat.
For a 7B model, you need at least 8 GB of RAM (16 GB recommended). For a 14B model, 32 GB. Most VPS providers offer these specs. See the guide on installing OpenClaw on a VPS for the full walkthrough, which covers running Ollama alongside OpenClaw on the same server.
The advantage of a VPS self-hosted LLM setup: your AI assistant is running remotely, you interact with it through Telegram or a web interface, and your local machine stays clean. The self-hosted part refers to the server ownership, not physical proximity.
Private LLM vs Cloud LLM: A Quick Comparison
| Factor | Self-Hosted LLM | Cloud LLM (OpenAI, Anthropic) |
|---|---|---|
| Data privacy | Complete. Nothing leaves your server. | Data goes to provider's servers. |
| Cost | One-time hardware or VPS monthly fee. | Per-token billing. Scales with usage. |
| Model quality | Behind frontier, closing fast. | Best available models. |
| Offline use | Yes. | No. |
| Setup effort | Moderate (one-time). | Minimal (just an API key). |
| Customization | Full. Fine-tune if needed. | Limited to provider options. |
Many OpenClaw users run a hybrid setup: a self-hosted LLM for routine tasks and private data, and a cloud model (Claude or GPT-4o) for complex reasoning that needs frontier-class capability. OpenClaw supports switching models per cron job, so you can automate this routing. See the multi-model AI guide for how that works.
Common Self-Hosted LLM Issues and Fixes
Model is slow
CPU inference is inherently slower than GPU. Try a smaller model first (7B instead of 14B). If you have a GPU, confirm Ollama is using it with ollama ps which shows the current model and which layers are on GPU.
Out of memory errors
The model is too large for your available RAM or VRAM. Pull a smaller version: ollama pull qwen2.5:7b instead of qwen2.5:14b. Ollama's quantized models (Q4 format) use roughly half the memory of full precision.
OpenClaw is not connecting to Ollama
Verify Ollama is running: curl http://localhost:11434 should return a response. Check that your OPENAI_BASE_URL matches where Ollama is listening. If OpenClaw runs on a different server from Ollama, replace localhost with the Ollama server's IP.
Summary
- A self-hosted LLM runs on your own hardware with no third-party API calls or data sharing
- Ollama is the easiest way to run local LLMs: one install, then
ollama pullany model - 7B models run on 8 GB RAM; 14B models need 16 GB; 70B needs 64 GB or high-VRAM GPU
- Top picks for 2026: Qwen2.5-14B (general), Qwen2.5-Coder (coding), DeepSeek-R1 (reasoning)
- OpenClaw connects to Ollama with a single model config line and OpenAI-compatible base URL
- VPS hosting gives you 24/7 uptime for a self-hosted LLM without leaving personal hardware running
Run OpenClaw with Your Own Private LLM
Install OpenClaw on a VPS or home server, connect it to Ollama, and get a full personal AI assistant with memory and automation. Your data never leaves your infrastructure.
Install OpenClaw Free →