Every time you send a message to ChatGPT, Claude, or Gemini, that text travels to a data center you have no visibility into. For most casual questions, that trade-off is fine. For anything involving business data, personal notes, client information, or financial details, it is not.
A private LLM solves the problem by running the language model entirely on your own machine or server. Nothing leaves your infrastructure. No API logs, no training data contributions, no third-party retention policies to audit.
This guide covers what a private LLM actually is, how to set one up, which models work best, and how OpenClaw turns a local private LLM into a fully capable AI assistant you can use every day.
What Is a Private LLM?
A private LLM is a large language model that runs inference locally on your own hardware. Instead of making an API call to OpenAI or Anthropic, your application talks to a model process running on your server, laptop, or VPS.
The key properties of a genuinely private LLM setup:
- No data leaves the host -- prompts, responses, and context stay on your machine
- No per-token billing -- you pay for hardware or a VPS, not for each query
- No rate limits -- inference is only limited by your compute
- Full model control -- you choose the weights, the version, the system prompt
This is different from "private mode" on a cloud AI product, which typically just disables training data use. Your data still transits their network. A self hosted private LLM never does.
How to Run LLM Privately: The Stack
The most practical route to a private LLM setup in 2026 is Ollama. It is a free, open-source tool that manages model downloads, serves a local HTTP API on port 11434, and works on Linux, macOS, and Windows.
Step 1: Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
That single command installs Ollama and the system service. It starts automatically on boot.
Step 2: Pull a Model
# Lightweight and capable -- good starting point
ollama pull llama3.1
# Smaller footprint for low-RAM servers
ollama pull phi3
# Strong coding ability
ollama pull qwen2.5-coder
Models are stored locally in ~/.ollama/models. Once downloaded, they run entirely offline.
Step 3: Test It
ollama run llama3.1 "Explain what a private LLM is in one paragraph"
If you see a response, your local private LLM is working. The model ran entirely on your hardware.
Hardware Requirements for a Private LLM Server
You do not need a GPU-powered workstation to run a private AI LLM. Modern quantized models run acceptably on CPU-only hardware:
| Model Size | Minimum RAM | CPU Inference Speed | Example Models |
|---|---|---|---|
| 3B parameters | 4 GB | Fast | Phi-3 Mini, Gemma 2 2B |
| 7B parameters | 8 GB | Moderate | Llama 3.1 8B, Mistral 7B |
| 13B parameters | 16 GB | Slow on CPU | Llama 3.1 13B, CodeLlama 13B |
| 70B parameters | 48 GB+ | GPU recommended | Llama 3.1 70B, Qwen2.5 72B |
A 4-8 GB VPS is enough for a 7B model. Hostinger VPS plans in that range cost roughly $7-12 per month, making a private LLM server accessible without significant hardware investment.
Which Models Work Best for a Private LLM Setup
Not all open-weight models are equal. These are the strongest choices for a local private LLM in 2026:
Llama 3.1 8B
Meta's flagship open model. Excellent general capability, strong instruction following, and a 128K context window. The 8B version runs comfortably on an 8 GB server and outperforms older 13B models. This is the recommended starting point for most private AI LLM deployments.
Phi-3 Mini (3.8B)
Microsoft's surprisingly capable small model. Fits on 4 GB RAM, responds fast, and handles everyday tasks well. Good choice if your server is tight on resources or you want low-latency responses for a chatbot.
Qwen 2.5 Coder
Alibaba's coding-focused model. If your private LLM server is mainly for writing and reviewing code, Qwen 2.5 Coder consistently outperforms Llama on programming tasks.
Mistral 7B
One of the earliest high-quality open models, still competitive. Particularly efficient at following structured outputs and JSON instructions, which makes it useful for automation workflows.
Connecting OpenClaw to Your Local Private LLM
OpenClaw is an open-source AI assistant platform designed to run on a self hosted private LLM or connect to cloud models. Pointing it at your local Ollama instance takes one config line:
# In your OpenClaw gateway config
model: ollama/llama3.1
With that change, every message you send through OpenClaw goes to your local model. No tokens reach Anthropic, OpenAI, or any other cloud provider.
OpenClaw adds the layer that raw Ollama lacks: persistent memory, scheduled tasks, Telegram and Discord integration, custom skills, and agentic workflows. The combination gives you a complete AI assistant that is both private and genuinely useful.
For the full installation walkthrough, see the guide on installing OpenClaw on a VPS.
Private LLM vs Cloud API: When to Use Each
A private LLM is not always the right answer. Here is a straightforward comparison:
Choose a private LLM when:
- Your prompts contain client data, financial records, or personal health information
- You send high query volumes where API costs would be significant
- You operate in a regulated industry with data residency requirements
- You want zero dependency on external service availability
Stick with cloud APIs when:
- You need the absolute best reasoning quality (frontier models still lead open weights)
- Your server has limited RAM and latency matters
- You are processing tasks that only run occasionally
- You rely on multimodal features like image analysis
OpenClaw handles both: set it to a local Ollama model for private work and switch to a cloud model for tasks where you want maximum capability. The multi-model guide explains how to route different jobs to different models.
Security Considerations for a Private LLM Server
Running Ollama on a VPS exposes port 11434 by default. That port should not be public. Bind Ollama to localhost only:
# In /etc/systemd/system/ollama.service
Environment="OLLAMA_HOST=127.0.0.1"
Then reload and restart:
sudo systemctl daemon-reload
sudo systemctl restart ollama
OpenClaw running on the same server communicates over localhost, so it still reaches the model. Nothing outside the server can.
For additional hardening, configure a firewall to block the port externally and use SSH tunnels if you need remote access to the Ollama API for development.
Getting Started: Your Private LLM Checklist
- Provision a VPS with at least 8 GB RAM (16 GB recommended for comfortable headroom)
- Install Ollama with the one-line script
- Pull Llama 3.1 8B or Phi-3 Mini as your first model
- Bind Ollama to localhost for security
- Install OpenClaw on the same server
- Set
model: ollama/llama3.1in your OpenClaw config - Connect OpenClaw to Telegram or Discord for daily use
Total setup time is around 30 minutes. Monthly cost is the VPS fee only, no API charges.
Summary
- A private LLM runs inference on your own hardware -- no data leaves your server
- Ollama is the simplest way to run LLM privately on Linux or a VPS
- 7B models like Llama 3.1 8B work well on an 8 GB server; smaller models fit 4 GB
- OpenClaw connects to Ollama with a single config line and adds memory, automations, and multi-channel access
- Bind Ollama to localhost and keep port 11434 off the public internet
Ready to Run Your Own Private LLM?
Install OpenClaw on your VPS and connect it to a local private LLM in under 30 minutes. Full control, zero cloud dependency.
Install OpenClaw Free →