Every time you send a message to ChatGPT, Claude, or Gemini, that text travels to a data center you have no visibility into. For most casual questions, that trade-off is fine. For anything involving business data, personal notes, client information, or financial details, it is not.

A private LLM solves the problem by running the language model entirely on your own machine or server. Nothing leaves your infrastructure. No API logs, no training data contributions, no third-party retention policies to audit.

This guide covers what a private LLM actually is, how to set one up, which models work best, and how OpenClaw turns a local private LLM into a fully capable AI assistant you can use every day.

What Is a Private LLM?

A private LLM is a large language model that runs inference locally on your own hardware. Instead of making an API call to OpenAI or Anthropic, your application talks to a model process running on your server, laptop, or VPS.

The key properties of a genuinely private LLM setup:

This is different from "private mode" on a cloud AI product, which typically just disables training data use. Your data still transits their network. A self hosted private LLM never does.

How to Run LLM Privately: The Stack

The most practical route to a private LLM setup in 2026 is Ollama. It is a free, open-source tool that manages model downloads, serves a local HTTP API on port 11434, and works on Linux, macOS, and Windows.

Step 1: Install Ollama

curl -fsSL https://ollama.com/install.sh | sh

That single command installs Ollama and the system service. It starts automatically on boot.

Step 2: Pull a Model

# Lightweight and capable -- good starting point
ollama pull llama3.1

# Smaller footprint for low-RAM servers
ollama pull phi3

# Strong coding ability
ollama pull qwen2.5-coder

Models are stored locally in ~/.ollama/models. Once downloaded, they run entirely offline.

Step 3: Test It

ollama run llama3.1 "Explain what a private LLM is in one paragraph"

If you see a response, your local private LLM is working. The model ran entirely on your hardware.

Hardware Requirements for a Private LLM Server

You do not need a GPU-powered workstation to run a private AI LLM. Modern quantized models run acceptably on CPU-only hardware:

Model Size Minimum RAM CPU Inference Speed Example Models
3B parameters 4 GB Fast Phi-3 Mini, Gemma 2 2B
7B parameters 8 GB Moderate Llama 3.1 8B, Mistral 7B
13B parameters 16 GB Slow on CPU Llama 3.1 13B, CodeLlama 13B
70B parameters 48 GB+ GPU recommended Llama 3.1 70B, Qwen2.5 72B

A 4-8 GB VPS is enough for a 7B model. Hostinger VPS plans in that range cost roughly $7-12 per month, making a private LLM server accessible without significant hardware investment.

Which Models Work Best for a Private LLM Setup

Not all open-weight models are equal. These are the strongest choices for a local private LLM in 2026:

Llama 3.1 8B

Meta's flagship open model. Excellent general capability, strong instruction following, and a 128K context window. The 8B version runs comfortably on an 8 GB server and outperforms older 13B models. This is the recommended starting point for most private AI LLM deployments.

Phi-3 Mini (3.8B)

Microsoft's surprisingly capable small model. Fits on 4 GB RAM, responds fast, and handles everyday tasks well. Good choice if your server is tight on resources or you want low-latency responses for a chatbot.

Qwen 2.5 Coder

Alibaba's coding-focused model. If your private LLM server is mainly for writing and reviewing code, Qwen 2.5 Coder consistently outperforms Llama on programming tasks.

Mistral 7B

One of the earliest high-quality open models, still competitive. Particularly efficient at following structured outputs and JSON instructions, which makes it useful for automation workflows.

Connecting OpenClaw to Your Local Private LLM

OpenClaw is an open-source AI assistant platform designed to run on a self hosted private LLM or connect to cloud models. Pointing it at your local Ollama instance takes one config line:

# In your OpenClaw gateway config
model: ollama/llama3.1

With that change, every message you send through OpenClaw goes to your local model. No tokens reach Anthropic, OpenAI, or any other cloud provider.

OpenClaw adds the layer that raw Ollama lacks: persistent memory, scheduled tasks, Telegram and Discord integration, custom skills, and agentic workflows. The combination gives you a complete AI assistant that is both private and genuinely useful.

For the full installation walkthrough, see the guide on installing OpenClaw on a VPS.

Private LLM vs Cloud API: When to Use Each

A private LLM is not always the right answer. Here is a straightforward comparison:

Choose a private LLM when:

Stick with cloud APIs when:

OpenClaw handles both: set it to a local Ollama model for private work and switch to a cloud model for tasks where you want maximum capability. The multi-model guide explains how to route different jobs to different models.

Security Considerations for a Private LLM Server

Running Ollama on a VPS exposes port 11434 by default. That port should not be public. Bind Ollama to localhost only:

# In /etc/systemd/system/ollama.service
Environment="OLLAMA_HOST=127.0.0.1"

Then reload and restart:

sudo systemctl daemon-reload
sudo systemctl restart ollama

OpenClaw running on the same server communicates over localhost, so it still reaches the model. Nothing outside the server can.

For additional hardening, configure a firewall to block the port externally and use SSH tunnels if you need remote access to the Ollama API for development.

Getting Started: Your Private LLM Checklist

  1. Provision a VPS with at least 8 GB RAM (16 GB recommended for comfortable headroom)
  2. Install Ollama with the one-line script
  3. Pull Llama 3.1 8B or Phi-3 Mini as your first model
  4. Bind Ollama to localhost for security
  5. Install OpenClaw on the same server
  6. Set model: ollama/llama3.1 in your OpenClaw config
  7. Connect OpenClaw to Telegram or Discord for daily use

Total setup time is around 30 minutes. Monthly cost is the VPS fee only, no API charges.

Summary

// get started

Ready to Run Your Own Private LLM?

Install OpenClaw on your VPS and connect it to a local private LLM in under 30 minutes. Full control, zero cloud dependency.

Install OpenClaw Free →