Running an LLM locally used to mean compiling code from GitHub, wrestling with CUDA drivers, and hoping your GPU had enough VRAM. That era is over. Ollama LLM changed the experience entirely.
With Ollama, you pull a model the same way you pull a Docker image: one command, and it runs. No configuration files. No Python environment headaches. No API keys. And crucially, no data leaving your machine.
This guide covers what Ollama is, how to install it, which models to run, and how to connect it to OpenClaw for a fully self-hosted personal AI assistant.
What Is Ollama?
Ollama is an open-source tool that lets you download and run large language models on your local machine or server. It handles model storage, quantisation, and serving, wrapping everything in a clean CLI and a local REST API.
You interact with it in two ways: directly via the terminal for quick queries, or through the API at localhost:11434 for programmatic use. Any application that can call an OpenAI-compatible API can talk to Ollama with a one-line config change.
Ollama supports macOS, Linux, and Windows. It works with NVIDIA and AMD GPUs, and falls back gracefully to CPU if no GPU is present (though speed will be slower).
How to Install Ollama
Installation is a single command on Linux and macOS:
curl -fsSL https://ollama.com/install.sh | sh
That installs the Ollama binary and starts the service automatically. On macOS, there is also a desktop app available from ollama.com if you prefer a GUI launcher.
On Windows, download the installer from the Ollama website. It sets up both the service and a system tray icon.
After installation, verify it is running:
ollama list
You should see an empty table, since no models have been downloaded yet. Now you are ready to pull one.
Best Ollama Models in 2026
The Ollama model library has grown significantly. Here are the models worth knowing, grouped by use case.
Best General-Purpose Model: Llama 3.1 8B
ollama pull llama3.1
Meta's Llama 3.1 at the 8B parameter size is the sweet spot for most users. It runs on a machine with 8 GB RAM, gives sharp results for chat, summarisation, and simple coding tasks, and responds quickly even on CPU. The 70B variant is dramatically stronger but requires a serious GPU or a lot of patience.
Best Coding Model: Qwen2.5-Coder
ollama pull qwen2.5-coder:7b
Alibaba's Qwen2.5-Coder line is the current local benchmark winner for code generation. The 7B model beats many larger models on coding benchmarks and fits comfortably in 8 GB VRAM. If you are using OpenClaw for code review, automation scripting, or building tools, this is the model to run locally.
Best Small Model: Phi-4 Mini
ollama pull phi4-mini
Microsoft's Phi-4 Mini is a 3.8B model that punches well above its weight. It runs on almost any modern laptop, responds fast even on CPU, and handles instruction-following tasks surprisingly well. Good for low-powered VPS instances where you want local inference without renting a GPU node.
Best for Long Context: Gemma 3
ollama pull gemma3:12b
Google's Gemma 3 at 12B supports a 128K context window in Ollama. If you need to feed it long documents, entire codebases, or extended conversation histories without truncation, this is the local option to reach for.
Best Multilingual Model: Qwen2.5
ollama pull qwen2.5:7b
The general Qwen2.5 line supports over 29 languages with strong quality. If you work in a non-English language or need your AI assistant to handle multilingual conversations, Qwen2.5 is the local pick.
Running an LLM Locally: Step by Step
Once Ollama is installed, running a model locally takes two commands:
# 1. Download the model (one-time)
ollama pull llama3.1
# 2. Start a chat session
ollama run llama3.1
The first command downloads the model file, which ranges from about 4 GB for a 7B model to 40 GB or more for larger variants. After that, ollama run launches an interactive chat right in your terminal.
To run a quick one-off query without an interactive session:
ollama run llama3.1 "What is a vector database and when should I use one?"
Models are stored locally, so subsequent runs start instantly after the first download. No internet connection required after the initial pull.
Using the Ollama API
For programmatic use, Ollama exposes a REST API at http://localhost:11434. It follows the OpenAI chat completions format, which means most OpenAI-compatible tools work with Ollama by just changing the base URL.
A basic API call looks like this:
curl http://localhost:11434/v1/chat/completions -H "Content-Type: application/json" -d '{
"model": "llama3.1",
"messages": [
{"role": "user", "content": "Summarise the concept of RAG in two sentences."}
]
}'
The response comes back in the same format as OpenAI's API. Any client library with an OpenAI-compatible mode, including LangChain, LlamaIndex, and most agent frameworks, can point at localhost:11434 with no other changes.
To make Ollama accessible from other machines on your network or from a remote server, start it with the host flag:
OLLAMA_HOST=0.0.0.0 ollama serve
Then reference it by IP from other devices. This is how you expose a local GPU machine's Ollama instance to other applications running on the same network or VPS.
Connecting Ollama to OpenClaw
OpenClaw's model-agnostic architecture means you can swap in any Ollama model with a single config change. This is covered in depth in the local AI with OpenClaw guide, but here is the short version.
In your OpenClaw gateway config, set the model to the Ollama format:
model: ollama/llama3.1
If Ollama is running on the same machine as OpenClaw, nothing else changes. The gateway calls localhost:11434 automatically. If Ollama is on a different machine, set the base URL:
model: ollama/llama3.1
ollama_base_url: http://192.168.1.50:11434
Once connected, OpenClaw runs entirely on your infrastructure. Your SOUL.md, MEMORY.md, skills, and cron jobs all work identically. The only difference is that inference happens on your hardware rather than a cloud provider's servers.
This setup is particularly powerful for privacy-sensitive use cases: handling personal documents, automating access to private systems, or running an AI assistant without any third-party data processing. The data sovereignty guide goes deeper on what this means in practice.
For users who want the best of both worlds, you can route routine low-stakes cron jobs through a cheap cloud model and reserve Ollama for tasks involving private data. OpenClaw's per-job model configuration makes this straightforward. See the multi-model setup guide for how to configure it.
Ollama vs LM Studio
Both Ollama and LM Studio are popular options for running LLMs locally. The difference comes down to interface and use case.
LM Studio is primarily a desktop GUI application. It is excellent for exploring models visually and has a built-in chat interface that feels polished. If you want a graphical model browser and are running on a personal laptop, LM Studio is worth trying.
Ollama is better for server use and programmatic access. Its CLI and API-first design makes it the right choice when you want to integrate local inference into another application, run it headlessly on a VPS, or call it from automation tools like OpenClaw or n8n. It is also simpler to keep running as a background service.
For OpenClaw users, Ollama is almost always the right choice. It is designed to be called programmatically, it runs as a service, and it handles all the model lifecycle management so OpenClaw does not have to.
Hardware Requirements
The minimum hardware depends on the model size you want to run:
- 3B models (Phi-4 Mini, Llama 3.2 3B): 4 GB RAM, any modern CPU. Will run on a basic VPS.
- 7B models (Llama 3.1 8B, Qwen2.5 7B): 8 GB RAM. CPU inference is slow but works. A 6 GB VRAM GPU makes it fast.
- 13B models (Gemma 3 12B): 16 GB RAM or 12 GB VRAM GPU for reasonable speed.
- 70B models (Llama 3.3 70B): 32 GB RAM minimum, or a dedicated GPU server. For most users, this size is better accessed via API.
GPU inference is roughly 5 to 10 times faster than CPU for most models. An RTX 3080 with 10 GB VRAM can run 7B models at 40 to 60 tokens per second, which feels close to real-time in conversation. CPU inference on a modern laptop gets around 5 to 10 tokens per second for the same models, which is usable for non-interactive tasks.
If you do not have a GPU but want fast local inference, cloud GPU rentals from RunPod or Vast.ai cost around $0.20 to $0.40 per hour for a machine that runs 7B models quickly. For occasional use, this can be cheaper than API costs at volume.
For running OpenClaw with a local Ollama instance on a VPS without a GPU, the 3B model class is the practical limit for real-time conversation. For batch processing or non-interactive automations where latency does not matter, 7B models on CPU work fine.
The Bottom Line
Ollama has removed most of the friction from running LLMs locally. The installation is trivial, the model library is broad, and the API is immediately compatible with the wider AI tooling ecosystem.
For OpenClaw users, Ollama is the recommended path to a fully private setup. You get persistent memory, custom skills, Telegram and Discord integration, and scheduled automation, all running on models that never touch a third-party server.
The combination of OpenClaw plus Ollama is the self-hosted AI stack that was theoretical two years ago. Today it is an afternoon's setup on a spare machine or cheap VPS.
Ready to Run Your Own Private AI Assistant?
Install OpenClaw on your VPS, connect it to Ollama, and have a fully private AI assistant running on your own hardware today.
Install OpenClaw Free