Running Large Language Models (LLMs) locally has evolved from a niche experiment for machine learning engineers into a practical standard for developers and power users. The barrier to entry was once high, requiring users to manage CUDA environments, compile C++ code, and juggle complex Python dependency trees. Today, that complexity is obsolete.
Ollama simplifies this process significantly. As an open-source tool designed to run LLMs locally, it eliminates the need for cloud-based API calls, thereby addressing data privacy concerns and removing monthly subscription fees. Since its initial release in November 2023, Ollama has been downloaded over 10 million times. It supports a wide variety of models—including Llama 3, Mistral, Gemma, and Phi-3—allowing users to download and run them with simple command-line instructions.
Under the hood, Ollama leverages the llama.cpp library for its core inference engine, written primarily in Go and C++. It utilizes the GGUF file format for model weights, which enables efficient quantization. This means you can run powerful models on consumer-grade hardware that would otherwise be insufficient for standard PyTorch formats.
Below is a step-by-step guide to setting up your own local inference server, covering everything from installation to custom model creation.
The first step is installing the software on your machine. Ollama is compatible with multiple operating systems, including macOS (Intel and Apple Silicon), Linux (x86_64 and ARM64), and Windows. There is no need to compile anything; you simply download the installer for your specific platform from the official site.
Once downloaded, execute the installation package. If you are on macOS or Linux, you can verify the installation directly in your terminal. Open a terminal window and type ollama --version. If the installation was successful, the command will display the current version number (e.g., 0.3.0).
Key Takeaway: Always verify the installation by checking the version. If the command isn’t recognized, ensure your PATH environment variable includes the Ollama binary location, or restart your terminal.
After confirming the binary exists, you need to ensure the background server is running. Ollama operates as a daemon that listens for requests on port 11434 by default. You can verify that this port is active using network tools like lsof -i :11434 (on macOS/Linux) or by checking the system tray icon on Windows to ensure the service is running. If the server isn't up, subsequent commands will fail to connect.
Ollama acts as a registry for model weights, so you don’t need to manually download large .gguf files from Hugging Face. Instead, use the pull command to fetch models automatically.
To get started, run ollama pull llama3. This command downloads the Llama 3 model weights to your local disk, displaying a progress bar that indicates download speed and total size.
It is important to understand why this process is efficient: Ollama uses the GGUF format for quantization. For example, the Llama 3 8B model, when quantized to Q4_K_M (a common default balance of speed and accuracy), requires approximately 4.7 GB of RAM or VRAM. This is significantly less memory than the unquantized version, making it accessible on laptops with 16GB of unified memory.
Once the download completes, check your installed models by running ollama list. This command confirms the download was successful and provides storage details, showing exactly how much disk space each model occupies.
Key Takeaway: Quantization levels matter. Q4_K_M is a good default for 8B models, but if you have more VRAM (e.g., 24GB on an RTX 3090), you can pull higher precision versions like Q8_0 for better accuracy.
Once the model is pulled, you can interact with it immediately without writing any code. Start a terminal-based chat session by running ollama run llama3. The prompt will change to indicate that you are now inside the model’s context.
You can type prompts directly into the terminal to test basic inference capabilities. For example, typing "Explain quantum computing in simple terms" will generate a response. The text appears as it is generated, giving you immediate feedback on the model's performance and latency. This is useful for a quick sanity check before integrating the model into an application.
When you are finished testing, exit the session using exit or by pressing Ctrl+D. This closes the interactive loop but keeps the model loaded in memory for a short period, allowing for fast subsequent calls.
Key Takeaway: The interactive mode is perfect for manual testing and prompt engineering. It provides real-time visibility into token generation speed and output quality without setting up a web server or API client.
Not everyone prefers typing in a terminal. Ollama includes a built-in web interface that allows users to interact with models via a browser without writing code. Access this UI at http://localhost:11434.
This browser-based interface provides a clean, no-code interaction experience. If you have pulled multiple models (e.g., llama3 and mistral), you can select different ones from the dropdown menu. The interface supports streaming responses, so you see the text appear word-by-word in real-time, similar to how it appears on chatbot websites.
Leverage this tool for quick demos to stakeholders or for non-technical users who want to try out local AI capabilities. It is also helpful for debugging issues, as you can verify if the API is responding correctly before building your own client.
Key Takeaway: The web UI is a powerful tool for demonstration and quick iteration. If the browser interface works, your server is running correctly, and any issues with custom code are likely in your implementation rather than the Ollama engine.
For production use or integration into your own applications, you will need to interact with Ollama via its REST API. The inference server runs on the default port (11434) and exposes endpoints accessible via HTTP requests from Python, JavaScript, Go, or any other language that handles HTTP.
The primary endpoints are /api/generate for single-turn completions and /api/chat for multi-turn conversations with context. You can send HTTP POST requests to http://localhost:11434/api/generate with a JSON body containing your prompt.
If you are using Python, the requests library is ideal for handling these requests. To handle streaming responses for real-time text generation, set stream=True in your request and iterate over the response lines, parsing the JSON chunks as they arrive. This allows your application to display text as it is generated, rather than waiting for the entire response to complete.
You can also pass parameters like temperature, top_k, and top_p directly in the API request body to control output behavior. For example, setting temperature: 0.1 makes the model more deterministic and fact-focused, while temperature: 0.9 encourages more creative and varied responses.
Key Takeaway: Always use streaming for user-facing applications. Waiting for a full LLM response can take seconds or minutes, leading to poor user experience. Streaming provides immediate feedback.
One of Ollama’s most powerful features is the ability to create custom models using Modelfiles. A Modelfile is a simple text file that defines a new model based on an existing base model.
Define a custom model by creating a Modelfile in your project directory. The first line specifies the base model, e.g., FROM llama3. Following that, you can set a specific system prompt using the SYSTEM directive. For example, if you want to build a legal assistant, you might write:
FROM llama3
SYSTEM You are a helpful legal assistant. Answer questions based on US law and cite relevant statutes when possible. Do not provide legal advice, only informational context.
PARAMETER temperature 0.2
PARAMETER top_k 40
This file sets the model’s role and generation parameters (low temperature for consistency, restricted top_k for stricter adherence to the prompt). Build the custom model using the command ollama create my-legal-bot -f Modelfile. This compiles the instructions into a new model that you can run just like any other: ollama run my-legal-bot.
Key Takeaway: Modelfiles allow you to specialize general-purpose models for specific tasks without fine-tuning. It’s a quick way to enforce persona, tone, and domain-specific constraints.
Running LLMs locally is resource-intensive. You must monitor RAM and VRAM usage, as running multiple models simultaneously is limited by hardware constraints. Ollama allows you to run multiple models, but typically only one is loaded into active memory at a time to maximize performance. If you switch between models, the previous one may be unloaded to free up memory for the new one.
To balance accuracy and memory efficiency, adjust quantization levels. A Q8_0 quantization uses more memory but preserves more of the original model's accuracy compared to Q4_K_M. If you are hitting memory limits, stick with lower quantizations. Conversely, if you have ample VRAM (e.g., 24GB+), higher quantizations will yield noticeably better reasoning capabilities.
Finally, manage your disk space. LLMs are large; a single 70B model can take up over 40GB of storage. Use ollama rm <model_name> to delete unused models and free up disk space. This command removes the weights from your local registry, ensuring you don't fill up your SSD with experimental models you no longer need.
Key Takeaway: Treat local LLMs like traditional software dependencies. Regularly clean up unused versions and choose quantization levels that fit your hardware budget to maintain system stability.
Can I run Ollama on a CPU-only machine without a GPU? Yes. Ollama is designed to work on CPUs, though inference will be significantly slower than on a dedicated GPU. For models up to 8B parameters, a modern CPU with sufficient RAM can handle interactive chat sessions at acceptable speeds. Larger models (70B+) on CPU are impractical for real-time interaction due to high latency.
How do I check which models are currently installed on my system?
Use the command ollama list in your terminal. This will display a table of all downloaded models, including their names, sizes, and creation dates. It helps you track what is taking up space on your disk.
Is Ollama free to use, and what about the licensing of the models themselves? Ollama itself is open-source and free to download and use. However, the licensing of the underlying models varies. For example, Llama 3 has a specific license agreement that restricts commercial use for very large companies (over 700 million monthly active users), while Mistral and Gemma have their own specific terms. Always check the license of the specific model you are pulling to ensure compliance with your intended use case.
How do I delete a model from Ollama to free up storage space?
Use the command ollama rm <model_name>. For example, ollama rm llama3 will remove the Llama 3 weights from your local drive. This is essential for managing disk space if you are experimenting with multiple different architectures.
Ready to bring AI to your local machine? Download Ollama today and start experimenting with LLMs without relying on cloud services. With full control over your data, hardware, and model parameters, local inference offers a level of privacy and customization that cloud APIs simply cannot match.