Installing and Configuring Ollama on Ubuntu

This guide details the process of installing Ollama on Ubuntu and effectively configuring its systemd service with custom parameters. Proper configuration ensures optimal performance, resource utilization, and control over your local Large Language Model (LLM) server.

Installation

Ollama provides a convenient and straightforward script for installation on Linux-based systems. This script handles the download, setup of the systemd service, and initial start of the Ollama server.

curl -fsSL https://ollama.com/install.sh | sh

Customizing Ollama Service Parameters

You can fine-tune Ollama’s behavior by modifying its systemd service configuration. This is typically done using a systemd override file to ensure your customizations persist across updates to the main service definition.

  1. Edit the systemd Service File: Access the override file for the Ollama service. This requires sudo privileges.

    sudo systemctl edit ollama.service

    This command opens an editor, allowing you to create or modify an override file, typically located at /etc/systemd/system/ollama.service.d/override.conf.

  2. Add or Modify Environment Variables: Within the [Service] section of the override file, you can define Environment variables that influence Ollama’s operation.

    [Service]
    Environment="OLLAMA_HOST=0.0.0.0"
    Environment="OLLAMA_NUM_THREAD=6"
    Environment="OLLAMA_CONTEXT_LENGTH=16384"
    Environment="OLLAMA_KEEP_ALIVE=24h"
    • OLLAMA_HOST=0.0.0.0: Configures Ollama to listen on all network interfaces. This is essential if you plan to access Ollama from other machines, containers, or through a local network. By default, Ollama often binds only to 127.0.0.1 (localhost).
    • OLLAMA_NUM_THREAD=6: Specifies the number of CPU threads that Ollama (or its underlying llama.cpp engine) should utilize. Adjust this value based on your CPU’s core count to optimize performance. Using 0 allows Ollama to attempt to auto-detect the optimal number of threads.
    • OLLAMA_CONTEXT_LENGTH=16384: Sets the maximum context window size for models loaded by Ollama. A larger context length allows the LLM to process and “remember” more input tokens, which is beneficial for longer conversations or complex prompts. This value should be adjusted based on your available GPU VRAM and system RAM.
    • OLLAMA_KEEP_ALIVE=24h: Determines how long loaded models remain in memory after their last use. Setting a higher value (e.g., 24h for 24 hours, 30m for 30 minutes) prevents models from being frequently unloaded and reloaded, thereby improving response times for subsequent requests.
  3. Reload systemd and Restart Ollama: After making changes to the service configuration, you must reload the systemd manager to register the new settings and then restart the Ollama service for the changes to take effect.

    sudo systemctl daemon-reload
    sudo systemctl restart ollama

Useful Ollama CLI Commands

Running Models

To download and run a model, simply use the ollama run command followed by the model name. If the model is not already downloaded locally, Ollama will automatically pull it from its registry.

ollama run mistral

Explicitly Pulling Models

You can pre-download models to your local system without immediately running them.

ollama pull llama2
ollama pull codellama

Listing Downloaded Models

To view all LLM models currently stored on your local Ollama server:

ollama list

Removing Models

To delete a locally stored model and free up disk space:

ollama rm mistral

Troubleshooting Ollama

Check Ollama Service Status

Verify if the Ollama service is running and inspect for any immediate errors.

systemctl status ollama

View Ollama Logs

Access real-time logs for the Ollama service, which are crucial for diagnosing issues.

journalctl -u ollama -f

Network Connectivity (Firewall)

If you are encountering issues accessing Ollama remotely, ensure that your system’s firewall is configured to allow incoming connections on the Ollama port (default 11434).

sudo ufw allow 11434/tcp