Installing and Configuring Ollama on Ubuntu
This guide details the process of installing Ollama on Ubuntu and effectively configuring its systemd service with custom parameters. Proper configuration ensures optimal performance, resource utilization, and control over your local Large Language Model (LLM) server.
Installation
Ollama provides a convenient and straightforward script for installation on Linux-based systems. This script handles the download, setup of the systemd service, and initial start of the Ollama server.
curl -fsSL https://ollama.com/install.sh | shCustomizing Ollama Service Parameters
You can fine-tune Ollama’s behavior by modifying its systemd service configuration. This is typically done using a systemd override file to ensure your customizations persist across updates to the main service definition.
-
Edit the
systemdService File: Access the override file for the Ollama service. This requiressudoprivileges.sudo systemctl edit ollama.serviceThis command opens an editor, allowing you to create or modify an override file, typically located at
/etc/systemd/system/ollama.service.d/override.conf. -
Add or Modify Environment Variables: Within the
[Service]section of the override file, you can defineEnvironmentvariables that influence Ollama’s operation.[Service] Environment="OLLAMA_HOST=0.0.0.0" Environment="OLLAMA_NUM_THREAD=6" Environment="OLLAMA_CONTEXT_LENGTH=16384" Environment="OLLAMA_KEEP_ALIVE=24h"OLLAMA_HOST=0.0.0.0: Configures Ollama to listen on all network interfaces. This is essential if you plan to access Ollama from other machines, containers, or through a local network. By default, Ollama often binds only to127.0.0.1(localhost).OLLAMA_NUM_THREAD=6: Specifies the number of CPU threads that Ollama (or its underlyingllama.cppengine) should utilize. Adjust this value based on your CPU’s core count to optimize performance. Using0allows Ollama to attempt to auto-detect the optimal number of threads.OLLAMA_CONTEXT_LENGTH=16384: Sets the maximum context window size for models loaded by Ollama. A larger context length allows the LLM to process and “remember” more input tokens, which is beneficial for longer conversations or complex prompts. This value should be adjusted based on your available GPU VRAM and system RAM.OLLAMA_KEEP_ALIVE=24h: Determines how long loaded models remain in memory after their last use. Setting a higher value (e.g.,24hfor 24 hours,30mfor 30 minutes) prevents models from being frequently unloaded and reloaded, thereby improving response times for subsequent requests.
-
Reload
systemdand Restart Ollama: After making changes to the service configuration, you must reload thesystemdmanager to register the new settings and then restart the Ollama service for the changes to take effect.sudo systemctl daemon-reload sudo systemctl restart ollama
Useful Ollama CLI Commands
Running Models
To download and run a model, simply use the ollama run command followed by the model name. If the model is not already downloaded locally, Ollama will automatically pull it from its registry.
ollama run mistralExplicitly Pulling Models
You can pre-download models to your local system without immediately running them.
ollama pull llama2
ollama pull codellamaListing Downloaded Models
To view all LLM models currently stored on your local Ollama server:
ollama listRemoving Models
To delete a locally stored model and free up disk space:
ollama rm mistralTroubleshooting Ollama
Check Ollama Service Status
Verify if the Ollama service is running and inspect for any immediate errors.
systemctl status ollamaView Ollama Logs
Access real-time logs for the Ollama service, which are crucial for diagnosing issues.
journalctl -u ollama -fNetwork Connectivity (Firewall)
If you are encountering issues accessing Ollama remotely, ensure that your system’s firewall is configured to allow incoming connections on the Ollama port (default 11434).
sudo ufw allow 11434/tcp