What Ollama is and what you need before starting

Ollama is a tool that lets you run large language models on your own computer instead of using a web service. You download Ollama, then use it to pull and run models like Llama 2, Mistral, or Neural Chat. Everything runs locally — the model files stay on your machine, and you control what data goes into them.

Before you install, check your hardware. Ollama works on Mac, Windows, and Linux. On Mac, any recent model works fine. On Windows and Linux, you need at least 8 GB of RAM to run most models comfortably, though 16 GB is better. If you have an Nvidia GPU, Ollama can use it to run models faster — this is optional but helpful if you plan to use models regularly.

You also need about 5 to 50 GB of free disk space depending on which models you want to run. Smaller models like Mistral take 4 to 7 GB. Larger ones like Llama 2 70B take 40+ GB. Start with one model and add more later if you want.

Key Takeaways

  • Download Ollama from ollama.ai, run the installer for your operating system, and follow the on-screen prompts to finish setup.
  • After installation, open a terminal or command prompt and type ollama pull modelname to download a model — start with ollama pull mistral for a smaller, faster option.
  • Run a model by typing ollama run modelname in the terminal, then type your questions or prompts at the prompt that appears.
  • Ollama listens on localhost:11434 by default, so you can connect other applications to it if you want to use the model in a different tool.
  • Models download to your computer and stay there — you do not need an internet connection after the initial download, though Ollama checks for updates when you run it.

Installing Ollama on Mac

Go to ollama.ai and click the Download button. The site detects your operating system and offers the right version. For Mac, you get a .dmg file. Open it and drag the Ollama icon into your Applications folder, just like any other Mac app.

After the drag completes, open Applications, find Ollama, and double-click it to launch. The first time you run it, Mac asks for permission. Click Open. Ollama installs a small background service and adds itself to your menu bar at the top of the screen. You do not need to do anything else — Ollama is now ready to use.

To start using it, open Terminal (find it in Applications > Utilities or search for "Terminal" in Spotlight). Type ollama pull mistral and press Enter. Ollama downloads the Mistral model, which takes a few minutes depending on your internet speed. When it finishes, type ollama run mistral to start chatting with the model.

Installing Ollama on Windows

Go to ollama.ai and download the Windows installer. It is a .exe file. Double-click it and Windows shows an installation wizard. Click through the prompts and choose where you want to install Ollama — the default location is fine. When the installer finishes, it automatically starts Ollama and adds it to your system tray (the icons at the bottom right of your screen).

Open Command Prompt or PowerShell. On Windows 10 and 11, right-click the Start button and select "Terminal" or search for "Command Prompt" in the Start menu. Type ollama pull mistral and press Enter. The model downloads to your computer. This takes several minutes on a typical home internet connection.

Once the download finishes, type ollama run mistral and press Enter. A prompt appears where you can type questions. Type your question and press Enter twice to send it. The model processes your input and prints a response. To exit, type /bye and press Enter.

Installing Ollama on Linux

Open a terminal and run this command: curl -fsSL https://ollama.ai/install.sh | sh. This downloads and runs an installation script. It asks for your password because it needs permission to install system-wide. Type your password and press Enter. The script installs Ollama and starts it as a background service.

After installation finishes, type ollama pull mistral to download a model. When that completes, type ollama run mistral to start using it. If you want to stop Ollama later, type sudo systemctl stop ollama. To start it again, type sudo systemctl start ollama.

On some Linux systems, you may need to add your user to the ollama group so you can run commands without typing sudo each time. If you see a permission error, ask your system administrator or check the Ollama documentation for your specific Linux distribution.

Downloading and running your first model

After Ollama is installed, the next step is to download a model. Open your terminal or command prompt and type ollama pull mistral. Mistral is a good starting point — it is smaller than many other models, runs reasonably fast on most computers, and produces good responses. The download is about 4 GB and takes 5 to 15 minutes depending on your internet speed.

While the model downloads, Ollama shows progress. When it finishes, you see a message saying the model is ready. Now type ollama run mistral and press Enter. Ollama loads the model into memory and shows a prompt where you can type. Type a question like "What is the capital of France?" and press Enter twice. The model thinks for a moment and prints an answer.

To try a different model later, type ollama pull llama2 or ollama pull neural-chat. Each model has different strengths — Llama 2 is larger and more capable but slower, while Neural Chat is smaller and faster. You can have multiple models installed at once and switch between them with ollama run modelname.

Connecting other applications to Ollama

Ollama runs a server on your computer that listens for requests on localhost:11434. This means other applications can send prompts to Ollama and get responses back. If you use a tool like Open WebUI, Lmstudio, or a custom Python script, you can point it at http://localhost:11434 and it will use your local Ollama models instead of a cloud service.

To check that Ollama is running and listening, open a terminal and type curl http://localhost:11434/api/tags. If Ollama is running, you see a list of models you have downloaded. If you get an error, Ollama may not be running — start it by typing ollama serve in a new terminal window.

Some applications need the model name in a specific format. When you run ollama pull mistral, the full model name is mistral:latest. If an application asks for a model name, use that format. You can also pull specific versions — for example, ollama pull mistral:7b pulls the 7-billion-parameter version of Mistral.

Troubleshooting common installation problems

If Ollama does not start after installation, try restarting your computer. On Mac, check that Ollama appears in your menu bar. On Windows, look for it in the system tray. On Linux, type sudo systemctl status ollama to see if the service is running.

If a model download fails or gets stuck, open a terminal and type ollama pull modelname again. Ollama resumes the download from where it left off. If you want to start fresh, type ollama rm modelname to delete the model, then pull it again.

If Ollama runs very slowly or your computer becomes unresponsive, you may not have enough RAM. Close other applications to free up memory. If that does not help, try a smaller model like Mistral instead of Llama 2. You can also limit how much memory Ollama uses by setting the OLLAMA_NUM_PARALLEL environment variable, though this requires some technical knowledge.

If you see an error about GPU support, Ollama is trying to use your graphics card but cannot find the right drivers. On Nvidia systems, install the Nvidia CUDA toolkit. On other systems, Ollama falls back to using your CPU, which is slower but still works.

Frequently Asked Questions

Do I need an internet connection to use Ollama after I download a model?

No. Once a model is downloaded and stored on your computer, you can run it offline. Ollama checks for updates when you start it, but the model itself runs without connecting to the internet. This is one of the main reasons people use Ollama instead of cloud-based services.

How much disk space do different models take up?

Mistral takes about 4 GB. Llama 2 7B takes about 4 GB. Llama 2 13B takes about 8 GB. Llama 2 70B takes about 40 GB. Neural Chat takes about 5 GB. Check the Ollama model library at ollama.ai/library to see the size of any model before you download it.

Can I run multiple models at the same time?

You can have multiple models installed, but running them simultaneously uses a lot of RAM and is usually slow. Ollama loads one model at a time into memory. If you switch to a different model while one is running, Ollama unloads the first one and loads the second. For most home computers, running one model at a time works best.

What is the difference between running Ollama in a terminal and using a web interface?

Running ollama run modelname in a terminal gives you a text-based chat. A web interface like Open WebUI adds a graphical window in your browser with a cleaner look and more features. Both connect to the same Ollama server running on your computer — the web interface is just a different way to send prompts and see responses.

Can I update Ollama after I install it?

Yes. On Mac, download the latest version from ollama.ai and drag it to Applications again — it replaces the old version. On Windows, download the new installer and run it. On Linux, type curl -fsSL https://ollama.ai/install.sh | sh again to update. Your downloaded models stay on your computer and do not need to be re-downloaded.