View all articles
Local AIOllamaHardwareCostsGuide

Local AI: what it is, what you need and when it pays off

JG
Jacobo González Jaspe
|
Local AI: what it is, what you need and when it pays off Measure and decide
This article is also available in Spanish:IA local: qué es, qué necesitas y cuándo compensa

By the end of this post you will know what local AI is, which pieces you need to try it this afternoon, and how to decide whether it beats a cloud API for you. You can start with the laptop you already have. With 8 GB of memory a small model already runs.

What local AI is

Local AI means the model runs on hardware you control. That can be your laptop, a mini PC in the office or a server in your rack. The prompt and the answer never leave that machine.

One distinction saves a lot of confusion: we are almost always talking about inference, not training. Inference is using a model that is already trained to get answers. Training a model from scratch takes data centres and is not a project for a small business. Adapting an existing model to your data (fine-tuning) is possible locally, but it is a second step and most people do not need it.

If you come from hospitality, think of it this way. A cloud API is eating out: you pay per plate and the kitchen is not yours. Local AI is your own kitchen: you buy the stove once and decide what comes through the door.

The three pieces you need

  1. A model. Open-weight models (Llama, Qwen, Gemma, Mistral, DeepSeek) are free to download. They usually come as GGUF files, and the file type decides how much memory they take. We explain it in GGUF quantization: pick the right file.
  2. A runtime, the program that loads the model and answers. Ollama is the simplest: free, no account, with a local API. llama.cpp is the engine underneath and gives you more control.
  3. Enough memory. The model file has to fit in RAM or VRAM, plus room for the conversation context. If it does not fit, it does not run slowly; it barely runs at all.

How much memory depends on model size and quantization. These figures come from the llama.cpp quantize README and bartowski’s table, collected in our GGUF guide (September 2026):

ModelQ4_K_M file (4-bit)F16 file (original)
Llama 3.1 8B4.58 GiB14.96 GiB
Llama 3.3 70B42.5 GB~141 GB

The rule of thumb that follows: at 4 bits, a little over half a GB per billion parameters, plus context. Before you download anything, work out your case in the VRAM explorer.

Try it in ten minutes

On Linux, Ollama installs with one line. On macOS and Windows, download it from ollama.com/download. The full step-by-step guide, with the usual problems: how to install Ollama and run your first local LLM.

bash
curl -fsSL https://ollama.com/install.sh | sh
ollama run llama3.1:8b "Explain local AI in three sentences."
ollama ps   # shows how much memory the loaded model uses

If your machine has 8 GB, swap llama3.1:8b for llama3.2:3b. If the download fails, check you have about 5 GB of free disk.

What we measured on our machine

These figures come from our own workstation, not a vendor sheet. It is far more machine than you need to start; it is here so you have a reference point.

Model (Ollama)Memory loadedGenerationMachine and date
llama3.1:8b9.2 GB41.3 tok/sDGX Spark GB10, 128 GB, 2026-10-04
qwen2.5-coder:7b6.6 GB41.0 tok/sDGX Spark GB10, 128 GB, 2026-10-04
gemma4:26b17 GB60.7 tok/sDGX Spark GB10, 128 GB, 2026-10-04
faster-whisper smalln/a30 min of audio in about 7.5 minDGX Spark GB10, 128 GB, 2026-10-04

Look at the third row. The 26B model generated faster than the 7B and 8B ones. Parameter count alone does not predict speed, which is why you measure. To measure it on your machine, use --verbose as the Ollama install guide explains.

What it is good at today, and what not

It works well on bounded, repetitive tasks:

  • Summarising, classifying and extracting data from documents.
  • Drafting emails, product sheets or reports.
  • Answering questions over your own documents with RAG (retrieval plus generation).
  • Completing and reviewing code with a 7B model.
  • Transcribing meetings with Whisper without uploading the audio anywhere.

When not to do this, or not only locally: hard multi-step reasoning, long-running agents and very specialised knowledge. The best cloud models are still ahead there. Locally, maintenance is also yours: updates, backups and security. The pattern that works best is hybrid: routine and private work on your machine, and a spend-capped cloud key for the exceptions.

Privacy and GDPR, in one paragraph

If the prompt and the answer stay inside your perimeter, there is no outside provider processing that data and no international transfer to justify. That fits the data protection by design principle in GDPR Article 25. It does not make you compliant on its own: you still need a legal basis, a record of processing and security measures. We cover it in GDPR Article 25: local AI inference is privacy by design and GDPR and AI in 2026: local deployment is the clean answer.

Local versus a cloud API

CriterionLocal AICloud API
Where the data livesOn your machineOn the provider’s servers
Cost shapeHardware once, plus electricity and your timeVariable, per token used
LatencyNo network round trip; depends on your hardwareDepends on the network and provider load
MaintenanceYours: updates, monitoring, backupsThe provider’s
Quality ceilingThe best open model that fits your memoryThe most capable models on the market
Getting startedInstall Ollama and download a modelAccount, card and API key

When it pays off

The cloud charges per token and your own machine is a fixed cost. So the answer almost always comes down to volume.

  • It pays off when several people use it daily, when volume is steady and when data must not leave the building.
  • It does not pay off on cost if you send a few short requests a day and a small cloud model handles the task. Then choose local only if privacy or predictability matter more.
  • Electricity counts. Worked example: a 30 W machine left on all month uses 30 W × 24 h × 30 days = 21.6 kWh. Multiply by your own tariff.

To put in your own numbers, use the script in cloud vs local AI: work out your break-even. For a full business case, continue with the local AI ROI framework. The hardware page includes a break-even calculator.

Verdict. Buy if you have daily volume or sensitive data. Wait if you do not yet know how many tokens you use: measure for two weeks. Skip if you ask ten questions a week.

Where to start, by profile

  1. Curious. Install Ollama on your laptop with the block above and ask it questions from your work. When you want to understand the pieces, the playground has short in-browser exercises on the context window, text similarity and whether a model fits in memory.
  2. Developer. Start at the developer hub, pick your file with the GGUF guide and measure speed with ollama run <model> --verbose.
  3. Company. See what each budget buys in the local AI hardware catalogue. Then work out the savings with the ROI framework.

Work with us

This for your company? We run the numbers on your real volumes before recommending any hardware, and we tell you when the cloud is enough. Let’s talk for 15 minutes or see how our consulting works.

Diagram
Share: LinkedIn X
Veredicto semanal

Get new guides before anyone else

Subscribe and we tell you when new guides, templates and workflows go up. One email a week, no spam.

Already published: 69 guides and 25 templates. All free, no signup.

Bonus: the local-AI starter pack PDF when you subscribe
Once a week No spam Unsubscribe anytime

See what you get

The EU AI Act now applies: a checklist you can complete

Tell us what you want to run

Tell us what you want to run and on what budget. We will tell you which hardware you need, which model fits, and what to expect from it, before you spend anything.

First call free, 15 min Local-first: your data stays on your network Open tools and guides

69 free guides · 17 compliance templates