For developers
Your coding model, on your machine
The same API you already use, no per-token bill, and your code never leaves your network. Ten minutes with Ollama and a 7B model.
41 tok/s qwen2.5-coder:7b · 6.6 GB loaded · measured on our DGX Spark, 2026-10-04
Start in 10 minutes
- Install Ollama from ollama.com.
- Pull the model and call the OpenAI-compatible API:
ollama pull qwen2.5-coder:7b
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "qwen2.5-coder:7b",
"messages": [{"role": "user", "content": "Write a Python function that reverses a string"}]}'From Python, change only the base URL of the official OpenAI client:
from openai import OpenAI
# The client requires a key; Ollama ignores it.
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
reply = client.chat.completions.create(
model="qwen2.5-coder:7b",
messages=[{"role": "user", "content": "Explain this error: KeyError on line 12"}],
)
print(reply.choices[0].message.content)Any editor or tool that accepts an OpenAI-compatible base URL can point at http://localhost:11434/v1. With 16 GB of RAM a 7B model is comfortable; for more quality, look at the 32B class in hardware.
What happens on your machine
Your editor talks to a local API identical to OpenAI's; the model runs on your GPU. There is no cable going out.
| Local Ollama on your machine | Cloud API an external vendor | |
|---|---|---|
| Where the model runs | On your machine | On the vendor's server |
| Your code leaves your network | No | Yes |
| Per-token bill | No | Yes |
| Works offline | Yes | No |
| Speed we measured | 41 tok/s | Not measured by us; we do not publish other people's figures. |
Machines that run a coding assistant
The cheapest that run it comfortably and at a practical speed (same rule as the catalogue).
Tools for agents, MCP and prompts
Agent builderDefine an agent and export its configuration.MCP builderScaffold an MCP server for your tools.Skill builderPackage reusable instructions for your assistant.Plugin builderBundle skills, agents and hooks into a plugin.Prompt writerStructure a prompt and score it.VRAM calculatorWhat fits on your GPU before you download.Dev Playground30 self-checking browser exercises: streaming, JSON, RAG, retries.
Guides for developers
- Point Claude Code at a Local Model with OllamaThree environment variables put Claude Code on a local Ollama model: the setup, what stays on your machine, what does not, and measurements with three models.
- CUDA on Fedora: a Local AI Workstation in an AfternoonInstall the NVIDIA driver from RPM Fusion, get Ollama on the GPU in five minutes, then build llama.cpp with the CUDA toolkit using the official Fedora guide.
- GGUF Quantization: Pick the Right File for Your MachineHow to read a GGUF file name, choose between Q4_K_M, Q5 and Q8 for the memory you have, and measure speed, memory and quality yourself with three commands.
- Llama 3.3 70B on Your Own Hardware: What It Takes to Run ItThe memory a 70B model really needs, measured and cited speeds on a 64 GB Mac, two 24 GB GPUs and our GB10, and the SME tasks where it beats an 8B model.
- llama.cpp on Intel Arc with SYCL: 16 GB of VRAM on a BudgetBuild llama.cpp with Intel's SYCL backend, run an 8B model on an Arc A770 or B580, and know from third-party numbers what speed to expect before you buy.
- Run Your Own AI Agent Gateway with OpenClaw and Ollama in One CommandInstall OpenClaw, point it at a local Ollama model, schedule your first automation and size the memory it needs, with numbers measured on our own gateway.
- Qwen2.5-72B-Instruct Locally: the 72B vs a 15x Faster 35BWhat Qwen2.5-72B needs in memory, how fast it really generates on a 128 GB workstation, and a measured Spanish-prompt comparison against qwen3.6:35b.
- Qwen 2.5 Coder 7B on Your Own Machine in 20 MinutesHands-on review of Qwen2.5-Coder-7B on Ollama: measured speed and memory, three SME tasks with exact prompts, honest limits, and VS Code and n8n wiring.
- Local RAG over Company Docs with n8n and OllamaStep-by-step tutorial to build a RAG pipeline with n8n, Ollama and ChromaDB on your own hardware, with commands, measured timings and zero cost per query.