For developers

Your coding model, on your machine

The same API you already use, no per-token bill, and your code never leaves your network. Ten minutes with Ollama and a 7B model.

41 tok/s qwen2.5-coder:7b · 6.6 GB loaded · measured on our DGX Spark, 2026-10-04

Start in 10 minutes

  1. Install Ollama from ollama.com.
  2. Pull the model and call the OpenAI-compatible API:
ollama pull qwen2.5-coder:7b

curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen2.5-coder:7b",
       "messages": [{"role": "user", "content": "Write a Python function that reverses a string"}]}'

From Python, change only the base URL of the official OpenAI client:

from openai import OpenAI

# The client requires a key; Ollama ignores it.
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")

reply = client.chat.completions.create(
    model="qwen2.5-coder:7b",
    messages=[{"role": "user", "content": "Explain this error: KeyError on line 12"}],
)
print(reply.choices[0].message.content)

Any editor or tool that accepts an OpenAI-compatible base URL can point at http://localhost:11434/v1. With 16 GB of RAM a 7B model is comfortable; for more quality, look at the 32B class in hardware.

What happens on your machine

Your editor talks to a local API identical to OpenAI's; the model runs on your GPU. There is no cable going out.

The local stack, from your editor to the GPU Your editor calls http://localhost:11434/v1, an OpenAI-compatible API served by Ollama or vLLM, which runs the model on your GPU. Everything stays inside your machine. Nothing leaves your machine Your editor, agent or script OpenAI client with a localbase_url http://localhost:11434/v1 OpenAI-compatible API Ollama / vLLM the runtime that serves the model qwen2.5-coder:7b 6.6 GB loaded · measured GPU and memory your hardware, on your desk
Does the model fit in 16 GB? qwen2.5-coder:7b takes 6.6 GB loaded, measured on our machine, against 16 GB of memory. It fits comfortably. qwen2.5-coder:7b, loaded 6.6 GB 0 16 GB Fits comfortably
Measured on our DGX Spark, 2026-10-04. Room is left for the OS and the editor.
Local Ollama on your machine Cloud API an external vendor
Where the model runs On your machine On the vendor's server
Your code leaves your network No Yes
Per-token bill No Yes
Works offline Yes No
Speed we measured 41 tok/s Not measured by us; we do not publish other people's figures.

Machines that run a coding assistant

The cheapest that run it comfortably and at a practical speed (same rule as the catalogue).

See every machine that works for code →

Tools for agents, MCP and prompts

Guides for developers