LLM engineering · Easy · 12 min

Does the model fit in memory?

The same sum our product pages use: weights, KV cache and runtime overhead against usable memory.

Write memoryNeededGb(params, quant, kvMbPerToken, contextTokens), which returns the GB needed:

- weights: params (billions) × BITS_PER_WEIGHT[quant] / 8; - KV cache: kvMbPerToken × contextTokens / 1000; - plus RUNTIME_OVERHEAD_GB (1.5 GB).

Also write fitStatus(neededGb, usableGb): return 'fits' when the need is <= 85 % of usable memory, 'tight' when it fits but goes past 85 %, and 'no' when it does not fit. Do not round anything: the results must match the site's.

Challenges 0/4

  • Llama 3.1 8B at Q4 with 8192 tokens needs 7.42 GB
  • Llama 3.1 70B at Q8 with 32,768 tokens: 87.26 GB
  • fits, tight and no verdicts around the 85 % line
  • On a DGX Spark (122 GB usable) the 70B does not fit in fp16 but fits in Q8

function memoryNeededGb(params, quant, kvMbPerToken, contextTokens) {
  const weights = (params * BITS_PER_WEIGHT[quant]) / 8;
  // add the KV cache and the runtime overhead
  return weights;
}

function fitStatus(neededGb, usableGb) {
  // 'fits' up to 85 % of usableGb, 'tight' up to 100 %, otherwise 'no'
  return neededGb <= usableGb ? 'fits' : 'no';
}

// Llama 3.1 8B at Q4 with 8192 tokens of context (0.131 MB of KV per token)
console.log(memoryNeededGb(8.0, 'q4', 0.131, 8192));
Console output appears here (console.log).

Go deeper: the related tool →

This in production, with your data? Let's talk for 15 minutes →