MINESHOP / TOOLS01 — LOCAL AI

CALCULATE FIRST. THEN CHOOSE HARDWARE.

How much memory
does your LLM need?

The parameter count is only the start. This planner estimates the model weights, the context memory and a runtime reserve separately, so you can check a hardware choice before you buy.

Free · No sign-up · Calculated on your device

Your assumptions

01 / INPUT

Illustrative dense transformers, not a guaranteed configuration of any specific model. Check the values in the model's config file.

Adjust the architecture and KV cache

A quantised KV cache must be supported by your inference engine; metadata and temporary buffers are not modelled separately here.

What this number means.

A

Weights are not all of the memory.

The raw calculation is parameters × bits ÷ 8. Quantised formats also store scales and metadata, and some tensors stay at higher precision. The adjustable overhead is an assumption; if you know it, the real size of the model files is a better starting point.

B

Long context costs extra.

For a classic transformer with a full attention cache we use: 2 × layers × KV heads × head dimension × tokens × parallel sequences × bytes per cache value. Prompt and generated tokens share the same context budget. Sliding windows, MLA, hybrid and other architectures can differ a lot.

C

Several GPUs are not one big memory pool.

Enough total capacity does not mean a model can be split across any cards. The engine, the layer split, the interconnect and the buffers on every GPU matter too. Offloading part of the model to the CPU adds system RAM, and bandwidth can then limit speed. MoE models that keep all experts in memory need more than their active parameters.

FROM ESTIMATE TO SYSTEM

Choose hardware deliberately.

First fix the model, quantisation, context and number of users, then compare offers. For example, 16 GB graphics cards (RTX 5060 Ti, RTX 5070 Ti, RTX 5080) run 8–14B models at 4-bit comfortably, while 70B models need several cards or a professional GPU such as the RTX PRO 6000 Blackwell with 96 GB. Graphics cards. This calculator is provided by the hardware retailer Mineshop.

See AI workstations ↗

Method, limits and sources

All values are in GiB (2³⁰ bytes); the GB figures manufacturers quote are not automatically the same unit. The calculator estimates memory for running a model (inference), not for training, optimiser states or a guaranteed number of tokens per second. Vision encoders, unmodelled activations and engine buffers must be added. The 15 % overhead and the 2 GiB reserve are adjustable starting assumptions, not universal constants.

Also available in German, French and Latvian.

No cookies, no external fonts, no analytics, and this calculator never sends the values you enter anywhere. The host may process technical access data. Updated: 30 September 2026.