Running a model is not training it
This guide is about using models that are already trained —what is called inference—: chatting, summarising, searching documents or drafting. Training or fine-tuning a model needs far more memory; we cover that in the guide to VRAM for training AI locally.
The rule that decides everything: the model has to fit
A model takes up roughly its number of parameters times the size of each parameter. With 4-bit quantisation, the usual choice for running models locally with very little loss of quality, each parameter takes a little over half a byte. On top of that comes the context memory: the text the model reads and writes for each query.
| Model | Approximate memory | Where it fits |
|---|---|---|
| 8 billion parameters (8B) | about 5 GB | Any GPU with 8 GB or more |
| gpt-oss-20b | about 13 GB | A 16 GB GPU |
| 32 billion (32B) | about 20 GB | A 24 or 32 GB GPU |
| 70 billion (70B) | about 40 GB | Two 32 GB GPUs or a professional 48 GB+ GPU |
| gpt-oss-120b | about 65 GB | An 80 to 96 GB professional GPU, or a 128 GB unified-memory machine |
Dedicated GPU or unified memory
With a dedicated graphics card the model lives in its memory (VRAM) and responds very fast, but the limit is hard: if it does not fit, it does not work well. Consumer cards reach 32 GB; professional ones, 96 GB.
Unified-memory machines share a large memory, up to 128 GB in compact machines, between the processor and the graphics. Much larger models fit in a small, quiet machine, at the cost of generating text more slowly than a high-end dedicated GPU. For one or a few people it is usually a good balance.
How many people and how much text
One person is not ten. When several queries overlap, each needs its own context space and the machine shares out its capacity. That is why, for a team, memory matters as much as GPU power: with long documents and several people at once, a 16 GB GPU runs out long before a 32 GB one.
Indicative configurations
| Scenario | Machine | Price |
|---|---|---|
| Up to 5 people, models up to about 20B | Ryzen 7 9700X · 32 GB · RTX 5060 Ti 16 GB | €2,903 |
| Up to 15 people, models up to about 20B | Ryzen 9 9900X · 64 GB · RTX 5070 Ti 16 GB | €4,727 |
| Up to 30 people or models of about 30B | Ryzen 9 9950X · 64 GB · 2× RTX 5070 Ti 16 GB | €6,509 |
| Up to 60 people or 70B models on GPU | Ryzen 9 9950X · 128 GB · 2× RTX PRO 4000 Blackwell 24 GB | €12,413 |
| Very large models | ACD-014 Enterprise Max · RTX PRO 6000 Blackwell 96 GB | €48,305 |
Prices excl. VAT; the first four come from the configurator. If what you want is the system already installed and configured, with search over your documents and no fees, see private AI for companies.
What we need to know
How many people will use it and how many at once, what kind of documents and how large, whether you need a specific model and whether the machine has to sit on a desk or in a cabinet. With that we propose the machine in writing.
Related hardware and services
Frequently asked questions
Can I use a laptop for local AI?
For small models and testing, yes. For daily work with documents, a desktop copes better with heat and continuous use, and takes more memory.
Which open model should I choose?
It depends on the language, the kind of task and the memory available. We test them with documents similar to yours before recommending one.
How much system RAM do I need?
With a dedicated GPU, at least as much as the GPU memory, and ideally double, to load models and documents comfortably. In unified-memory machines, that memory is the model’s memory.
Shall we help you choose the machine?
Tell us how many people will use it and with which documents. We propose the configuration in writing.