Home
Hardware
Configurator
Services
OEM Guides About us Warranty Contact Versión en español →
Local AI · Hardware · Language models

Computer for local AI: what hardware you need to run language models

You do not need a supercomputer to run an AI model locally, but you do need to understand one rule: the model has to fit in memory. From there, it is a trade-off between speed, model size and number of people.

ACD&CO technical team26 September 20264 min read

Running a model is not training it

This guide is about using models that are already trained —what is called inference—: chatting, summarising, searching documents or drafting. Training or fine-tuning a model needs far more memory; we cover that in the guide to VRAM for training AI locally.

The rule that decides everything: the model has to fit

A model takes up roughly its number of parameters times the size of each parameter. With 4-bit quantisation, the usual choice for running models locally with very little loss of quality, each parameter takes a little over half a byte. On top of that comes the context memory: the text the model reads and writes for each query.

Approximate size of some open models
ModelApproximate memoryWhere it fits
8 billion parameters (8B)about 5 GBAny GPU with 8 GB or more
gpt-oss-20babout 13 GBA 16 GB GPU
32 billion (32B)about 20 GBA 24 or 32 GB GPU
70 billion (70B)about 40 GBTwo 32 GB GPUs or a professional 48 GB+ GPU
gpt-oss-120babout 65 GBAn 80 to 96 GB professional GPU, or a 128 GB unified-memory machine
These are indicative figures for the model itself. With long documents, the context may need several extra GB for each query being served at the same time.

Dedicated GPU or unified memory

With a dedicated graphics card the model lives in its memory (VRAM) and responds very fast, but the limit is hard: if it does not fit, it does not work well. Consumer cards reach 32 GB; professional ones, 96 GB.

Unified-memory machines share a large memory, up to 128 GB in compact machines, between the processor and the graphics. Much larger models fit in a small, quiet machine, at the cost of generating text more slowly than a high-end dedicated GPU. For one or a few people it is usually a good balance.

How many people and how much text

One person is not ten. When several queries overlap, each needs its own context space and the machine shares out its capacity. That is why, for a team, memory matters as much as GPU power: with long documents and several people at once, a 16 GB GPU runs out long before a 32 GB one.

Indicative configurations

ACD&CO machines for local AI
ScenarioMachinePrice
Up to 5 people, models up to about 20BRyzen 7 9700X · 32 GB · RTX 5060 Ti 16 GB€2,903
Up to 15 people, models up to about 20BRyzen 9 9900X · 64 GB · RTX 5070 Ti 16 GB€4,727
Up to 30 people or models of about 30BRyzen 9 9950X · 64 GB · 2× RTX 5070 Ti 16 GB€6,509
Up to 60 people or 70B models on GPURyzen 9 9950X · 128 GB · 2× RTX PRO 4000 Blackwell 24 GB€12,413
Very large modelsACD-014 Enterprise Max · RTX PRO 6000 Blackwell 96 GB€48,305

Prices excl. VAT; the first four come from the configurator. If what you want is the system already installed and configured, with search over your documents and no fees, see private AI for companies.

What we need to know

How many people will use it and how many at once, what kind of documents and how large, whether you need a specific model and whether the machine has to sit on a desk or in a cabinet. With that we propose the machine in writing.

Frequently asked questions

Can I use a laptop for local AI?

For small models and testing, yes. For daily work with documents, a desktop copes better with heat and continuous use, and takes more memory.

Which open model should I choose?

It depends on the language, the kind of task and the memory available. We test them with documents similar to yours before recommending one.

How much system RAM do I need?

With a dedicated GPU, at least as much as the GPU memory, and ideally double, to load models and documents comfortably. In unified-memory machines, that memory is the model’s memory.

Shall we help you choose the machine?

Tell us how many people will use it and with which documents. We propose the configuration in writing.

Request a quote