How Much VRAM Do You Actually Need for AI & Machine Learning?

How much VRAM do you really need for AI? This guide explains how parameter count, quantisation, context length and workload type affect GPU memory, helping you choose the right GPU for AI and machine learning.

Table of Contents

 

AI & MACHINE LEARNING PC BUYING GUIDE

How Much VRAM Do You Actually Need?

Learn how parameter count, quantisation, context length and workload type affect GPU memory so you can choose the right VRAM capacity for local AI, machine learning, fine-tuning and model inference.

GPU VRAM is one of the most important specifications to check when buying a workstation for artificial intelligence and machine learning. A faster GPU can complete calculations more quickly, but if the model, its cache and the working data do not fit in GPU memory, performance can fall sharply or the workload may fail with an out-of-memory error.

The difficult part is that model size alone does not tell you the complete VRAM requirement. Parameter count sets the starting point, quantisation changes how much memory the model weights consume, context length increases cache requirements during generation, and fine-tuning or training adds gradients, optimizer states and activations. This guide explains how those pieces fit together so UK buyers can choose a realistic GPU memory budget.

✓

Quick Answer

Start with the largest model you expect to run, multiply its parameter count by the approximate bytes used per parameter, then add headroom for KV cache, framework allocations and temporary tensors. As a simple weight estimate, FP16/BF16 uses about 2 bytes per parameter, INT8 about 1 byte, and 4-bit quantisation about 0.5 byte. Real workloads need more than the raw weight figure, especially with long context windows, larger batches or fine-tuning.

GPU Memory First

Why VRAM Matters So Much for AI Workloads

AI applications continually move model weights, activations, caches and working tensors through GPU memory. The amount of VRAM available therefore sets a practical ceiling on the model size, context length, batch size and training method you can use without offloading work to slower system memory.

For local AI, choosing enough VRAM can be more important than buying a faster GPU with insufficient memory for the models you actually use.
▣

Model Fit

The model weights need to fit in memory at the precision or quantisation level you plan to use.

↔

Context Headroom

Longer prompts and generated sequences increase KV-cache memory, so the same model can require more VRAM at larger context lengths.

⚡

Stable Performance

Spare VRAM gives frameworks room for temporary tensors, cache allocations and workload spikes instead of running at the limit.

Parameter Count

How Model Parameters Translate Into VRAM

Parameter count is the easiest starting point for estimating model-weight memory. The figures below are simplified weight-only estimates before KV cache, framework overhead, temporary tensors, adapters or other runtime allocations are added.

Model Size FP16 / BF16 Weights INT8 Weights 4-Bit Weights
3B parameters Approx. 6GB Approx. 3GB Approx. 1.5GB
8B parameters Approx. 16GB Approx. 8GB Approx. 4GB
14B parameters Approx. 28GB Approx. 14GB Approx. 7GB
32B parameters Approx. 64GB Approx. 32GB Approx. 16GB
70B parameters Approx. 140GB Approx. 70GB Approx. 35GB
100B parameters Approx. 200GB Approx. 100GB Approx. 50GB

These numbers are planning estimates, not guaranteed minimums. Quantised formats include metadata and implementation overhead, while inference also needs memory for cache and execution. A GPU that matches the weight-only number exactly can therefore still be too small for the real workload.

Capacity Guide

What Can Different VRAM Capacities Handle?

These categories are practical starting points rather than fixed limits. The exact model that fits depends on architecture, quantisation format, context length, batch size, framework and whether you are running inference, image generation or fine-tuning.

8GB VRAM AI workstation planning Entry AI

8GB VRAM

Suitable for learning, smaller quantised language models, embeddings and lighter GPU-accelerated AI workloads where memory requirements are modest.

8GB guidance →
12GB VRAM AI workstation planning Everyday Local AI

12GB VRAM

A useful entry point for local inference with many small and mid-size quantised models, plus general AI development and moderate image-generation workflows.

12GB guidance →
16GB VRAM AI workstation planning Strong All-Rounder

16GB VRAM

Provides more comfortable headroom for 7B–14B-class quantised models, larger contexts and heavier diffusion or development workflows.

16GB guidance →
24GB VRAM AI workstation planning Advanced Local AI

24GB VRAM

A popular high-capacity target for serious local AI, larger quantised models, higher-resolution generation and selected LoRA or QLoRA fine-tuning workloads.

24GB guidance →
32GB VRAM AI workstation planning Large-Model Headroom

32GB VRAM

Gives 32B-class quantised models substantially more operating room and supports larger contexts, batches and more demanding professional AI workflows.

32GB guidance →
48GB or more VRAM professional AI workstation Professional AI

48GB+ VRAM

Designed for larger models, higher-precision inference, long-context serving, bigger fine-tuning jobs and professional workloads that need substantial memory headroom.

48GB+ guidance →
Parameter count and VRAM requirements for AI models
Factor One

Parameter Count Sets the Starting Point

A model with more parameters stores more learned values, so its weight memory increases roughly in proportion to parameter count. An 8-billion- parameter model therefore needs roughly twice the weight memory of a 4-billion-parameter model when both use the same numeric precision.

  • More parameters normally mean more model-weight memory
  • FP16 or BF16 weights use roughly 2 bytes per parameter
  • Weight memory is only the first part of the VRAM budget
  • Always check the exact model architecture and runtime
Quantisation reducing AI model VRAM requirements
Factor Two

Quantisation Can Change the Answer Completely

Quantisation stores model weights using fewer bits. Moving from 16-bit weights to 8-bit can roughly halve model-weight memory, while 4-bit quantisation can reduce it to around one quarter of the original 16-bit weight footprint. This is why a large model that would never fit at BF16 may become practical for local inference in a 4-bit format.

  • 8-bit quantisation can roughly halve weight memory versus 16-bit
  • 4-bit quantisation can reduce weight memory to roughly one quarter
  • Quantised formats still carry metadata and runtime overhead
  • Lower memory use may involve quality or performance trade-offs
Context length and KV cache GPU memory requirements
Factor Three

Context Length Adds KV-Cache Memory

During autoregressive text generation, the model stores key and value tensors from previous tokens so it does not have to recompute the entire conversation every time a new token is produced. This KV cache grows as the active sequence becomes longer, so a model that runs comfortably at a short context can need noticeably more VRAM with long documents, extended chats or multiple concurrent requests.

  • Longer context usually means a larger KV cache
  • More concurrent sequences or bigger batches increase cache demand
  • Some architectures use sliding or chunked attention to limit growth
  • Long-context users should prioritise extra VRAM headroom
AI inference fine-tuning and training GPU memory
Workload Type

Inference and Fine-Tuning Need Different VRAM Budgets

Inference mainly needs model weights, cache and execution workspace. Training and fine-tuning can need far more because gradients, optimizer states and forward activations must also be stored. Techniques such as LoRA, QLoRA, gradient checkpointing and lower-precision optimizers can reduce memory use, but the required capacity still depends heavily on sequence length, batch size and the exact training method.

  • Inference usually has the lowest memory requirement
  • LoRA and QLoRA can make fine-tuning more memory efficient
  • Full training stores gradients, activations and optimizer states
  • Long sequences and larger batches can drive peak memory much higher
Multi-GPU AI workstation with combined VRAM
Scaling Up

Does Two GPUs Mean Double the Usable VRAM?

Not automatically. Multiple GPUs only behave like one larger memory pool when the software and model are deliberately split across devices using techniques such as tensor parallelism, pipeline parallelism or model sharding. With ordinary data-parallel training, each GPU may hold its own copy of the model, so adding another card improves throughput without simply doubling the maximum model size.

  • Model sharding can spread weights across multiple GPUs
  • Tensor parallelism can split suitable inference workloads
  • Data parallelism often replicates the model on each GPU
  • Software support and interconnect behaviour matter as much as capacity
Buying Advice

How to Work Out the VRAM You Actually Need

Use these four steps before choosing an AI workstation GPU.

1

Name Your Largest Model

Start with the largest parameter count you genuinely expect to run locally rather than sizing the system around today’s smallest job.

2

Choose the Precision

Decide whether you need BF16/FP16, 8-bit or 4-bit weights. This choice can change model-weight memory by several times.

3

Add Context and Workload Headroom

Allow extra memory for KV cache, longer prompts, larger batches, runtime allocations, image tensors or training activations.

4

Leave Room for the Next Project

If two GPU options are close in price, extra VRAM can extend the useful life of an AI workstation as model sizes and context needs grow.

Choose a Cortex AI Workstation Around Your Models

Tell us the model size, framework, quantisation and context length you plan to use. House of Computers can size the GPU, VRAM, system memory and storage around your real AI workload before your workstation is built.

Explore Cortex AI Workstations

Frequently Asked Questions

Is 8GB of VRAM enough for AI?

8GB can be enough for learning, smaller quantised language models, embeddings and lighter AI workloads. It is much more restrictive for larger local models, long contexts, high-resolution diffusion and serious fine-tuning, so users planning to grow into heavier AI work should consider more VRAM where budget allows.

Is 16GB of VRAM enough for local LLMs?

16GB is a strong general-purpose capacity for many local AI users. It can comfortably support numerous 7B- and 14B-class quantised models, although exact fit depends on quantisation, architecture, context length, batch size and runtime overhead.

Can a 24GB GPU run a 32B model?

A 32B model has a theoretical 4-bit weight footprint of about 16GB, so many 4-bit implementations can fit within 24GB with room left for runtime use. However, long context, large batches, model architecture and quantisation overhead can materially increase the requirement. Check the exact model and software before buying.

How much VRAM does a 70B model need?

The raw weight estimate is roughly 140GB at FP16/BF16, 70GB at 8-bit and 35GB at 4-bit. Real inference requires additional memory for cache and runtime overhead, so a 70B-class model usually needs more than the 4-bit weight figure alone and may benefit from a 48GB-class GPU or a properly configured multi-GPU system.

Does context length affect VRAM usage?

Yes. During text generation, the KV cache stores information from previous tokens. Longer active sequences generally require more cache memory, although the exact growth depends on the model architecture and attention implementation.

Do two 24GB GPUs give me 48GB of VRAM?

Only when the workload and software can split the model or tensors across both GPUs. Some multi-GPU methods combine capacity for model sharding, while other methods replicate the model on each card. Two GPUs should therefore not be treated as one 48GB GPU by default.

Final Verdict

The right amount of VRAM is determined by the complete workload, not just the GPU name. Parameter count tells you how large the model weights are, quantisation changes that footprint, context length adds cache memory and fine-tuning introduces additional training tensors. The safest approach is to calculate the weight requirement first and then leave meaningful headroom for everything the model needs while it is actually running.

For buyers choosing between otherwise similar AI workstation configurations, additional VRAM is often one of the most useful long-term upgrades because it allows larger models, longer contexts and more demanding projects to stay on the GPU instead of relying on slower offloading or an early system upgrade.