How Much VRAM Do You Actually Need for AI & Machine Learning?
Table of Contents
How Much VRAM Do You Actually Need?
Learn how parameter count, quantisation, context length and workload type affect GPU memory so you can choose the right VRAM capacity for local AI, machine learning, fine-tuning and model inference.
GPU VRAM is one of the most important specifications to check when buying a workstation for artificial intelligence and machine learning. A faster GPU can complete calculations more quickly, but if the model, its cache and the working data do not fit in GPU memory, performance can fall sharply or the workload may fail with an out-of-memory error.
The difficult part is that model size alone does not tell you the complete VRAM requirement. Parameter count sets the starting point, quantisation changes how much memory the model weights consume, context length increases cache requirements during generation, and fine-tuning or training adds gradients, optimizer states and activations. This guide explains how those pieces fit together so UK buyers can choose a realistic GPU memory budget.
Quick Answer
Start with the largest model you expect to run, multiply its parameter count by the approximate bytes used per parameter, then add headroom for KV cache, framework allocations and temporary tensors. As a simple weight estimate, FP16/BF16 uses about 2 bytes per parameter, INT8 about 1 byte, and 4-bit quantisation about 0.5 byte. Real workloads need more than the raw weight figure, especially with long context windows, larger batches or fine-tuning.
Why VRAM Matters So Much for AI Workloads
AI applications continually move model weights, activations, caches and working tensors through GPU memory. The amount of VRAM available therefore sets a practical ceiling on the model size, context length, batch size and training method you can use without offloading work to slower system memory.
Model Fit
The model weights need to fit in memory at the precision or quantisation level you plan to use.
Context Headroom
Longer prompts and generated sequences increase KV-cache memory, so the same model can require more VRAM at larger context lengths.
Stable Performance
Spare VRAM gives frameworks room for temporary tensors, cache allocations and workload spikes instead of running at the limit.
How Model Parameters Translate Into VRAM
Parameter count is the easiest starting point for estimating model-weight memory. The figures below are simplified weight-only estimates before KV cache, framework overhead, temporary tensors, adapters or other runtime allocations are added.
| Model Size | FP16 / BF16 Weights | INT8 Weights | 4-Bit Weights |
|---|---|---|---|
| 3B parameters | Approx. 6GB | Approx. 3GB | Approx. 1.5GB |
| 8B parameters | Approx. 16GB | Approx. 8GB | Approx. 4GB |
| 14B parameters | Approx. 28GB | Approx. 14GB | Approx. 7GB |
| 32B parameters | Approx. 64GB | Approx. 32GB | Approx. 16GB |
| 70B parameters | Approx. 140GB | Approx. 70GB | Approx. 35GB |
| 100B parameters | Approx. 200GB | Approx. 100GB | Approx. 50GB |
These numbers are planning estimates, not guaranteed minimums. Quantised formats include metadata and implementation overhead, while inference also needs memory for cache and execution. A GPU that matches the weight-only number exactly can therefore still be too small for the real workload.
What Can Different VRAM Capacities Handle?
These categories are practical starting points rather than fixed limits. The exact model that fits depends on architecture, quantisation format, context length, batch size, framework and whether you are running inference, image generation or fine-tuning.
Entry AI8GB VRAM
Suitable for learning, smaller quantised language models, embeddings and lighter GPU-accelerated AI workloads where memory requirements are modest.
8GB guidance →
Everyday Local AI12GB VRAM
A useful entry point for local inference with many small and mid-size quantised models, plus general AI development and moderate image-generation workflows.
12GB guidance →
Strong All-Rounder16GB VRAM
Provides more comfortable headroom for 7B–14B-class quantised models, larger contexts and heavier diffusion or development workflows.
16GB guidance →
Advanced Local AI24GB VRAM
A popular high-capacity target for serious local AI, larger quantised models, higher-resolution generation and selected LoRA or QLoRA fine-tuning workloads.
24GB guidance →
Large-Model Headroom32GB VRAM
Gives 32B-class quantised models substantially more operating room and supports larger contexts, batches and more demanding professional AI workflows.
32GB guidance →
Professional AI48GB+ VRAM
Designed for larger models, higher-precision inference, long-context serving, bigger fine-tuning jobs and professional workloads that need substantial memory headroom.
48GB+ guidance →
Parameter Count Sets the Starting Point
A model with more parameters stores more learned values, so its weight memory increases roughly in proportion to parameter count. An 8-billion- parameter model therefore needs roughly twice the weight memory of a 4-billion-parameter model when both use the same numeric precision.
- More parameters normally mean more model-weight memory
- FP16 or BF16 weights use roughly 2 bytes per parameter
- Weight memory is only the first part of the VRAM budget
- Always check the exact model architecture and runtime

Quantisation Can Change the Answer Completely
Quantisation stores model weights using fewer bits. Moving from 16-bit weights to 8-bit can roughly halve model-weight memory, while 4-bit quantisation can reduce it to around one quarter of the original 16-bit weight footprint. This is why a large model that would never fit at BF16 may become practical for local inference in a 4-bit format.
- 8-bit quantisation can roughly halve weight memory versus 16-bit
- 4-bit quantisation can reduce weight memory to roughly one quarter
- Quantised formats still carry metadata and runtime overhead
- Lower memory use may involve quality or performance trade-offs

Context Length Adds KV-Cache Memory
During autoregressive text generation, the model stores key and value tensors from previous tokens so it does not have to recompute the entire conversation every time a new token is produced. This KV cache grows as the active sequence becomes longer, so a model that runs comfortably at a short context can need noticeably more VRAM with long documents, extended chats or multiple concurrent requests.
- Longer context usually means a larger KV cache
- More concurrent sequences or bigger batches increase cache demand
- Some architectures use sliding or chunked attention to limit growth
- Long-context users should prioritise extra VRAM headroom

Inference and Fine-Tuning Need Different VRAM Budgets
Inference mainly needs model weights, cache and execution workspace. Training and fine-tuning can need far more because gradients, optimizer states and forward activations must also be stored. Techniques such as LoRA, QLoRA, gradient checkpointing and lower-precision optimizers can reduce memory use, but the required capacity still depends heavily on sequence length, batch size and the exact training method.
- Inference usually has the lowest memory requirement
- LoRA and QLoRA can make fine-tuning more memory efficient
- Full training stores gradients, activations and optimizer states
- Long sequences and larger batches can drive peak memory much higher

Does Two GPUs Mean Double the Usable VRAM?
Not automatically. Multiple GPUs only behave like one larger memory pool when the software and model are deliberately split across devices using techniques such as tensor parallelism, pipeline parallelism or model sharding. With ordinary data-parallel training, each GPU may hold its own copy of the model, so adding another card improves throughput without simply doubling the maximum model size.
- Model sharding can spread weights across multiple GPUs
- Tensor parallelism can split suitable inference workloads
- Data parallelism often replicates the model on each GPU
- Software support and interconnect behaviour matter as much as capacity
How to Work Out the VRAM You Actually Need
Use these four steps before choosing an AI workstation GPU.
Name Your Largest Model
Start with the largest parameter count you genuinely expect to run locally rather than sizing the system around today’s smallest job.
Choose the Precision
Decide whether you need BF16/FP16, 8-bit or 4-bit weights. This choice can change model-weight memory by several times.
Add Context and Workload Headroom
Allow extra memory for KV cache, longer prompts, larger batches, runtime allocations, image tensors or training activations.
Leave Room for the Next Project
If two GPU options are close in price, extra VRAM can extend the useful life of an AI workstation as model sizes and context needs grow.
Choose a Cortex AI Workstation Around Your Models
Tell us the model size, framework, quantisation and context length you plan to use. House of Computers can size the GPU, VRAM, system memory and storage around your real AI workload before your workstation is built.
Explore Cortex AI WorkstationsFrequently Asked Questions
Is 8GB of VRAM enough for AI?
8GB can be enough for learning, smaller quantised language models, embeddings and lighter AI workloads. It is much more restrictive for larger local models, long contexts, high-resolution diffusion and serious fine-tuning, so users planning to grow into heavier AI work should consider more VRAM where budget allows.
Is 16GB of VRAM enough for local LLMs?
16GB is a strong general-purpose capacity for many local AI users. It can comfortably support numerous 7B- and 14B-class quantised models, although exact fit depends on quantisation, architecture, context length, batch size and runtime overhead.
Can a 24GB GPU run a 32B model?
A 32B model has a theoretical 4-bit weight footprint of about 16GB, so many 4-bit implementations can fit within 24GB with room left for runtime use. However, long context, large batches, model architecture and quantisation overhead can materially increase the requirement. Check the exact model and software before buying.
How much VRAM does a 70B model need?
The raw weight estimate is roughly 140GB at FP16/BF16, 70GB at 8-bit and 35GB at 4-bit. Real inference requires additional memory for cache and runtime overhead, so a 70B-class model usually needs more than the 4-bit weight figure alone and may benefit from a 48GB-class GPU or a properly configured multi-GPU system.
Does context length affect VRAM usage?
Yes. During text generation, the KV cache stores information from previous tokens. Longer active sequences generally require more cache memory, although the exact growth depends on the model architecture and attention implementation.
Do two 24GB GPUs give me 48GB of VRAM?
Only when the workload and software can split the model or tensors across both GPUs. Some multi-GPU methods combine capacity for model sharding, while other methods replicate the model on each card. Two GPUs should therefore not be treated as one 48GB GPU by default.
Final Verdict
The right amount of VRAM is determined by the complete workload, not just the GPU name. Parameter count tells you how large the model weights are, quantisation changes that footprint, context length adds cache memory and fine-tuning introduces additional training tensors. The safest approach is to calculate the weight requirement first and then leave meaningful headroom for everything the model needs while it is actually running.
For buyers choosing between otherwise similar AI workstation configurations, additional VRAM is often one of the most useful long-term upgrades because it allows larger models, longer contexts and more demanding projects to stay on the GPU instead of relying on slower offloading or an early system upgrade.