Local AI Models: Hardware Tiers and Sizing Guide
A hardware-to-model selection guide mapping local AI models across 8GB to 24GB+ VRAM tiers, quantization footprints, and verified licensing terms.
Selecting viable local AI models requires matching model parameter scale and precision directly to available hardware memory tiers. Quantizing weights from 16-bit to 4-bit precision cuts raw footprint requirements by roughly 75%, allowing capable architectures to execute inside standard consumer hardware.1
VRAM Tiers: Mapping Model Sizes from 8GB to 24GB+
Hardware video memory dictates parameter ceilings for local inferencing. Discrete GPU tiers define clear boundaries for stable execution across parameter sizes.1
An 8GB VRAM tier serves entry-level 3B to 7B models. A 12GB VRAM tier hosts daily-use 8B architectures like Qwen3 and Llama 3.1 8B Instruct, as well as mixture-of-experts builds like Qwen3-Coder 30B (MoE) spanning 12GB to 16GB tiers. Moving up to 16GB VRAM provides capacity for complex 14B to 20B architectures such as Phi-4 and gpt-oss-20b, while 24GB+ pools target power-user workloads like Gemma 4 26B and Qwen3 VL 32B.1
Apple Silicon handles memory differently through unified system pools. Operating system and application overhead shifts the practical boundaries: a 16GB Mac aligns with the 8GB VRAM tier, and a 32GB Mac works within the 16GB VRAM tier.1
Memory Math and Quantization: Sizing Models for Local Execution
Raw weights in full precision (FP16) require approximately 2 bytes per parameter. Loading an unquantized 8-billion parameter model requires roughly 16GB of memory simply to load into address space. Similarly, a 7B parameter model takes roughly 14 GB, and a 70B parameter variant takes around 140 GB.12
Quantizing weight precision down to 4-bit (such as Q4_K_M) shrinks the model footprint by approximately 75%. Under 4-bit compression, a 7B parameter model reduces down to about 4 GB. However, memory planning must account for framework reserves: frameworks and CUDA drivers permanently reserve another 0.5GB to 1GB of memory overhead.12
Memory Bandwidth and Throughput Across Hardware Platforms
Local token generation speeds depend directly on underlying memory bus speeds. A good local setup produces 20 to 80 tokens per second depending on model size and specific hardware configurations.2
Local hardware platforms fall into three distinct memory bandwidth tiers.2
NVIDIA GPUs provide 300 to 1000+ GB/s of bandwidth. Apple Silicon unified memory reaches 100 to 800 GB/s, while host system DDR5 RAM offers 50 to 80 GB/s.2
Model Family Profiles and Benchmark Performance
Open-weight families vary widely in parameter options and verified benchmark results. Meta Llama 3.2 is offered in 1B, 3B, 8B, 70B, and 405B parameter sizes with multimodal options, while Microsoft Phi-3 and Phi-4 target small language model categories at 3.8B and 14B parameters. Google Gemma 2 provides 2B, 9B, and 27B parameter sizes.2
Looking at standardized evaluations, Llama 3.1 8B Instruct achieves a score of 84.5 on GSM8K and 30.4 on GPQA Diamond. Google Gemma 2 9B IT achieves 39.4 on the GPQA Diamond benchmark and 23.8 on GPQA Main.34
Commercial Licensing: Permissive Standards vs. Community Licenses
Model weights carry specific legal licensing boundaries that dictate commercial utility. Phi-4 14B is distributed under the MIT license, while Ministral 3 8B, gpt-oss-20b, Qwen3-Coder 30B (MoE), and Qwen3 VL 32B are licensed under Apache 2.0. Multimodal models like Gemma 4 12B and near-frontier generalist models like Gemma 4 26B also operate under Apache 2.0 following Google's licensing shift in April 2026.1
Older models maintain customized licenses. Google Gemma 2 9B IT is governed under the specific gemma license. Meta Llama models, including Llama 3.1 8B Instruct, operate under a Community license with strict usage restrictions.14
Local Inference Runtimes and Backend Compatibility
Dedicated inference engines determine how model weights load and execute across supported CPU and GPU hardware backends.5
llama.cpp operates as a pure C/C++ engine running GGML and GGUF formats across both CPU and GPU hardware. Koboldcpp provides a C/C++ runtime supporting GGML models with a built-in user interface on both CPU and GPU. For GPU-exclusive setups, ExLlamaV2 provides a Python/C++ inference library optimized for consumer-class GPUs supporting GPTQ and EXL2 formats, while SGLang supports Safetensor, AWQ, and GPTQ formats, delivering 3-5x higher throughput than vLLM via RadixAttention, control flow, and KV cache reuse.5

Broader runtime platforms offer continuous engine integrations. LocalAI released version 4.3.0 in May 2026, enabling the llama.cpp prompt cache by default. In June 2026, LocalAI introduced native C++/ggml biometric backends including voice-detect.cpp for speaker recognition.6
Match your deployment choice directly to your addressable memory: select an 8B model like Ministral 3 8B or Llama 3.1 8B Instruct on 12GB GPUs, step up to Phi-4 14B or gpt-oss-20b on 16GB pools, and reserve 24GB+ hardware for models such as Gemma 4 26B.
Questions
What are the different types of AI models?
Open local models vary by architecture, including standard dense transformers like Llama 3.2 and Gemma 2, mixture-of-experts architectures such as Qwen3-Coder 30B (MoE), and vision-language variants like Qwen3 VL 32B.
What are the best open source AI tools?
Standard execution tools include inference backends like llama.cpp for C/C++ cross-platform support, ExLlamaV2 for GPU environments, SGLang for high-throughput serving, Koboldcpp for combined UI execution, and LocalAI for runtime management.
What is the memory footprint difference between FP16 and 4-bit quantization?
Raw FP16 models require approximately 2 bytes per parameter, meaning an 8B model requires roughly 16GB just to load. Quantizing weights to 4-bit precision reduces the model's memory footprint by approximately 75%.