4 Things to Know About Running Llama 3.1 8B at 4-bit
16GB of RAM is enough for Llama 3.1 8B at 4-bit quantization. The model uses about 5-6GB, leaving 9-10GB free for your OS and light multitasking.
8GB of RAM will make the model borderline unusable. After loading the model, you’ll have only about 2GB free, causing heavy swapping and severe slowdowns.
With 8GB, you can try a smaller quant (3-bit) or a smaller model (7B or 3B), but the 8B at 4-bit on 8GB is a bad experience. Don’t expect smooth performance.
If you’re buying RAM, 32GB is better for larger models or heavier multitasking. But 16GB is the minimum for comfortable use with this specific model.