“Using int4 means each number is represented using only 4 bits – a 4x reduction in data size compared to BF16.”

Google’s quantized Gemma 3 27B fits in about 14GB, enough for one RTX 3090. Quantization-aware training cuts the quality loss roughly in half. The memory figures cover weights only, and long contexts need more. Running a capable model at home keeps getting cheaper.