A Q4 variant may store much of a model at roughly four-bit precision instead of sixteen-bit floating point.
Quantisation changes how numerical values are represented. It is not the same as deleting half a model's concepts, and it does not replace matching the model to the task. Hugging Face describes the resource and accuracy trade-off.
Labels such as Q4 and Q8 describe families of lower-precision representations. Extra suffixes identify schemes that can use blocks, scaling information and mixed precision. The label alone does not tell you the exact file size or complete runtime memory use.
Use a format your runtime supports, then compare outputs on material that matters to you. Inspect whether a small model extracts the correct invoice fields before scaling up. Lower precision can make local use more practical, but a smaller file is not a promise of greater speed on every processor. llama.cpp documents potential accuracy loss.
Sources
Primary references checked 11 October 2026. This explanation is not a benchmark of a particular model or computer.