A GPU with 8 GB of dedicated memory is different from a PC with 32 GB of system RAM.
For a discrete graphics card, dedicated memory is separate from ordinary system RAM. Some computers use a unified memory arrangement instead, so interpret the hardware report for that machine. A GPU model name alone is not a usable capacity check.
Model weights are only part of the allocation. The runtime needs working buffers, and language-model attention caches can become substantial for long contexts. Hugging Face explains cache memory trade-offs.
Leave room for the display and other running applications. If a runtime moves parts of a model into system memory, it may fit but behave differently from an entirely GPU-resident run. Check actual allocation and supported hardware instead of relying on a file-size comparison. Ollama's FAQ explains its CPU/GPU placement display.
Sources
Primary references checked 11 October 2026. This explanation is not a benchmark of a particular model or computer.