Start with the job and model
A PC that runs a small text model may be unsuitable for a large video workflow. Write down the outcome: short chat replies, document extraction, one image or a clip with particular dimensions and duration. Identify a model that supports the task and a runtime that supports the model. “Local AI” covers many workloads.
Use the model author's requirements and the matching example. ComfyUI's model documentation explains that workflows can need several weight files and particular loaders. Count all required components before downloading. A small workflow JSON tells you little about the size of its dependencies.
Make a hardware note
Record your operating system, processor, system RAM, graphics-card model, dedicated graphics memory and free disk space. On a unified-memory machine, record total memory and leave headroom for ordinary applications. In Windows, Task Manager can help identify the memory and GPU; check manufacturer specifications if labels are ambiguous.
Check the runtime's current compatibility list rather than assuming any GPU is supported. Ollama's hardware documentation distinguishes NVIDIA, AMD, Apple and other backends. Driver and operating-system requirements matter alongside capacity. Read the installation route for your exact platform before changing software.
Estimate memory without promising a fit
For a simple dense model with eight billion stored parameters, sixteen bits per parameter gives a weight-only estimate of 8 billion × 16 ÷ 8 = 16 billion bytes, roughly 16 GB in decimal units. Four bits gives roughly 4 GB before additional information and overhead. This illustrates arithmetic, not requirements for a named model or a measured allocation.
Quantisation can reduce the weight footprint, but buffers and language-model caches still need room. Hugging Face describes lower-precision loading. The actual file, context setting and runtime behaviour are more useful than “8 GB is enough”. Do not allocate all nominal memory to weights.
Try the smallest relevant workload
Choose one documented configuration and keep its default example unchanged. Check its source, dependencies and licence before downloading. For a text test, use a short prompt and modest output limit. For media, follow supported small settings; arbitrary dimensions or frame counts can be invalid.
Record whether it loads, whether output is produced, peak memory if available and elapsed time. In Ollama, ollama ps reports loaded-model placement across CPU and GPU memory; its FAQ explains the column. Loading successfully does not prove acceptable speed or quality.
Decide from the result
If memory is the obstacle, try a supported smaller model or precision variant, or reduce a configurable context/workload within documented limits. If compatibility is the obstacle, choose a supported runtime route rather than endlessly changing prompts. If an occasional task cannot run acceptably, compare a hosted route's actual terms and costs before buying hardware.
This page is a checklist, not an automatic PC scanner. We have not benchmarked these choices on your machine. A capacity estimate cannot guarantee generation.
Sources
Primary references checked 11 October 2026. This explanation is not a benchmark of a particular model or computer.