What is AI inference?
Inference is running a trained AI model on an input to produce an output. It is the computation that happens when you ask a model a question or generate an image.
Glossary
Definitions, examples and related explanations. Understand the language behind the tools, then follow it into a practical guide.
16 entries
Inference is running a trained AI model on an input to produce an output. It is the computation that happens when you ask a model a question or generate an image.
A checkpoint is saved model state, usually including weights. In image workflows, the term often refers to the main model file selected by a loader.
ComfyUI is software for assembling and running AI media workflows through a graph of connected nodes.
The context window is the amount of token context a model can use for a request, subject to the model and runtime configuration.
Diffusion models learn to reverse a noise process and can generate media through repeated refinement steps.
In generative AI, hallucination describes fabricated or unsupported content presented as though it were established information.
Image-to-video generates video frames conditioned on one or more images, sometimes with text or other controls.
An LLM is a machine-learning model trained on large amounts of language data to perform tasks such as generating, classifying or transforming text.
LoRA, or low-rank adaptation, adapts a model by training comparatively small update matrices while keeping the base weights frozen.
A prompt is input used to guide a generative model, including instructions and relevant context. Some models also accept images or other media.
Quantisation represents model numbers with lower precision to reduce storage or memory requirements, with possible changes to output quality.
RAG retrieves relevant material from an external source and supplies it to a model as context for generating an answer.
A token is a unit of input or output processed by a language model. It can be a word fragment, punctuation or another encoded unit.
Training adjusts model parameters using data and an optimisation process. Fine-tuning continues training from an existing trained model.
VRAM is memory available to a graphics processor. Local AI uses accelerator memory for model data and working state.
A workflow is a connected set of processing steps, their settings and inputs. In ComfyUI it is a node graph that can be saved and shared.
No matching entry yet. Try a shorter term or browse another content type.