Glossary
Glossary terms that lower the jargon barrier without flattening the nuance.
A compact index for local AI terms. Search, jump by first letter, or scan by category without turning every guide into a wall of definitions.
Glossary group
Hardware terms
Thunderbolt, USB4, and OCuLink
External connectivity options that shape whether an external GPU path is practical or fragile.
Thunderbolt, USB4, and OCuLink are high-speed connection paths that can carry external storage, displays, or GPU traffic, but they differ in bandwidth, compatibility, and setup expectations.
Tokens per second
A rough way to feel whether a local model interaction will stay comfortable enough for real use.
Tokens per second is a practical performance signal that tells you how quickly a model can generate output during inference.
Unified Memory
A Mac/Apple Silicon buying concept that affects how much memory is available to both CPU and GPU-style workloads.
Unified memory is shared system memory used by Apple Silicon CPU, GPU, and neural-engine workflows instead of separate system RAM and GPU VRAM.
VRAM
Memory on the graphics side that often becomes the real limit before people realize it.
VRAM is the fast memory a GPU uses to hold models and inference work while they run.
Glossary group
Models terms
Embedding
A numeric representation of text or other content used for search, clustering, and retrieval workflows.
An embedding is a vector representation that lets a retrieval system compare pieces of content by meaning rather than only by exact words.
GGUF
A common local model file format used across quantized local AI workflows.
GGUF is a model file format commonly used for quantized local models.
Glossary group
Runtime / formats terms
Local-first
A stance about where the default value should happen, not a purity test.
Local-first means the setup prioritizes running, storing, or processing key parts locally by default, even if some cloud use still exists around the edges.
Quantization
One of the main ways local models get smaller and more practical to run.
Quantization means storing or running a model in a lower-precision format to reduce memory use and make local inference more practical.
Runtime
The execution layer that actually loads and runs a local model.
A runtime is the engine or service responsible for loading a model and serving responses locally.
Glossary group
Guides terms
Context window
How much text or prompt history a model can consider in one go.
Context window is the amount of input history or working text a model can keep in view during one request.
RAG
Short for retrieval-augmented generation, but the practical question is whether it improves a real document workflow.
RAG combines a model with a retrieval step that brings in relevant external documents or knowledge before generating an answer.