Llama 3.2 1B
Model size details
Family
Llama
Open the family page when adjacent sizes change the decision.
Maker
Meta
Provider or project that owns the model lane.
Size
1B
Needs source review
VRAM
CPU / tiny
Planning band for local quantized runs, not a benchmark guarantee.
Tasks
Chat
Primary local jobs this row should help decide.
Context
128K
Context length affects VRAM and practical throughput.
Runtime path
Ollama, LM Studio, llama.cpp
Confirm current runtime and quant support before choosing a setup.
Quantization
Q4, Q8, GGUF
The available quant changes memory use, quality, and throughput.
Decision summary and limitation
A small text instruction model for local chat, rewriting, and retrieval-assistant experiments.
Limitation: A 1B model is not a device-fit guarantee: quantization, context KV cache, runtime, and system overhead change the actual requirement.
Source, freshness and evidence
Primary source: official-huggingface-meta-llama
Freshness: Official model card checked 2026-08-06
Evidence: Official Meta model card; Llama 3.2 Community License applies.
Family variants
Adjacent model sizes in the same family.
| Model | Tasks | VRAM | Fit band |
|---|---|---|---|
| Llama 3.2 1B | Chat | CPU / tiny | Tiny local |
| Llama 3.3 70B | Chat | 48 GB+ | 48GB+ workstation |