Llama
Llama model size lanes for local AI decisions.
Llama model sizes
Compare Llama rows by task, VRAM band, and context before opening a size page.
Memory fit
View
Tasks
Llama 3.2 1B
Meta · 1B
A small text instruction model for local chat, rewriting, and retrieval-assistant experiments.
Limitation: A 1B model is not a device-fit guarantee: quantization, context KV cache, runtime, and system overhead change the actual requirement.
Source: Primary model source · Official model card checked 2026-08-06
Official Meta model card; Llama 3.2 Community License applies.
Llama 3.3 70B
Meta · 70B
A large text instruction model for high-memory local assistant experiments.
Limitation: A 70B model needs a high-memory deployment plan. Review license, attribution, redistribution, quantization, and throughput before selecting it.
Source: Primary model source · Official model card checked 2026-08-06
Official Meta model card; Llama 3.3 Community License applies.
VRAM bands are planning estimates for local fit exploration. Validate context length and sustained performance before making a hardware decision.