Make AI Fit
EN

Llama 3.2 1B

LlamaMeta1BCPU / tiny128KTiny local

Model size details

Family

Llama

Open the family page when adjacent sizes change the decision.

Maker

Meta

Provider or project that owns the model lane.

Size

1B

Needs source review

VRAM

CPU / tiny

Planning band for local quantized runs, not a benchmark guarantee.

Tasks

Chat

Primary local jobs this row should help decide.

Context

128K

Context length affects VRAM and practical throughput.

Runtime path

Ollama, LM Studio, llama.cpp

Confirm current runtime and quant support before choosing a setup.

Quantization

Q4, Q8, GGUF

The available quant changes memory use, quality, and throughput.

Decision summary and limitation

A small text instruction model for local chat, rewriting, and retrieval-assistant experiments.

Limitation: A 1B model is not a device-fit guarantee: quantization, context KV cache, runtime, and system overhead change the actual requirement.

Source, freshness and evidence

Primary source: official-huggingface-meta-llama

Freshness: Official model card checked 2026-08-06

Evidence: Official Meta model card; Llama 3.2 Community License applies.

Family variants

Adjacent model sizes in the same family.

Open family
Model Tasks VRAM Fit band
Llama 3.2 1B Chat CPU / tiny Tiny local
Llama 3.3 70B Chat 48 GB+ 48GB+ workstation