Make AI Fit
EN

Llama 3.3 70B

LlamaMeta70B48 GB+128K48GB+ workstation

Model size details

Family

Llama

Open the family page when adjacent sizes change the decision.

Maker

Meta

Provider or project that owns the model lane.

Size

70B

Needs source review

VRAM

48 GB+

Planning band for local quantized runs, not a benchmark guarantee.

Tasks

Chat

Primary local jobs this row should help decide.

Context

128K

Context length affects VRAM and practical throughput.

Runtime path

vLLM

Confirm current runtime and quant support before choosing a setup.

Quantization

Q5, FP16

The available quant changes memory use, quality, and throughput.

Decision summary and limitation

A large text instruction model for high-memory local assistant experiments.

Limitation: A 70B model needs a high-memory deployment plan. Review license, attribution, redistribution, quantization, and throughput before selecting it.

Source, freshness and evidence

Primary source: official-huggingface-meta-llama

Freshness: Official model card checked 2026-08-06

Evidence: Official Meta model card; Llama 3.3 Community License applies.

Family variants

Adjacent model sizes in the same family.

Open family
Model Tasks VRAM Fit band
Llama 3.2 1B Chat CPU / tiny Tiny local
Llama 3.3 70B Chat 48 GB+ 48GB+ workstation