Gemma 3 4B
Model size details
Family
Gemma
Open the family page when adjacent sizes change the decision.
Maker
Provider or project that owns the model lane.
Size
4B
Needs source review
VRAM
8 GB
Planning band for local quantized runs, not a benchmark guarantee.
Tasks
Chat
Primary local jobs this row should help decide.
Context
131K
Context length affects VRAM and practical throughput.
Runtime path
LM Studio, llama.cpp
Confirm current runtime and quant support before choosing a setup.
Quantization
Q4, Q8, GGUF
The available quant changes memory use, quality, and throughput.
Decision summary and limitation
A compact image-and-text instruction model for local assistant experiments.
Limitation: Its 128K capability is not a promise that a 128K workload fits a particular machine; validate runtime, quantization, and context overhead.
Source, freshness and evidence
Primary source: official-huggingface-google
Freshness: Official model card checked 2026-08-06
Evidence: Official Google model card; Gemma terms require agreement.
Family variants
Adjacent model sizes in the same family.
| Model | Tasks | VRAM | Fit band |
|---|---|---|---|
| Gemma 3 4B | Chat | 8 GB | Laptop / edge start |