Make AI Fit
EN

Gemma 3 4B

GemmaGoogle4B8 GB131KLaptop / edge start

Model size details

Family

Gemma

Open the family page when adjacent sizes change the decision.

Maker

Google

Provider or project that owns the model lane.

Size

4B

Needs source review

VRAM

8 GB

Planning band for local quantized runs, not a benchmark guarantee.

Tasks

Chat

Primary local jobs this row should help decide.

Context

131K

Context length affects VRAM and practical throughput.

Runtime path

LM Studio, llama.cpp

Confirm current runtime and quant support before choosing a setup.

Quantization

Q4, Q8, GGUF

The available quant changes memory use, quality, and throughput.

Decision summary and limitation

A compact image-and-text instruction model for local assistant experiments.

Limitation: Its 128K capability is not a promise that a 128K workload fits a particular machine; validate runtime, quantization, and context overhead.

Source, freshness and evidence

Primary source: official-huggingface-google

Freshness: Official model card checked 2026-08-06

Evidence: Official Google model card; Gemma terms require agreement.

Family variants

Adjacent model sizes in the same family.

Open family
Model Tasks VRAM Fit band
Gemma 3 4B Chat 8 GB Laptop / edge start