Make AI Fit
EN

Models

Choose models by task and hardware fit.

Start with a concrete task and a memory planning band. Families remain useful metadata, not a navigation detour.

All model sizes

Compare concrete model-size rows by task, VRAM band, and context before opening a size page.

6 sizes3 families

Memory fit

View

Tasks

6 of 6 model sizesall tasksAll VRAM

Llama 3.2 1B

Meta · 1B

CPU / tiny

A small text instruction model for local chat, rewriting, and retrieval-assistant experiments.

Limitation: A 1B model is not a device-fit guarantee: quantization, context KV cache, runtime, and system overhead change the actual requirement.

Source: Primary model source · Official model card checked 2026-08-06

Official Meta model card; Llama 3.2 Community License applies.

Gemma 3 4B

Google · 4B

8 GB

A compact image-and-text instruction model for local assistant experiments.

Limitation: Its 128K capability is not a promise that a 128K workload fits a particular machine; validate runtime, quantization, and context overhead.

Source: Primary model source · Official model card checked 2026-08-06

Official Google model card; Gemma terms require agreement.

DeepSeek-R1-Distill-Qwen-7B

Alibaba · 7.6B

8 GB

A small-to-mid text reasoning model for local exploration, with upstream lineage kept visible.

Limitation: Reasoning output can be long and slow; do not treat this entry as an accuracy, safety, or device-fit claim.

Source: Primary model source · Official model card checked 2026-08-06

Official DeepSeek model card; MIT license with Qwen Apache-2.0 lineage noted.

Qwen 2 5 Coder 7b

Alibaba Qwen · 7B

8 GB

A 7B text instruction model for local code generation, explanation, review, and fixing experiments.

Limitation: No performance, agent reliability, or memory-fit promise is verified; third-party GGUFs are not canonical evidence for this record.

Source: Primary model source · Official model card and Apache-2.0 license checked 2026-08-06

Official Qwen model card and license.

Qwen 3 14B

Alibaba · 14B

16 GB

A larger dense text model for general chat and reasoning exploration after local deployment planning.

Limitation: Official sources confirm a 128K dense 14B text model, not a runtime, quant, throughput, or device-fit recommendation; validate those choices locally.

Source: Primary model source · Official Qwen3 model card and release checked 2026-08-06

Official Qwen3 14B model card and release; Apache-2.0.

Llama 3.3 70B

Meta · 70B

48 GB+

A large text instruction model for high-memory local assistant experiments.

Limitation: A 70B model needs a high-memory deployment plan. Review license, attribution, redistribution, quantization, and throughput before selecting it.

Source: Primary model source · Official model card checked 2026-08-06

Official Meta model card; Llama 3.3 Community License applies.

VRAM bands are planning estimates for local fit exploration. Validate context length and sustained performance before making a hardware decision.