Models
Choose models by task and hardware fit.
Start with a concrete task and a memory planning band. Families remain useful metadata, not a navigation detour.
All model sizes
Compare concrete model-size rows by task, VRAM band, and context before opening a size page.
Memory fit
View
Tasks
Llama 3.2 1B
Meta · 1B
A small text instruction model for local chat, rewriting, and retrieval-assistant experiments.
Limitation: A 1B model is not a device-fit guarantee: quantization, context KV cache, runtime, and system overhead change the actual requirement.
Source: Primary model source · Official model card checked 2026-08-06
Official Meta model card; Llama 3.2 Community License applies.
Gemma 3 4B
Google · 4B
A compact image-and-text instruction model for local assistant experiments.
Limitation: Its 128K capability is not a promise that a 128K workload fits a particular machine; validate runtime, quantization, and context overhead.
Source: Primary model source · Official model card checked 2026-08-06
Official Google model card; Gemma terms require agreement.
DeepSeek-R1-Distill-Qwen-7B
Alibaba · 7.6B
A small-to-mid text reasoning model for local exploration, with upstream lineage kept visible.
Limitation: Reasoning output can be long and slow; do not treat this entry as an accuracy, safety, or device-fit claim.
Source: Primary model source · Official model card checked 2026-08-06
Official DeepSeek model card; MIT license with Qwen Apache-2.0 lineage noted.
Qwen 2 5 Coder 7b
Alibaba Qwen · 7B
A 7B text instruction model for local code generation, explanation, review, and fixing experiments.
Limitation: No performance, agent reliability, or memory-fit promise is verified; third-party GGUFs are not canonical evidence for this record.
Source: Primary model source · Official model card and Apache-2.0 license checked 2026-08-06
Official Qwen model card and license.
Qwen 3 14B
Alibaba · 14B
A larger dense text model for general chat and reasoning exploration after local deployment planning.
Limitation: Official sources confirm a 128K dense 14B text model, not a runtime, quant, throughput, or device-fit recommendation; validate those choices locally.
Source: Primary model source · Official Qwen3 model card and release checked 2026-08-06
Official Qwen3 14B model card and release; Apache-2.0.
Llama 3.3 70B
Meta · 70B
A large text instruction model for high-memory local assistant experiments.
Limitation: A 70B model needs a high-memory deployment plan. Review license, attribution, redistribution, quantization, and throughput before selecting it.
Source: Primary model source · Official model card checked 2026-08-06
Official Meta model card; Llama 3.3 Community License applies.
VRAM bands are planning estimates for local fit exploration. Validate context length and sustained performance before making a hardware decision.