Back to LLM catalog

Small Language Models — Local Dev Catalog

Small language models, held to the same standard. Efficiency isn't an excuse for opacity.

Models

13

CPU-Only Models

12

Min RAM Required

0.5GB

Frameworks

6

Why SLMs for Local Dev?

You don't need a data center to build AI-powered apps. These models run on a laptop, a Raspberry Pi, or even in a browser — no GPU required, no cloud costs, no API keys.

Get Started in Minutes

Install Ollama or LM Studio, pick a model below, and run the quick-start command. Your first local LLM in under 5 minutes.

Privacy by Default

All inference runs locally. Your data never leaves your machine — ideal for sensitive domains, offline environments, or just personal projects.

Family:
RAM:
Use Case:

13 models — sorted by parameter count (smallest first)

SmolLM2360M
HuggingFace
0.36B
params

Runs in a browser via WebAssembly. The go-to for truly embedded or browser-native AI experiences.

Browser AIUltra CompactEmbedded
RAM
0.5GB
VRAM
CPU OK
Context
8K
Qwen 2.50.5B
Alibaba
0.5B
params

Half a billion parameters — the most resource-efficient model in the catalog. Surprisingly capable for simple tasks.

Minimal ResourcesCPU FriendlyMicro
RAM
1GB
VRAM
CPU OK
Context
33K
Gemma 41B
Google DeepMind
1B
params

Google's latest ultra-compact Gemma 4 with multimodal vision support. Designed for edge and on-device deployment.

Ultra CompactMultimodalEdge AI
RAM
2GB
VRAM
CPU OK
Context
33K
Gemma 31B
Google DeepMind
1B
params

The smallest Gemma 3 — perfect for truly constrained environments. Runs on a Raspberry Pi 5.

Ultra CompactEdge AICPU Friendly
RAM
2GB
VRAM
CPU OK
Context
33K
Llama 3.21B
Meta
1B
params

Meta's official 1B model with an extraordinary 128K context window. Optimized for on-device and mobile deployment.

Long ContextEdge AICPU Friendly
RAM
2GB
VRAM
CPU OK
Context
131K
TinyLlama1.1B
StatNLP Research
1.1B
params

The classic entry-point SLM. Extremely fast on CPU, minimal RAM usage — great for learning and prototyping pipelines.

Beginner FriendlyUltra FastCPU Friendly
RAM
2GB
VRAM
CPU OK
Context
2K
Qwen 2.51.5B
Alibaba
1.5B
params

A solid 1.5B model with surprisingly strong multilingual capabilities and a 32K context window.

MultilingualCPU FriendlyCompact
RAM
2GB
VRAM
CPU OK
Context
33K
SmolLM21.7B
HuggingFace
1.7B
params

HuggingFace's best-in-class small model for its size. Trained on high-quality curated data for impressive coherence.

High QualityCPU FriendlyCompact
RAM
2GB
VRAM
CPU OK
Context
8K
Qwen 2.53B
Alibaba
3B
params

Excellent instruction-following at 3B. One of the best multilingual SLMs for local development.

MultilingualCodingCPU Friendly
RAM
4GB
VRAM
CPU OK
Context
33K
Llama 3.23B
Meta
3B
params

Meta's 3B model balances capability and footprint well. Great starting point for RAG and agentic pipelines on local hardware.

RAG ReadyLong ContextCPU Friendly
RAM
4GB
VRAM
CPU OK
Context
131K
Phi-4Mini (3.8B)
Microsoft
3.8B
params

Exceptionally capable small model from Microsoft. Punches well above its weight class on reasoning and coding tasks.

Best for BeginnersCodingCPU Friendly
RAM
4GB
VRAM
CPU OK
Context
16K
Phi-3Mini (3.8B)
Microsoft
3.8B
params

Microsoft's compact powerhouse with a massive 128K context window. Ideal for long-document tasks on consumer hardware.

Long ContextCodingCPU Friendly
RAM
4GB
VRAM
CPU OK
Context
128K
Gemma 44B
Google DeepMind
4B
params

Balanced multimodal SLM from the Gemma 4 family. Strong vision-language understanding in a 4B footprint.

MultimodalVisionBalanced
RAM
5GB
VRAM
4GB
Context
33K