Model guide

Gemma 4 E2B

Gemma 4 · 5B dense · up to 128K context.

Local inference profile

Architecture
Dense
Parameters read per token
2B
Maximum context
131,072 tokens
FP16 KV cache per 1K tokens
0.015 GB

Published quantisations

FormatEffective B/paramEstimated weights
BF16/FP162.0010.0 GB
Q8_01.085.4 GB
Q6_K0.844.2 GB
Q5_K_M0.703.5 GB
Q4_K_M0.572.8 GB
Q3_K_M0.462.3 GB
Q2_K0.341.7 GB