Model guide

Qwen3.6 35B-A3B

Qwen3.6 · 35B total, 3B active per token · up to 256K context.

Local inference profile

Architecture
Mixture of experts
Parameters read per token
3B
Maximum context
262,144 tokens
FP16 KV cache per 1K tokens
0.020 GB

Published quantisations

FormatEffective B/paramEstimated weights
BF16/FP162.0070.0 GB
Q8_01.0837.8 GB
Q6_K0.8429.4 GB
Q5_K_M0.7024.5 GB
Q4_K_M0.5719.9 GB
Q3_K_M0.4616.1 GB
Q2_K0.3411.9 GB

Measured benchmarks

These rows calibrate decode and first-token estimates in the calculator.

RTX 3090 llama.cpp benchmark · IQ4_XS (4.25 bpw)
ContextPrefillDecode
5123246 tok/s136.1 tok/s
2,0483172 tok/s134.4 tok/s
4,0963112 tok/s131.0 tok/s
8,1923027 tok/s126.9 tok/s
16,3842989 tok/s118.9 tok/s
32,7682848 tok/s103.6 tok/s
65,5362529 tok/s— tok/s