Local inference profile
- Architecture
- Mixture of experts
- Parameters read per token
- 3B
- Maximum context
- 262,144 tokens
- FP16 KV cache per 1K tokens
- 0.020 GB
Published quantisations
| Format | Effective B/param | Estimated weights |
|---|---|---|
| BF16/FP16 | 2.00 | 70.0 GB |
| Q8_0 | 1.08 | 37.8 GB |
| Q6_K | 0.84 | 29.4 GB |
| Q5_K_M | 0.70 | 24.5 GB |
| Q4_K_M | 0.57 | 19.9 GB |
| Q3_K_M | 0.46 | 16.1 GB |
| Q2_K | 0.34 | 11.9 GB |
Measured benchmarks
These rows calibrate decode and first-token estimates in the calculator.
| Context | Prefill | Decode |
|---|---|---|
| 512 | 3246 tok/s | 136.1 tok/s |
| 2,048 | 3172 tok/s | 134.4 tok/s |
| 4,096 | 3112 tok/s | 131.0 tok/s |
| 8,192 | 3027 tok/s | 126.9 tok/s |
| 16,384 | 2989 tok/s | 118.9 tok/s |
| 32,768 | 2848 tok/s | 103.6 tok/s |
| 65,536 | 2529 tok/s | — tok/s |