Measured results

Benchmarks, not estimates.

Results measured on the lab workstation. For hardware, models, or settings not tested here, use the calculator—and treat its output as an estimate.

Test system

Two discrete GPUs, not unified memory.

Two discrete 24GB VRAM pools; model data is distributed by llama.cpp, not held in one unified 48GB pool.

Hardware

2 × NVIDIA RTX 3090 24GB

Threadripper PRO 3945WX workstation
275 W limit per GPU
Driver NVIDIA 595.84

Runtime

llama.cpp

llama-bench / llama-batched-bench
Build fc6545d (build 50)
Flash Attention enabled

Memory placement

Per configuration

7 configurations are fully GPU-resident. 2 larger MoE configurations offload expert layers to CPU and system RAM.

Measured configurations

Model, quant, runtime, and placement.

Different quants and tuning are shown explicitly; these are practical system results, not a controlled model-quality comparison.

ModelQuantBatch / ubatchPlacementDetail
Llama 3.3 70BQ4_K_S1,024 / 512Fully GPU-residentAll layers on GPU; tensor split 1:1
Llama 3.3 Nemotron Super 49BQ5_K_M4,096 / 256Fully GPU-residentAll layers on GPU; layer split 0.7:1.3
Qwen3.6 27B MTPQ4_K_M2,048 / 2,048Fully GPU-residentAll layers on GPU; tensor split 1:1
Qwen3.5 122B-A10BUD-Q4_K_S4,096 / 256CPU offloadContext: automatic fit with 2,048 MiB free per GPU; concurrency: 20 MoE expert layers on CPU, remaining layers split 65:35
GLM-4.5 Air 106B-A12Bi1-Q4_K_S2,048 / 512CPU offload18 MoE expert layers on CPU; remaining layers split 65:35
Gemma 4 31B QAT MTPQ4_K_M2,048 / 512Fully GPU-residentAll layers on GPU; tensor split 1:1
Mistral Small 3.2 24Bi1-Q5_K_M2,048 / 512Fully GPU-residentAll layers on GPU; tensor split 1:1
Qwen3.8 27B CRACKQ8_02,048 / 2,048Fully GPU-residentAll layers on GPU; tensor split 1:1
Qwen3.8 27B Uncensored OrcaRouterQ6_K2,048 / 2,048Fully GPU-residentAll layers on GPU; tensor split 1:1

Context scaling

PP and TG at prefilled KV-cache depth.

Tokens per second. A dash means that depth was not measured for that configuration.

Prompt processing (PP)

Model02K4K8K16K32K64K128K256K
Llama 3.3 70B786.6746.5721.2673.7598.4486.0
Llama 3.3 Nemotron Super 49B962.7913.8879.8815.0706.9550.4
Qwen3.6 27B MTP1580.01529.11508.61465.61387.41251.81046.6774.1498.1
Qwen3.5 122B-A10B312.4306.3298.8294.7289.9279.6
GLM-4.5 Air 106B-A12B284.7279.2274.3261.0239.9206.0
Gemma 4 31B QAT MTP1517.61425.31392.31325.31213.21025.7788.6541.2332.5
Mistral Small 3.2 24B2284.12216.62153.02038.31855.11553.6
Qwen3.8 27B CRACK1645.01581.51563.41507.41419.91277.61056.0780.8497.2
Qwen3.8 27B Uncensored OrcaRouter1476.41430.11404.61371.81297.21175.1992.2748.2484.9

Token generation (TG)

Model02K4K8K16K32K64K128K256K
Llama 3.3 70B31.9130.2728.9526.7323.3018.50
Llama 3.3 Nemotron Super 49B22.9922.1621.5420.3218.2915.15
Qwen3.6 27B MTP59.0558.7158.1156.9955.0451.2545.0835.8625.41
Qwen3.5 122B-A10B31.1130.9030.7530.3529.3227.36
GLM-4.5 Air 106B-A12B27.6026.1725.0022.4619.5915.55
Gemma 4 31B QAT MTP50.8049.2148.6047.3845.2541.4935.5627.6619.08
Mistral Small 3.2 24B72.7070.8768.8265.4459.9451.02
Qwen3.8 27B CRACK45.0344.7144.3543.6542.4840.3036.3230.3522.65
Qwen3.8 27B Uncensored OrcaRouter48.2847.9547.4146.5045.1442.4238.0931.4222.95

Concurrency

One, two, and four simultaneous streams.

Aggregate TG grows with concurrency while TG per stream falls. Rates are measured tokens per second.

ModelStreamsAggregate PPAggregate TGTG / streamAggregate all
Llama 3.3 70B1750.431.0131.0185.94
Llama 3.3 70B2752.549.8424.92132.03
Llama 3.3 70B4751.862.9315.73161.71
Llama 3.3 Nemotron Super 49B1956.822.6822.6864.96
Llama 3.3 Nemotron Super 49B2962.839.9019.95110.55
Llama 3.3 Nemotron Super 49B4841.554.7913.70145.44
Qwen3.6 27B MTP11483.559.2259.22164.53
Qwen3.6 27B MTP21607.498.3749.19262.93
Qwen3.6 27B MTP41658.2128.7232.18334.28
Qwen3.5 122B-A10B1255.635.2935.2982.96
Qwen3.5 122B-A10B2256.735.2517.6382.97
Qwen3.5 122B-A10B4259.949.9712.49108.22
GLM-4.5 Air 106B-A12B1289.627.1427.1468.57
GLM-4.5 Air 106B-A12B2289.918.569.2849.34
GLM-4.5 Air 106B-A12B4295.827.516.8869.59
Gemma 4 31B QAT MTP11442.150.1050.10140.53
Gemma 4 31B QAT MTP21446.685.7342.87229.94
Gemma 4 31B QAT MTP41473.1114.9928.75298.39
Mistral Small 3.2 24B12170.472.2972.29203.31
Mistral Small 3.2 24B22198.7119.8259.91324.14
Mistral Small 3.2 24B42232.6165.5841.39432.57
Qwen3.8 27B CRACK11504.344.6644.66126.47
Qwen3.8 27B CRACK21652.079.9539.97218.68
Qwen3.8 27B CRACK41715.2130.8632.71340.60
Qwen3.8 27B Uncensored OrcaRouter11376.948.2048.20135.14
Qwen3.8 27B Uncensored OrcaRouter21495.384.9842.49228.92
Qwen3.8 27B Uncensored OrcaRouter41539.4121.4830.37314.76

Methodology and provenance

A short, reproducible protocol.

  • PP: 512-token prompt processing.
  • TG: 256-token generation.
  • KV cache: Q8_0 K/V, with Flash Attention.
  • Runs: built-in warmup, then 3 measured runs; arithmetic mean reported.
  • Context depth: prefilled KV-cache depth before the measured work.
  • Timing boundary: llama-bench excludes tokenization and sampling.

Quants, batch settings, layer placement, and model architectures vary. This corpus compares tuned configurations on one system; it is not a controlled comparison of model quality.

Full narrative, tuning notes, interpretation, and reproducibility details are in the source RTX 3090 article.

Need a hypothetical configuration? Open the calculator; its outputs are estimates, not measurements.