Tools

Local AI Calculator

Choose hardware or a model first. Then inspect memory fit, calibrated decode speed, DDR offload, and compatible alternatives.

HardwareArchitecture names such as Ampere identify a GPU generation. Memory capacity determines what fits; memory bandwidth strongly influences decode speed.

24 GB · 936 GB/s · Ampere

Used $800–$1,000 · 26.7 GB VRAM/$1k · Read the hardware guide

Primary use
1510203050100200400

Recommended for interactive chat: 12-30+ tok/s

Only models meeting this estimated decode speed are shown.

What to run on RTX 3090

Three picks from the curated open-weight catalog that fit this system and speed filter. Select up to three to compare.

Fastest at Q4

Gemma 4 E2B details

Q4_K_M234–294 tok/sExpected: 263 tok/s

Suitable for short, straightforward conversations

The quickest eligible model at the balanced Q4 setting.

Best balance

Qwen3.6 35B-A3B details

Q4_K_M146–184 tok/sExpected: 164 tok/s

Stronger instruction following and reasoning

Balances model capability with responsive decode.

Largest viable

Qwen3 32B details

Q4_K_M29–37 tok/sExpected: 33 tok/s

Slower, but with higher dense-model capacity

The largest model this system can run at a usable single-stream decode speed.

About these estimates

tok/s is estimated single-stream decode speed after prompt processing — not multi-user throughput. Prompt speed and time to first token appear only for exact benchmark-covered pairings.

Speed and memory start from published model size, KV-cache structure, device capacity and bandwidth, then apply architecture, runtime, selected DDR generation/channel count, partial offload, and multi-device factors. The confidence label flags interpolation.

Prices are indicative US whole-system street-price bands, updated July 2026; recommendation cards identify used-component mixes. Verify any shortlist on your own workload before buying.