Performance

Atlas runs the model on your machine, so speed is a property of your hardware, not of our servers. These are measured medians from people who opted into anonymous performance reporting — real computers, not our benchmarks on our machines, published whether or not they flatter us.

202 responses · 11 installs · last 90 days · Atlas 2.3.0 · updated 2026-08-08

Generation speed

best result per machine · tokens per second · higher is better

Time to first token

How long before it starts replying · lower is better
Median and 95th-percentile generation speed and time to first token, by hardware, model, and Atlas version.
Hardware Model Atlas tok/s p95 First token
RTX 3090 · 24 GBLinux Qwen3 4Bq4_k_m 2.3.0 86.1 105.3 1.41 s
RTX 3090 · 24 GBLinux Mistral Nemo 12Bq4_k_m 2.3.0 67.2 74.6 8.63 s
RTX 3090 · 24 GBLinux Ministral 3 14Bq4_k_m 2.3.0 56.7 58.0 1.47 s
RTX 3090 · 24 GBLinux · thinking mode Qwen3 14Bq4_k_m 2.3.0 55.3 58.1 6.21 s
Apple M4 Pro · 48 GB unifiedmacOS Qwen3 4B4bit 2.3.0 54.5 54.6 1.62 s
RTX 3090 · 24 GBLinux Qwen3 14Bq4_k_m 2.3.0 52.5 61.8 0.78 s
Apple M4 Pro · 48 GB unifiedmacOS · thinking mode Qwen3 4B4bit 2.3.0 48.6 54.0 12.50 s
Apple M4 Pro · 48 GB unifiedmacOS · thinking mode Qwen3 14B4bit 2.3.0 23.7 27.4 17.84 s
Apple M4 · 16 GB unifiedmacOS Qwen3 4B4bit 2.3.0 23.7 37.6 3.52 s
Apple M3 · 24 GB unifiedmacOS · thinking mode Qwen3 4B4bit 2.3.0 16.9 34.8 36.30 s
Apple M3 · 24 GB unifiedmacOS Qwen3 4B4bit 2.3.0 9.1 35.0 4.40 s

Medians across every reported response, so half of real replies are faster and half slower. Each bar is one machine at its best measured configuration, named beneath it — the table lists every model that machine ran. Time to first token is recorded only for streamed replies, and excludes the first response after a model loads, which carries one-time warmup. Your own speed will vary with model size, context length, and how much of your vault the answer needs. Memory is dedicated graphics memory on a discrete card, and the machine’s total unified memory on Apple Silicon, where one pool is shared between the system and the model. Atlas Local is new, so these figures come from 11 installs in total — treat them as early signal, not a settled benchmark. How this data is collected.

Explore the data

One dot per reported configuration — a machine running a particular model, in a particular mode. Switch the comparison, recolor by any dimension, and click a key to filter.

Does a machine that starts replying quickly also keep replying quickly? Up and to the left is better.

0.50s1.0s2.0s5.0s10s20s50s0.020406080100 Time to first token — seconds (log scale) Generation speed — tokens per second

Key click to filter