Performance
Atlas runs the model on your machine, so speed is a property of your hardware, not of our servers. These are measured medians from people who opted into anonymous performance reporting — real computers, not our benchmarks on our machines, published whether or not they flatter us.
202 responses · 11 installs · last 90 days · Atlas 2.3.0 · updated 2026-08-08
Generation speed
best result per machine · tokens per second · higher is betterTime to first token
How long before it starts replying · lower is better| Hardware | Model | Atlas | tok/s | p95 | First token |
|---|---|---|---|---|---|
| RTX 3090 · 24 GBLinux | Qwen3 4Bq4_k_m | 2.3.0 | 86.1 | 105.3 | 1.41 s |
| RTX 3090 · 24 GBLinux | Mistral Nemo 12Bq4_k_m | 2.3.0 | 67.2 | 74.6 | 8.63 s |
| RTX 3090 · 24 GBLinux | Ministral 3 14Bq4_k_m | 2.3.0 | 56.7 | 58.0 | 1.47 s |
| RTX 3090 · 24 GBLinux · thinking mode | Qwen3 14Bq4_k_m | 2.3.0 | 55.3 | 58.1 | 6.21 s |
| Apple M4 Pro · 48 GB unifiedmacOS | Qwen3 4B4bit | 2.3.0 | 54.5 | 54.6 | 1.62 s |
| RTX 3090 · 24 GBLinux | Qwen3 14Bq4_k_m | 2.3.0 | 52.5 | 61.8 | 0.78 s |
| Apple M4 Pro · 48 GB unifiedmacOS · thinking mode | Qwen3 4B4bit | 2.3.0 | 48.6 | 54.0 | 12.50 s |
| Apple M4 Pro · 48 GB unifiedmacOS · thinking mode | Qwen3 14B4bit | 2.3.0 | 23.7 | 27.4 | 17.84 s |
| Apple M4 · 16 GB unifiedmacOS | Qwen3 4B4bit | 2.3.0 | 23.7 | 37.6 | 3.52 s |
| Apple M3 · 24 GB unifiedmacOS · thinking mode | Qwen3 4B4bit | 2.3.0 | 16.9 | 34.8 | 36.30 s |
| Apple M3 · 24 GB unifiedmacOS | Qwen3 4B4bit | 2.3.0 | 9.1 | 35.0 | 4.40 s |
Medians across every reported response, so half of real replies are faster and half slower. Each bar is one machine at its best measured configuration, named beneath it — the table lists every model that machine ran. Time to first token is recorded only for streamed replies, and excludes the first response after a model loads, which carries one-time warmup. Your own speed will vary with model size, context length, and how much of your vault the answer needs. Memory is dedicated graphics memory on a discrete card, and the machine’s total unified memory on Apple Silicon, where one pool is shared between the system and the model. Atlas Local is new, so these figures come from 11 installs in total — treat them as early signal, not a settled benchmark. How this data is collected.
Explore the data
One dot per reported configuration — a machine running a particular model, in a particular mode. Switch the comparison, recolor by any dimension, and click a key to filter.
Does a machine that starts replying quickly also keep replying quickly? Up and to the left is better.
Key click to filter