Performance
Atlas runs the model on your machine, so speed is a property of your hardware, not of our servers. These are measured medians from people who opted into anonymous performance reporting — real computers, not our benchmarks on our machines, published whether or not they flatter us.
391 responses · 14 installs · last 90 days · Atlas 2.3.0 · updated 2026-08-21
Generation speed
best result per machine · tokens per second · higher is betterTime to first token
How long before it starts replying · lower is better| Hardware | Model | Atlas | tok/s | p95 | First token |
|---|---|---|---|---|---|
| RTX 3090 · 24 GBLinux | Qwen3 4Bq4_k_m | 2.3.0 | 86.1 | 105.3 | 1.41 s |
| RTX 2080 · 8 GBLinux · thinking mode | Qwen3 4Bq4_k_m | 2.3.0 | 72.6 | 78.7 | 8.29 s |
| RTX 2080 · 8 GBLinux | Qwen3 4Bq4_k_m | 2.3.0 | 69.5 | 97.0 | 1.31 s |
| RTX 3090 · 24 GBLinux | Mistral Nemo 12Bq4_k_m | 2.3.0 | 67.2 | 74.6 | 8.63 s |
| RTX 3090 · 24 GBLinux | Ministral 3 14Bq4_k_m | 2.3.0 | 56.7 | 58.0 | 1.47 s |
| RTX 3090 · 24 GBLinux · thinking mode | Qwen3 14Bq4_k_m | 2.3.0 | 56.1 | 61.5 | 6.21 s |
| Apple M4 Pro · 48 GB unifiedmacOS | Qwen3 4B4bit | 2.3.0 | 54.5 | 54.6 | 1.62 s |
| RTX 3090 · 24 GBLinux | Qwen3 14Bq4_k_m | 2.3.0 | 52.5 | 61.8 | 0.78 s |
| Apple M4 Pro · 48 GB unifiedmacOS · thinking mode | Qwen3 4B4bit | 2.3.0 | 48.6 | 54.0 | 12.50 s |
| Apple M4 · 16 GB unifiedmacOS | Qwen3 4B4bit | 2.3.0 | 24.2 | 37.9 | 3.30 s |
| Apple M4 Pro · 48 GB unifiedmacOS · thinking mode | Qwen3 14B4bit | 2.3.0 | 23.7 | 27.4 | 17.84 s |
| Apple M3 · 24 GB unifiedmacOS · thinking mode | Qwen3 4B4bit | 2.3.0 | 16.9 | 34.8 | 36.30 s |
| Apple M4 · 16 GB unifiedmacOS | Qwen3 8B4bit | 2.3.0 | 15.3 | 20.6 | 5.20 s |
| Apple M3 · 24 GB unifiedmacOS | Qwen3 4B4bit | 2.3.0 | 9.1 | 35.0 | 4.40 s |
Medians across every reported response, so half of real replies are faster and half slower. Each bar is one machine at its best measured configuration, named beneath it — the table lists every model that machine ran. Time to first token is recorded only for streamed replies, and excludes the first response after a model loads, which carries one-time warmup. Your own speed will vary with model size, context length, and how much of your vault the answer needs. Memory is dedicated graphics memory on a discrete card, and the machine’s total unified memory on Apple Silicon, where one pool is shared between the system and the model. Atlas Local is new, so these figures come from 14 installs in total — treat them as early signal, not a settled benchmark. How this data is collected.
Explore the data
One dot per reported configuration — a machine running a particular model, in a particular mode. Describe the machine you have and the matching results stay filled in while the rest hollow out, so you can see which model is worth running on it. Hover any dot for its numbers.
All 14 reported configurations — pick your machine to highlight it.
Does a machine that starts replying quickly also keep replying quickly? Up and to the left is better.
Key click to filter
On iPhone and iPad
Atlas on iOS runs Apple’s on-device Foundation Models. The system supplies the model, so there is nothing to pick and nothing to size — the only question is how fast your device answers.
9 responses · 1 install · last 90 days · Atlas 2.3.0 · updated 2026-08-21
Speed against first token
one dot per device · up and to the left is betterHow quickly a device starts replying, against how quickly it keeps replying. Hover a dot for its numbers.
Mode click to filter
| Device | Model | Atlas | tok/s | p95 | First token |
|---|---|---|---|---|---|
| iPhone 15 Pro · 6 GBiOS · iPhone16,1 | Foundation Models | 2.3.0 | 38.9 | 95.6 | 3.25 s |
Medians across every reported response, so half of real replies are faster and half slower. On iPhone and iPad, Atlas runs Apple's on-device Foundation Models — the system supplies the model, so unlike the desktop there is nothing to choose, and these are the numbers the device delivers as it ships. Each dot is one device at its best measured result. Time to first token is recorded only for streamed replies, and excludes the first response after a model loads, which carries one-time warmup. Devices report Apple’s hardware identifier rather than a product name — iOS publishes no product name to apps, and the device name it does expose is user-editable, so Atlas never reads it — and the identifier is listed beside the name it was matched to. Your own speed will vary with context length and how much of your vault the answer needs. These figures come from 1 install in total — treat them as early signal, not a settled benchmark. How this data is collected.