Performance

Atlas runs the model on your machine, so speed is a property of your hardware, not of our servers. These are measured medians from people who opted into anonymous performance reporting — real computers, not our benchmarks on our machines, published whether or not they flatter us.

391 responses · 14 installs · last 90 days · Atlas 2.3.0 · updated 2026-08-21

Generation speed

best result per machine · tokens per second · higher is better

Time to first token

How long before it starts replying · lower is better
Median and 95th-percentile generation speed and time to first token, by hardware, model, and Atlas version.
Hardware Model Atlas tok/s p95 First token
RTX 3090 · 24 GBLinux Qwen3 4Bq4_k_m 2.3.0 86.1 105.3 1.41 s
RTX 2080 · 8 GBLinux · thinking mode Qwen3 4Bq4_k_m 2.3.0 72.6 78.7 8.29 s
RTX 2080 · 8 GBLinux Qwen3 4Bq4_k_m 2.3.0 69.5 97.0 1.31 s
RTX 3090 · 24 GBLinux Mistral Nemo 12Bq4_k_m 2.3.0 67.2 74.6 8.63 s
RTX 3090 · 24 GBLinux Ministral 3 14Bq4_k_m 2.3.0 56.7 58.0 1.47 s
RTX 3090 · 24 GBLinux · thinking mode Qwen3 14Bq4_k_m 2.3.0 56.1 61.5 6.21 s
Apple M4 Pro · 48 GB unifiedmacOS Qwen3 4B4bit 2.3.0 54.5 54.6 1.62 s
RTX 3090 · 24 GBLinux Qwen3 14Bq4_k_m 2.3.0 52.5 61.8 0.78 s
Apple M4 Pro · 48 GB unifiedmacOS · thinking mode Qwen3 4B4bit 2.3.0 48.6 54.0 12.50 s
Apple M4 · 16 GB unifiedmacOS Qwen3 4B4bit 2.3.0 24.2 37.9 3.30 s
Apple M4 Pro · 48 GB unifiedmacOS · thinking mode Qwen3 14B4bit 2.3.0 23.7 27.4 17.84 s
Apple M3 · 24 GB unifiedmacOS · thinking mode Qwen3 4B4bit 2.3.0 16.9 34.8 36.30 s
Apple M4 · 16 GB unifiedmacOS Qwen3 8B4bit 2.3.0 15.3 20.6 5.20 s
Apple M3 · 24 GB unifiedmacOS Qwen3 4B4bit 2.3.0 9.1 35.0 4.40 s

Medians across every reported response, so half of real replies are faster and half slower. Each bar is one machine at its best measured configuration, named beneath it — the table lists every model that machine ran. Time to first token is recorded only for streamed replies, and excludes the first response after a model loads, which carries one-time warmup. Your own speed will vary with model size, context length, and how much of your vault the answer needs. Memory is dedicated graphics memory on a discrete card, and the machine’s total unified memory on Apple Silicon, where one pool is shared between the system and the model. Atlas Local is new, so these figures come from 14 installs in total — treat them as early signal, not a settled benchmark. How this data is collected.

Explore the data

One dot per reported configuration — a machine running a particular model, in a particular mode. Describe the machine you have and the matching results stay filled in while the rest hollow out, so you can see which model is worth running on it. Hover any dot for its numbers.

Vendor
Chip
Memory
Model
Mode
System

All 14 reported configurations — pick your machine to highlight it.

Does a machine that starts replying quickly also keep replying quickly? Up and to the left is better.

0.50s1.0s2.0s5.0s10s20s50s0.020406080100 Time to first token — seconds (log scale) Generation speed — tokens per second

Key click to filter

On iPhone and iPad

Atlas on iOS runs Apple’s on-device Foundation Models. The system supplies the model, so there is nothing to pick and nothing to size — the only question is how fast your device answers.

9 responses · 1 install · last 90 days · Atlas 2.3.0 · updated 2026-08-21

Speed against first token

one dot per device · up and to the left is better

How quickly a device starts replying, against how quickly it keeps replying. Hover a dot for its numbers.

2.0s5.0s3638404244 Time to first token — seconds (log scale) Generation speed — tokens per second iPhone 15 Pro

Mode click to filter

Median and 95th-percentile generation speed and time to first token, by device and Atlas version.
Device Model Atlas tok/s p95 First token
iPhone 15 Pro · 6 GBiOS · iPhone16,1 Foundation Models 2.3.0 38.9 95.6 3.25 s

Medians across every reported response, so half of real replies are faster and half slower. On iPhone and iPad, Atlas runs Apple's on-device Foundation Models — the system supplies the model, so unlike the desktop there is nothing to choose, and these are the numbers the device delivers as it ships. Each dot is one device at its best measured result. Time to first token is recorded only for streamed replies, and excludes the first response after a model loads, which carries one-time warmup. Devices report Apple’s hardware identifier rather than a product name — iOS publishes no product name to apps, and the device name it does expose is user-editable, so Atlas never reads it — and the identifier is listed beside the name it was matched to. Your own speed will vary with context length and how much of your vault the answer needs. These figures come from 1 install in total — treat them as early signal, not a settled benchmark. How this data is collected.