Skip to content
PocketWebToolsPocketWebTools
PocketWebTools for Mac

The best local LLM for your Mac, by RAM

The right local model is mostly decided by how much memory your Mac has. This is the comparison we ran to choose the defaults in PocketWebTools for Mac: six open models, the same prompts for each, on a 16 GB M1 Pro and, for the two 32 GB models, a 32 GB Mac mini M4. It covers the models the app ships, not every model that exists.

Measured on a 16 GB M1 Pro and a 32 GB Mac mini M4Updated

Six models, side by side

Loaded is the file plus the buffers llama.cpp allocates at the context window the app uses. A dash means we have not measured it.

ModelMac memoryLoadedReplyReads promptTest
Qwen3.8 27B32 GB and up20.4 GB5.3 to 5.8About 568.5 of 9
Qwen3.6 35B‑A3BOur pick32 GB and up23.7 GB25.1 to 27.3About 4009 of 9
Qwen3.5 9BOur pick16 GB and up7.1 GB18.6 to 19About 2209 of 9
Gemma 4 12B16 GB and up8.8 GB13.8 to 14.7About 1508.5 of 9
Bonsai 27B (1‑bit)8 GB and up5.1 GB14 to 16.3About 778.5 of 9
Qwen3.5 2BOur pick8 GB and up2.0 GB73 to 76-9 of 9

Reply and Reads prompt are in tokens per second. The two 32 GB models ran on a 32 GB Mac mini M4, a newer chip than the M1 Pro, so compare those two rows with each other.

8 GB Mac: Qwen3.5 2B first, Bonsai 27B when you can wait

Qwen3.5 2B is a 1.3 GB download that needs about 2.0 GB once loaded and replies at 73 to 76 tokens per second, fast enough that the answer is on screen before you finish reading the first line. It scored 9 of 9 on our test prompts, which are everyday jobs: summarize, rewrite, answer questions about a memo, write a small function.

Bonsai 27B is the bigger option: a 27B model with 1‑bit weights that needs about 5.1 GB once loaded. It replies at 14 to 16 tokens per second and reads prompts slowly, about 77 tokens per second, so a long pasted document means a long wait. On an 8 GB Mac, 5.1 GB leaves little room for anything else.

16 GB Mac: Qwen3.5 9B

Qwen3.5 9B won on every column we measured. It replied at 19 tokens per second against 14 to 15 for Gemma 4 12B, read prompts at about 220 tokens per second against 150, scored 9 of 9 against 8.5, and is the smaller download: 6.0 GB against 7.4 GB.

Gemma 4 12B is worth installing as a second opinion, and it has one real advantage: regenerating a reply reuses the cache, where the Qwen models read the whole chat again. It needs one llama.cpp setting to fit at all. With the default full-size sliding-window cache it took 5.4 GB for the cache alone at a 16K context and ran the Mac out of graphics memory; window-sized, the cache takes 0.7 GB.

32 GB Mac: Qwen3.6 35B‑A3B

Qwen3.6 35B‑A3B is the one to install first. On a 32 GB Mac mini M4 it replied at 25 to 27 tokens per second, read prompts at about 400 tokens per second and scored 9 of 9 on our test. It is a mixture-of-experts model: a 22.4 GB file, about 23.7 GB once loaded at a 32K context, that runs roughly 3B of its 35B parameters for each token.

Qwen3.8 27B is a dense model in a smaller 17.6 GB file, about 20.4 GB loaded, with every parameter working on every token. On the same Mac that meant 5.3 to 5.8 tokens per second and 13.8 seconds to read a 774-token prompt. It scored 8.5 of 9: its summary called a planned bus link a rail link. It leaves about 3 GB more memory free, and nine prompts is too small a test to settle which of the two writes better, so it is worth installing if you can wait for the answer.

What a tokens-per-second number leaves out

A 16 GB Mac does not give a model 16 GB

macOS caps how much graphics memory one app may use. On our 16 GB M1 Pro that cap was about 12.7 GB. Add the buffers to the file size before deciding a model fits: Gemma 4 12B is a 7.4 GB file and 8.8 GB loaded.

Reading speed matters as much as reply speed

Before the first word appears the model has to read your whole prompt. On our M1 Pro that ran from 77 tokens per second (Bonsai 27B) to 220 (Qwen3.5 9B), so the same 774-token prompt took 10.0 seconds on one and 3.5 on the other. If you paste documents, compare this column, not the reply speed.

Follow-ups should be nearly instant

With the prompt cache kept between turns, a follow-up in the same chat started in 0.18 to 0.59 seconds on all five models we timed. If every message in a long chat is slow to start, the front end is reading the whole conversation again each time.

Context length is a memory setting

The cache that holds the conversation grows with the context window. The app runs the 8 GB models at 8K, the 16 GB models at 16K and the 32 GB models at 32K for that reason. Raising it costs memory you may not have.

How we tested

  • The test Macs: a 16 GB M1 Pro for the 8 and 16 GB models and a 32 GB Mac mini M4 for the two 32 GB models, both on macOS 26 with other apps closed.
  • One engine: llama.cpp on Metal with every layer on the GPU, the way the app runs it.
  • Nine prompts, the same for every model: a three-turn summarize and rewrite chat, two questions about an 800-token memo, a list-format instruction, a TypeScript function, a reasoning puzzle and an email rewrite. Greedy sampling, so a rerun gives the same answer.
  • Each model at the context window the app uses for it, in the 4‑bit file the app ships (1‑bit for Bonsai 27B).
  • Nine prompts is a small test. It was enough to separate these models for everyday writing and reading; it is not a leaderboard.

Run all six in PocketWebTools for Mac

The app picks the models that fit your Mac, downloads them with one click and runs them with the settings on this page. $99 once, free lifetime updates.

Questions

What is the best local LLM for a 16 GB Mac?
Of the models we measured, Qwen3.5 9B. On a 16 GB M1 Pro it replied at 19 tokens per second, read prompts at about 220 tokens per second, scored 9 of 9 on our test prompts and needs about 7.1 GB once loaded. Gemma 4 12B is the second pick at 14 to 15 tokens per second and 8.8 GB loaded.
What is the best local LLM for an 8 GB Mac?
Qwen3.5 2B for speed: about 2.0 GB loaded and 73 to 76 tokens per second on our M1 Pro. Bonsai 27B is the larger option at about 5.1 GB loaded and 14 to 16 tokens per second, measured on the same 16 GB Mac.
What is the best local LLM for a 32 GB Mac?
Of the two we measured, Qwen3.6 35B‑A3B. On a 32 GB Mac mini M4 it replied at 25 to 27 tokens per second, scored 9 of 9 on our test prompts and needs about 23.7 GB once loaded. Qwen3.8 27B needs less memory, about 20.4 GB, and replied at 5.3 to 5.8 tokens per second on the same Mac.
How much RAM does a local LLM need on a Mac?
The file size plus the buffers the engine allocates, which was 0.7 to 2.8 GB for the six models we measured. In practice: about 2.0 GB for a 2B model, 7.1 GB for a 9B, 8.8 GB for a 12B, 20.4 GB for the 27B and 23.7 GB for the 35B. macOS also caps the graphics memory one app can use, at about 12.7 GB on our 16 GB Mac.
Is a bigger model always better?
Not on the same Mac. On our 16 GB M1 Pro the 9B model beat the 12B on speed and on our test, and the 1‑bit 27B scored 8.5 of 9 where the 9B scored 9. More parameters only help when the model fits with room to spare and keeps enough precision.
Which Mac did you test on?
A 16 GB M1 Pro for the 8 and 16 GB models, and a 32 GB Mac mini M4 for the two 32 GB models. The M4 is a newer chip, so speeds from the two Macs are not directly comparable. The memory figures do not depend on the chip.
Why are only six models compared?
They are the six chat models PocketWebTools for Mac ships, and the only ones we have run through the same test. Each is licensed for commercial use under Apache 2.0.