Qwen3.6 35B‑A3B: RAM, speed and settings on a Mac
Qwen3.6 35B‑A3B is a mixture-of-experts model: 35B parameters, about 3B of them active for each token. The file is 22.4 GB and needs about 23.7 GB of memory at the app's 32K context. Below is what we measured running it on a 32 GB Mac mini M4, and the settings PocketWebTools for Mac uses.
Which Mac runs it
The chat models the app ships, by the memory your Mac has. Loaded is the file plus the buffers llama.cpp allocates at the app's context window. Speeds are from a 16 GB M1 Pro; the two 32 GB models ran on a 32 GB Mac mini M4.
| Model | Mac memory | Download | Loaded | Reply speed |
|---|---|---|---|---|
| Qwen3.8 27B | 32 GB and up | 17.6 GB | 20.4 GB | 5.3 to 5.8 tokens/s |
| Qwen3.6 35B‑A3B | 32 GB and up | 22.4 GB | 23.7 GB | 25.1 to 27.3 tokens/s |
| Qwen3.5 9B | 16 GB and up | 6.0 GB | 7.1 GB | 18.6 to 19 tokens/s |
| Gemma 4 12B | 16 GB and up | 7.4 GB | 8.8 GB | 13.8 to 14.7 tokens/s |
| Bonsai 27B (1‑bit) | 8 GB and up | 3.8 GB | 5.1 GB | 14 to 16.3 tokens/s |
| Qwen3.5 2B | 8 GB and up | 1.3 GB | 2.0 GB | 73 to 76 tokens/s |
Qwen3.6 35B‑A3B speed on a 32 GB Mac mini M4
Measured in the app's own engine: llama.cpp on Metal, every layer on the GPU, other apps closed.
- Load, file already on disk
- 9.1 s
- Reply speed
- 25 to 27 tokens/s
Qwen3.8 27B replied at 5.3 to 5.8 tokens per second on the same Mac.
- Reading a prompt
- About 400 tokens/s
A 774-token prompt took 1.9 s before the first word appeared.
- Follow-up in the same chat
- 0.28 s to the first word
458 of 489 prompt tokens came from the cache.
- Our nine-prompt test
- 9 of 9
Summarizing, rewriting, questions about an 800-token memo, list formatting, TypeScript and a reasoning puzzle. Qwen3.8 27B scored 8.5 on the same Mac.
Settings the app runs Qwen3.6 35B‑A3B with
The same values work in any llama.cpp front end.
- File
- Qwen3.6-35B‑A3B-UD-Q4_K_XL.gguf
Unsloth's 4‑bit dynamic quantization, pinned to one revision and checked against its SHA-256 after download.
- Context window
- 32,768 tokens
- GPU layers
- All, on Metal
- Batch size
- 512
- Temperature
- 0.7
- Top-p / top-k / min-p
- 0.9 / 40 / 0.05
- Thinking
- Off
The reply starts after an empty think block, so the answer begins straight away.
- Prompt cache
- Kept between turns
A follow-up only reads the new message, not the whole chat again.
What the app does with Qwen3.6 35B‑A3B
Chat
Threads saved on your Mac, replies rendered as Markdown with code blocks and tables, and a meter that shows how much of the 32K window a thread has used.
Clip
Reads a video's transcript and picks the moments worth cutting. When it is installed and fits, the app prefers it over the 27B for picking, after the two 16 GB models.
Extract, then ask
Send to Chat hands the text of a scan or PDF to the chat screen. The 32K window holds a long document in one go.
Where it falls short
- Every speed here comes from one 32 GB Mac mini M4. The 8 and 16 GB models on these pages were timed on a 16 GB M1 Pro, an older chip, so the two sets are not directly comparable.
- It is filed under 32 GB Macs, and that is tight. On our 32 GB Mac, macOS let one app use 26.8 GB of graphics memory and this model takes 23.7 GB. The app checks before loading and refuses a model that does not fit, so with other apps open it may ask you to close them first.
- Nine prompts is a small test. It checks everyday writing and reading jobs; it is not a leaderboard.
Run Qwen3.6 35B‑A3B in PocketWebTools for Mac
One app for chat, transcription, documents, voices, photos and video, with every model on your Mac. $99 once, free lifetime updates.
Qwen3.6 35B‑A3B questions
- How much RAM does Qwen3.6 35B‑A3B need?
- About 23.7 GB at a 32,768-token context: the 22.4 GB file (Unsloth's UD-Q4_K_XL quantization) plus 1.3 GB of buffers, measured from llama.cpp's own log on a 32 GB Mac mini M4. PocketWebTools for Mac files it under 32 GB Macs.
- Does Qwen3.6 35B‑A3B run on a 16 GB Mac?
- No. The file alone is 22.4 GB. On a 16 GB Mac the app offers Qwen3.5 9B, a 6.0 GB download that replied at 19 tokens per second on our 16 GB M1 Pro.
- What does A3B mean?
- About 3B active parameters. It is a mixture-of-experts model: all 35B parameters have to be in memory, but each token only runs through roughly 3B of them. That is why it needs the memory of a 35B model while doing the per-token work of a much smaller one.
- How fast is Qwen3.6 35B‑A3B on a Mac?
- 25 to 27 tokens per second when replying and about 400 tokens per second when reading a prompt, on a 32 GB Mac mini M4 with every layer on the GPU. It loads in about 9 seconds and a follow-up in the same chat starts in about a quarter of a second.
- Qwen3.6 35B‑A3B or Qwen3.8 27B?
- Both are filed under 32 GB Macs. Qwen3.8 27B is a dense model in a smaller 17.6 GB file, with every parameter working on every token. Qwen3.6 35B‑A3B is a larger 22.4 GB file that runs about 3B parameters per token. On the same 32 GB Mac mini M4 it replied at 25 to 27 tokens per second against 5.3 to 5.8, and scored 9 of 9 on our test against 8.5. Qwen3.8 27B needs about 3 GB less memory.
- Can I use Qwen3.6 35B‑A3B commercially?
- Yes. It is released under Apache 2.0, which allows commercial use. The app ships the license text and adds no restrictions of its own.