Skip to content
PocketWebToolsPocketWebTools
PocketWebTools for Mac

Qwen3.6 35B‑A3B: RAM, speed and settings on a Mac

Qwen3.6 35B‑A3B is a mixture-of-experts model: 35B parameters, about 3B of them active for each token. The file is 22.4 GB and needs about 23.7 GB of memory at the app's 32K context. Below is what we measured running it on a 32 GB Mac mini M4, and the settings PocketWebTools for Mac uses.

Apache 2.0Measured on a 32 GB Mac mini M4Updated
22.4GB
Download
23.7GB
Memory once loaded
25 to 27tokens/s
Reply speed
32Ktokens
Context window in the app

Which Mac runs it

The chat models the app ships, by the memory your Mac has. Loaded is the file plus the buffers llama.cpp allocates at the app's context window. Speeds are from a 16 GB M1 Pro; the two 32 GB models ran on a 32 GB Mac mini M4.

ModelMac memoryDownloadLoadedReply speed
Qwen3.8 27B32 GB and up17.6 GB20.4 GB5.3 to 5.8 tokens/s
Qwen3.6 35B‑A3BThis page32 GB and up22.4 GB23.7 GB25.1 to 27.3 tokens/s
Qwen3.5 9B16 GB and up6.0 GB7.1 GB18.6 to 19 tokens/s
Gemma 4 12B16 GB and up7.4 GB8.8 GB13.8 to 14.7 tokens/s
Bonsai 27B (1‑bit)8 GB and up3.8 GB5.1 GB14 to 16.3 tokens/s
Qwen3.5 2B8 GB and up1.3 GB2.0 GB73 to 76 tokens/s
See which one we would pick for your Mac

Qwen3.6 35B‑A3B speed on a 32 GB Mac mini M4

Measured in the app's own engine: llama.cpp on Metal, every layer on the GPU, other apps closed.

Load, file already on disk
9.1 s
Reply speed
25 to 27 tokens/s

Qwen3.8 27B replied at 5.3 to 5.8 tokens per second on the same Mac.

Reading a prompt
About 400 tokens/s

A 774-token prompt took 1.9 s before the first word appeared.

Follow-up in the same chat
0.28 s to the first word

458 of 489 prompt tokens came from the cache.

Our nine-prompt test
9 of 9

Summarizing, rewriting, questions about an 800-token memo, list formatting, TypeScript and a reasoning puzzle. Qwen3.8 27B scored 8.5 on the same Mac.

Settings the app runs Qwen3.6 35B‑A3B with

The same values work in any llama.cpp front end.

File
Qwen3.6-35B‑A3B-UD-Q4_K_XL.gguf

Unsloth's 4‑bit dynamic quantization, pinned to one revision and checked against its SHA-256 after download.

Context window
32,768 tokens
GPU layers
All, on Metal
Batch size
512
Temperature
0.7
Top-p / top-k / min-p
0.9 / 40 / 0.05
Thinking
Off

The reply starts after an empty think block, so the answer begins straight away.

Prompt cache
Kept between turns

A follow-up only reads the new message, not the whole chat again.

What the app does with Qwen3.6 35B‑A3B

Chat

Threads saved on your Mac, replies rendered as Markdown with code blocks and tables, and a meter that shows how much of the 32K window a thread has used.

Clip

Reads a video's transcript and picks the moments worth cutting. When it is installed and fits, the app prefers it over the 27B for picking, after the two 16 GB models.

Extract, then ask

Send to Chat hands the text of a scan or PDF to the chat screen. The 32K window holds a long document in one go.

Where it falls short

  • Every speed here comes from one 32 GB Mac mini M4. The 8 and 16 GB models on these pages were timed on a 16 GB M1 Pro, an older chip, so the two sets are not directly comparable.
  • It is filed under 32 GB Macs, and that is tight. On our 32 GB Mac, macOS let one app use 26.8 GB of graphics memory and this model takes 23.7 GB. The app checks before loading and refuses a model that does not fit, so with other apps open it may ask you to close them first.
  • Nine prompts is a small test. It checks everyday writing and reading jobs; it is not a leaderboard.

Run Qwen3.6 35B‑A3B in PocketWebTools for Mac

One app for chat, transcription, documents, voices, photos and video, with every model on your Mac. $99 once, free lifetime updates.

See everything in the Mac app

Qwen3.6 35B‑A3B questions

How much RAM does Qwen3.6 35B‑A3B need?
About 23.7 GB at a 32,768-token context: the 22.4 GB file (Unsloth's UD-Q4_K_XL quantization) plus 1.3 GB of buffers, measured from llama.cpp's own log on a 32 GB Mac mini M4. PocketWebTools for Mac files it under 32 GB Macs.
Does Qwen3.6 35B‑A3B run on a 16 GB Mac?
No. The file alone is 22.4 GB. On a 16 GB Mac the app offers Qwen3.5 9B, a 6.0 GB download that replied at 19 tokens per second on our 16 GB M1 Pro.
What does A3B mean?
About 3B active parameters. It is a mixture-of-experts model: all 35B parameters have to be in memory, but each token only runs through roughly 3B of them. That is why it needs the memory of a 35B model while doing the per-token work of a much smaller one.
How fast is Qwen3.6 35B‑A3B on a Mac?
25 to 27 tokens per second when replying and about 400 tokens per second when reading a prompt, on a 32 GB Mac mini M4 with every layer on the GPU. It loads in about 9 seconds and a follow-up in the same chat starts in about a quarter of a second.
Qwen3.6 35B‑A3B or Qwen3.8 27B?
Both are filed under 32 GB Macs. Qwen3.8 27B is a dense model in a smaller 17.6 GB file, with every parameter working on every token. Qwen3.6 35B‑A3B is a larger 22.4 GB file that runs about 3B parameters per token. On the same 32 GB Mac mini M4 it replied at 25 to 27 tokens per second against 5.3 to 5.8, and scored 9 of 9 on our test against 8.5. Qwen3.8 27B needs about 3 GB less memory.
Can I use Qwen3.6 35B‑A3B commercially?
Yes. It is released under Apache 2.0, which allows commercial use. The app ships the license text and adds no restrictions of its own.