Qwen3.5 9B: RAM, speed and settings on a Mac
Qwen3.5 9B is the model PocketWebTools for Mac recommends for 16 GB Macs: a 6.0 GB file that needs about 7.1 GB of memory once it is loaded. Below is what we measured running it in the app on a 16 GB M1 Pro, and the settings the app uses.
Which Mac runs it
The chat models the app ships, by the memory your Mac has. Loaded is the file plus the buffers llama.cpp allocates at the app's context window. Speeds are from a 16 GB M1 Pro; the two 32 GB models ran on a 32 GB Mac mini M4.
| Model | Mac memory | Download | Loaded | Reply speed |
|---|---|---|---|---|
| Qwen3.8 27B | 32 GB and up | 17.6 GB | 20.4 GB | 5.3 to 5.8 tokens/s |
| Qwen3.6 35B‑A3B | 32 GB and up | 22.4 GB | 23.7 GB | 25.1 to 27.3 tokens/s |
| Qwen3.5 9B | 16 GB and up | 6.0 GB | 7.1 GB | 18.6 to 19 tokens/s |
| Gemma 4 12B | 16 GB and up | 7.4 GB | 8.8 GB | 13.8 to 14.7 tokens/s |
| Bonsai 27B (1‑bit) | 8 GB and up | 3.8 GB | 5.1 GB | 14 to 16.3 tokens/s |
| Qwen3.5 2B | 8 GB and up | 1.3 GB | 2.0 GB | 73 to 76 tokens/s |
Qwen3.5 9B speed on an M1 Pro
Measured in the app's own engine: llama.cpp on Metal, every layer on the GPU, other apps closed.
- Load, file already on disk
- 4.9 to 5.5 s
- Reply speed
- 18.6 to 19.0 tokens/s
The fastest of the three models we compared on a 16 GB Mac.
- Reading a prompt
- About 220 tokens/s
A 774-token prompt took 3.5 s before the first word appeared.
- Follow-up in the same chat
- 0.18 s to the first word
427 of 458 prompt tokens came from the cache.
- Our nine-prompt test
- 9 of 9
Summarizing, rewriting, questions about an 800-token memo, list formatting, TypeScript and a reasoning puzzle. Gemma 4 12B and Bonsai 27B scored 8.5.
- Picking clips from a video
- 2 min 36 s for a 25.7-minute talk
Transcription with Parakeet included. It returned 5 clips.
Settings the app runs Qwen3.5 9B with
The same values work in any llama.cpp front end.
- File
- Qwen3.5-9B-UD-Q4_K_XL.gguf
Unsloth's 4‑bit dynamic quantization, pinned to one revision and checked against its SHA-256 after download.
- Context window
- 16,384 tokens
- GPU layers
- All, on Metal
- Batch size
- 512
- Temperature
- 0.7
- Top-p / top-k / min-p
- 0.9 / 40 / 0.05
- Thinking
- Off
The reply starts after an empty think block, so the answer begins straight away.
- Prompt cache
- Kept between turns
A follow-up only reads the new message, not the whole chat again.
What the app does with Qwen3.5 9B
Chat
Threads saved on your Mac, replies rendered as Markdown with code blocks and tables, and a meter that shows how much of the 16K window a thread has used.
Clip
The app's first choice for reading a video's transcript and picking the moments worth cutting. A 25.7-minute talk was transcribed and picked in 2 minutes 36 seconds.
Extract, then ask
Send to Chat hands the text of a scan or PDF to the chat screen, so you can ask the model about the document.
Where it falls short
- Regenerating a reply or editing an earlier message reads the whole chat again. The model keeps recurrent state that cannot be rewound, so only continuing a chat uses the cache. Gemma 4 12B does not have this limit.
- Long pasted documents still take a while to read. It reads about 220 tokens per second, so the wait for the first word grows with what you paste.
- It is filed under 16 GB Macs. At 7.1 GB loaded it is too large for an 8 GB Mac, where the app offers Bonsai 27B and Qwen3.5 2B.
Run Qwen3.5 9B in PocketWebTools for Mac
One app for chat, transcription, documents, voices, photos and video, with every model on your Mac. $99 once, free lifetime updates.
Qwen3.5 9B questions
- How much RAM does Qwen3.5 9B need?
- About 7.1 GB once loaded: the 6.0 GB file (Unsloth's UD-Q4_K_XL quantization) plus 1.1 GB of buffers at a 16,384-token context, measured from llama.cpp's own log on an M1 Pro. PocketWebTools for Mac files it under 16 GB Macs.
- How fast is Qwen3.5 9B on an M1 Pro?
- 18.6 to 19.0 tokens per second when replying and about 220 tokens per second when reading a prompt, on a 16 GB M1 Pro with every layer on the GPU. A follow-up in the same chat starts in under a fifth of a second.
- Qwen3.5 9B or Gemma 4 12B on a 16 GB Mac?
- Qwen3.5 9B is our first pick. On the same Mac it replied at 19 tokens per second against 14 to 15 for Gemma 4 12B, scored 9 of 9 on our test against 8.5, and is a smaller download (6.0 GB against 7.4 GB). Gemma 4 12B is worth having as a second opinion, and it can regenerate a reply without reading the chat again.
- Does Qwen3.5 9B run on an 8 GB Mac?
- Not in the app. It needs about 7.1 GB once loaded, which leaves nothing for macOS on an 8 GB Mac. The app offers Bonsai 27B (5.1 GB loaded) and Qwen3.5 2B (2.0 GB loaded) there.
- Can I use Qwen3.5 9B commercially?
- Yes. It is released under Apache 2.0, which allows commercial use. The app ships the license text and adds no restrictions of its own.