Skip to content
PocketWebToolsPocketWebTools
PocketWebTools for Mac

Qwen3.5 9B: RAM, speed and settings on a Mac

Qwen3.5 9B is the model PocketWebTools for Mac recommends for 16 GB Macs: a 6.0 GB file that needs about 7.1 GB of memory once it is loaded. Below is what we measured running it in the app on a 16 GB M1 Pro, and the settings the app uses.

Apache 2.0Measured on a 16 GB M1 ProUpdated
6.0GB
Download
7.1GB
Memory once loaded
19 to 19tokens/s
Reply speed
16Ktokens
Context window in the app

Which Mac runs it

The chat models the app ships, by the memory your Mac has. Loaded is the file plus the buffers llama.cpp allocates at the app's context window. Speeds are from a 16 GB M1 Pro; the two 32 GB models ran on a 32 GB Mac mini M4.

ModelMac memoryDownloadLoadedReply speed
Qwen3.8 27B32 GB and up17.6 GB20.4 GB5.3 to 5.8 tokens/s
Qwen3.6 35B‑A3B32 GB and up22.4 GB23.7 GB25.1 to 27.3 tokens/s
Qwen3.5 9BThis page16 GB and up6.0 GB7.1 GB18.6 to 19 tokens/s
Gemma 4 12B16 GB and up7.4 GB8.8 GB13.8 to 14.7 tokens/s
Bonsai 27B (1‑bit)8 GB and up3.8 GB5.1 GB14 to 16.3 tokens/s
Qwen3.5 2B8 GB and up1.3 GB2.0 GB73 to 76 tokens/s
See which one we would pick for your Mac

Qwen3.5 9B speed on an M1 Pro

Measured in the app's own engine: llama.cpp on Metal, every layer on the GPU, other apps closed.

Load, file already on disk
4.9 to 5.5 s
Reply speed
18.6 to 19.0 tokens/s

The fastest of the three models we compared on a 16 GB Mac.

Reading a prompt
About 220 tokens/s

A 774-token prompt took 3.5 s before the first word appeared.

Follow-up in the same chat
0.18 s to the first word

427 of 458 prompt tokens came from the cache.

Our nine-prompt test
9 of 9

Summarizing, rewriting, questions about an 800-token memo, list formatting, TypeScript and a reasoning puzzle. Gemma 4 12B and Bonsai 27B scored 8.5.

Picking clips from a video
2 min 36 s for a 25.7-minute talk

Transcription with Parakeet included. It returned 5 clips.

Settings the app runs Qwen3.5 9B with

The same values work in any llama.cpp front end.

File
Qwen3.5-9B-UD-Q4_K_XL.gguf

Unsloth's 4‑bit dynamic quantization, pinned to one revision and checked against its SHA-256 after download.

Context window
16,384 tokens
GPU layers
All, on Metal
Batch size
512
Temperature
0.7
Top-p / top-k / min-p
0.9 / 40 / 0.05
Thinking
Off

The reply starts after an empty think block, so the answer begins straight away.

Prompt cache
Kept between turns

A follow-up only reads the new message, not the whole chat again.

What the app does with Qwen3.5 9B

Chat

Threads saved on your Mac, replies rendered as Markdown with code blocks and tables, and a meter that shows how much of the 16K window a thread has used.

Clip

The app's first choice for reading a video's transcript and picking the moments worth cutting. A 25.7-minute talk was transcribed and picked in 2 minutes 36 seconds.

Extract, then ask

Send to Chat hands the text of a scan or PDF to the chat screen, so you can ask the model about the document.

Where it falls short

  • Regenerating a reply or editing an earlier message reads the whole chat again. The model keeps recurrent state that cannot be rewound, so only continuing a chat uses the cache. Gemma 4 12B does not have this limit.
  • Long pasted documents still take a while to read. It reads about 220 tokens per second, so the wait for the first word grows with what you paste.
  • It is filed under 16 GB Macs. At 7.1 GB loaded it is too large for an 8 GB Mac, where the app offers Bonsai 27B and Qwen3.5 2B.

Run Qwen3.5 9B in PocketWebTools for Mac

One app for chat, transcription, documents, voices, photos and video, with every model on your Mac. $99 once, free lifetime updates.

See everything in the Mac app

Qwen3.5 9B questions

How much RAM does Qwen3.5 9B need?
About 7.1 GB once loaded: the 6.0 GB file (Unsloth's UD-Q4_K_XL quantization) plus 1.1 GB of buffers at a 16,384-token context, measured from llama.cpp's own log on an M1 Pro. PocketWebTools for Mac files it under 16 GB Macs.
How fast is Qwen3.5 9B on an M1 Pro?
18.6 to 19.0 tokens per second when replying and about 220 tokens per second when reading a prompt, on a 16 GB M1 Pro with every layer on the GPU. A follow-up in the same chat starts in under a fifth of a second.
Qwen3.5 9B or Gemma 4 12B on a 16 GB Mac?
Qwen3.5 9B is our first pick. On the same Mac it replied at 19 tokens per second against 14 to 15 for Gemma 4 12B, scored 9 of 9 on our test against 8.5, and is a smaller download (6.0 GB against 7.4 GB). Gemma 4 12B is worth having as a second opinion, and it can regenerate a reply without reading the chat again.
Does Qwen3.5 9B run on an 8 GB Mac?
Not in the app. It needs about 7.1 GB once loaded, which leaves nothing for macOS on an 8 GB Mac. The app offers Bonsai 27B (5.1 GB loaded) and Qwen3.5 2B (2.0 GB loaded) there.
Can I use Qwen3.5 9B commercially?
Yes. It is released under Apache 2.0, which allows commercial use. The app ships the license text and adds no restrictions of its own.