Gemma 4 12B: RAM, speed and settings on a Mac
Gemma 4 12B is Google's 12B instruct model: a 7.4 GB file that needs about 8.8 GB of memory once it is loaded, with one llama.cpp setting that decides whether it fits a 16 GB Mac at all. Below is what we measured running it in PocketWebTools for Mac on a 16 GB M1 Pro, and the settings the app uses.
Which Mac runs it
The chat models the app ships, by the memory your Mac has. Loaded is the file plus the buffers llama.cpp allocates at the app's context window. Speeds are from a 16 GB M1 Pro; the two 32 GB models ran on a 32 GB Mac mini M4.
| Model | Mac memory | Download | Loaded | Reply speed |
|---|---|---|---|---|
| Qwen3.8 27B | 32 GB and up | 17.6 GB | 20.4 GB | 5.3 to 5.8 tokens/s |
| Qwen3.6 35B‑A3B | 32 GB and up | 22.4 GB | 23.7 GB | 25.1 to 27.3 tokens/s |
| Qwen3.5 9B | 16 GB and up | 6.0 GB | 7.1 GB | 18.6 to 19 tokens/s |
| Gemma 4 12B | 16 GB and up | 7.4 GB | 8.8 GB | 13.8 to 14.7 tokens/s |
| Bonsai 27B (1‑bit) | 8 GB and up | 3.8 GB | 5.1 GB | 14 to 16.3 tokens/s |
| Qwen3.5 2B | 8 GB and up | 1.3 GB | 2.0 GB | 73 to 76 tokens/s |
Gemma 4 12B speed on an M1 Pro
Measured in the app's own engine: llama.cpp on Metal, every layer on the GPU, other apps closed.
- Load, file already on disk
- 2.6 s
- Reply speed
- 13.8 to 14.7 tokens/s
14.1 to 14.7 in the August 2026 bake-off, 13.8 after the October engine update.
- Reading a prompt
- About 150 tokens/s
A 774-token prompt took 5.1 s before the first word appeared.
- Follow-up in the same chat
- 0.25 s to the first word
395 of 426 prompt tokens came from the cache.
- Regenerating a reply
- Uses the cache
395 of 426 tokens were reused after rewinding a chat. The Qwen models and Bonsai 27B read the whole chat again.
- Our nine-prompt test
- 8.5 of 9
Lost half a point for attaching a number to the wrong item in a rewrite. Qwen3.5 9B scored 9 of 9 on the same prompts.
Settings the app runs Gemma 4 12B with
The same values work in any llama.cpp front end.
- File
- gemma-4-12b-it-UD-Q4_K_XL.gguf
Unsloth's 4‑bit dynamic quantization, pinned to one revision and checked against its SHA-256 after download.
- Context window
- 16,384 tokens
- Sliding-window cache
- Window-sized (swa_full off)
llama.cpp's default full-size cache took 5.4 GB at this context and ran a 16 GB Mac out of graphics memory. Window-sized, it takes 0.7 GB.
- GPU layers
- All, on Metal
- Batch size
- 512
- Temperature
- 0.7
- Top-p / top-k / min-p
- 0.9 / 40 / 0.05
- Thinking
- Off
The reply starts after an empty thinking channel, so the answer begins straight away.
- Prompt cache
- Kept between turns
A follow-up only reads the new message, not the whole chat again.
What the app does with Gemma 4 12B
Chat
Threads saved on your Mac, replies rendered as Markdown with code blocks and tables, and a meter that shows how much of the 16K window a thread has used.
Clip
Reads a video's transcript and picks the moments worth cutting. On our test podcast it found 3 of 4 planted highlights in 45 seconds. It is the app's second choice for picking, after Qwen3.5 9B.
Extract, then ask
Send to Chat hands the text of a scan or PDF to the chat screen. On a 16 GB Mac, the GLM-OCR document model ran in 3.6 seconds with Gemma still loaded.
Where it falls short
- It is slower than Qwen3.5 9B on the same Mac: 14 to 15 tokens per second against 19, and about 150 tokens per second reading a prompt against 220.
- It is the largest thing a 16 GB Mac holds comfortably. With Gemma loaded, the app refused to load the 3.2 GB dots.ocr document model beside it and named Gemma as the model to unload.
- Its chat format was newer than the llama.cpp build we started with, and every reply failed until the app rendered the format itself. In another front end, check that Gemma 4 is supported before blaming the model for a failed reply.
Run Gemma 4 12B in PocketWebTools for Mac
One app for chat, transcription, documents, voices, photos and video, with every model on your Mac. $99 once, free lifetime updates.
Gemma 4 12B questions
- How much RAM does Gemma 4 12B need?
- About 8.8 GB once loaded: the 7.4 GB file (Unsloth's UD-Q4_K_XL quantization) plus 1.4 GB of buffers at a 16,384-token context, measured from llama.cpp's own log on an M1 Pro. PocketWebTools for Mac files it under 16 GB Macs.
- How fast is Gemma 4 12B on an M1 Pro?
- 13.8 to 14.7 tokens per second when replying and about 150 tokens per second when reading a prompt, on a 16 GB M1 Pro with every layer on the GPU. A follow-up in the same chat starts in a quarter of a second.
- Why does Gemma 4 12B run out of memory on a 16 GB Mac?
- Most likely the sliding-window cache. Gemma 4 has 40 sliding-window layers, and llama.cpp's default allocates a full-size cache for each: 5.4 GB at a 16K context, on top of the 7.4 GB file. Turning swa_full off sizes the cache to the window instead, 0.7 GB in our run, and the model then fits.
- Gemma 4 12B or Qwen3.5 9B on a 16 GB Mac?
- Qwen3.5 9B is our first pick: faster (19 tokens per second against 14 to 15), a smaller download and 9 of 9 on our test against 8.5. Gemma 4 12B is the one to add for a second opinion, and it regenerates a reply from the cache where the Qwen models read the chat again.
- Can I use Gemma 4 12B commercially?
- Yes. Gemma 4 is released under Apache 2.0, which allows commercial use. The app ships the license text and adds no restrictions of its own.