Skip to content
PocketWebToolsPocketWebTools
PocketWebTools for Mac

Gemma 4 12B: RAM, speed and settings on a Mac

Gemma 4 12B is Google's 12B instruct model: a 7.4 GB file that needs about 8.8 GB of memory once it is loaded, with one llama.cpp setting that decides whether it fits a 16 GB Mac at all. Below is what we measured running it in PocketWebTools for Mac on a 16 GB M1 Pro, and the settings the app uses.

Apache 2.0Measured on a 16 GB M1 ProUpdated
7.4GB
Download
8.8GB
Memory once loaded
14 to 15tokens/s
Reply speed
16Ktokens
Context window in the app

Which Mac runs it

The chat models the app ships, by the memory your Mac has. Loaded is the file plus the buffers llama.cpp allocates at the app's context window. Speeds are from a 16 GB M1 Pro; the two 32 GB models ran on a 32 GB Mac mini M4.

ModelMac memoryDownloadLoadedReply speed
Qwen3.8 27B32 GB and up17.6 GB20.4 GB5.3 to 5.8 tokens/s
Qwen3.6 35B‑A3B32 GB and up22.4 GB23.7 GB25.1 to 27.3 tokens/s
Qwen3.5 9B16 GB and up6.0 GB7.1 GB18.6 to 19 tokens/s
Gemma 4 12BThis page16 GB and up7.4 GB8.8 GB13.8 to 14.7 tokens/s
Bonsai 27B (1‑bit)8 GB and up3.8 GB5.1 GB14 to 16.3 tokens/s
Qwen3.5 2B8 GB and up1.3 GB2.0 GB73 to 76 tokens/s
See which one we would pick for your Mac

Gemma 4 12B speed on an M1 Pro

Measured in the app's own engine: llama.cpp on Metal, every layer on the GPU, other apps closed.

Load, file already on disk
2.6 s
Reply speed
13.8 to 14.7 tokens/s

14.1 to 14.7 in the August 2026 bake-off, 13.8 after the October engine update.

Reading a prompt
About 150 tokens/s

A 774-token prompt took 5.1 s before the first word appeared.

Follow-up in the same chat
0.25 s to the first word

395 of 426 prompt tokens came from the cache.

Regenerating a reply
Uses the cache

395 of 426 tokens were reused after rewinding a chat. The Qwen models and Bonsai 27B read the whole chat again.

Our nine-prompt test
8.5 of 9

Lost half a point for attaching a number to the wrong item in a rewrite. Qwen3.5 9B scored 9 of 9 on the same prompts.

Settings the app runs Gemma 4 12B with

The same values work in any llama.cpp front end.

File
gemma-4-12b-it-UD-Q4_K_XL.gguf

Unsloth's 4‑bit dynamic quantization, pinned to one revision and checked against its SHA-256 after download.

Context window
16,384 tokens
Sliding-window cache
Window-sized (swa_full off)

llama.cpp's default full-size cache took 5.4 GB at this context and ran a 16 GB Mac out of graphics memory. Window-sized, it takes 0.7 GB.

GPU layers
All, on Metal
Batch size
512
Temperature
0.7
Top-p / top-k / min-p
0.9 / 40 / 0.05
Thinking
Off

The reply starts after an empty thinking channel, so the answer begins straight away.

Prompt cache
Kept between turns

A follow-up only reads the new message, not the whole chat again.

What the app does with Gemma 4 12B

Chat

Threads saved on your Mac, replies rendered as Markdown with code blocks and tables, and a meter that shows how much of the 16K window a thread has used.

Clip

Reads a video's transcript and picks the moments worth cutting. On our test podcast it found 3 of 4 planted highlights in 45 seconds. It is the app's second choice for picking, after Qwen3.5 9B.

Extract, then ask

Send to Chat hands the text of a scan or PDF to the chat screen. On a 16 GB Mac, the GLM-OCR document model ran in 3.6 seconds with Gemma still loaded.

Where it falls short

  • It is slower than Qwen3.5 9B on the same Mac: 14 to 15 tokens per second against 19, and about 150 tokens per second reading a prompt against 220.
  • It is the largest thing a 16 GB Mac holds comfortably. With Gemma loaded, the app refused to load the 3.2 GB dots.ocr document model beside it and named Gemma as the model to unload.
  • Its chat format was newer than the llama.cpp build we started with, and every reply failed until the app rendered the format itself. In another front end, check that Gemma 4 is supported before blaming the model for a failed reply.

Run Gemma 4 12B in PocketWebTools for Mac

One app for chat, transcription, documents, voices, photos and video, with every model on your Mac. $99 once, free lifetime updates.

See everything in the Mac app

Gemma 4 12B questions

How much RAM does Gemma 4 12B need?
About 8.8 GB once loaded: the 7.4 GB file (Unsloth's UD-Q4_K_XL quantization) plus 1.4 GB of buffers at a 16,384-token context, measured from llama.cpp's own log on an M1 Pro. PocketWebTools for Mac files it under 16 GB Macs.
How fast is Gemma 4 12B on an M1 Pro?
13.8 to 14.7 tokens per second when replying and about 150 tokens per second when reading a prompt, on a 16 GB M1 Pro with every layer on the GPU. A follow-up in the same chat starts in a quarter of a second.
Why does Gemma 4 12B run out of memory on a 16 GB Mac?
Most likely the sliding-window cache. Gemma 4 has 40 sliding-window layers, and llama.cpp's default allocates a full-size cache for each: 5.4 GB at a 16K context, on top of the 7.4 GB file. Turning swa_full off sizes the cache to the window instead, 0.7 GB in our run, and the model then fits.
Gemma 4 12B or Qwen3.5 9B on a 16 GB Mac?
Qwen3.5 9B is our first pick: faster (19 tokens per second against 14 to 15), a smaller download and 9 of 9 on our test against 8.5. Gemma 4 12B is the one to add for a second opinion, and it regenerates a reply from the cache where the Qwen models read the chat again.
Can I use Gemma 4 12B commercially?
Yes. Gemma 4 is released under Apache 2.0, which allows commercial use. The app ships the license text and adds no restrictions of its own.