Bonsai 27B: RAM, speed and settings on a Mac
Bonsai 27B is PrismML's 27B model with 1‑bit weights: a 3.8 GB file that needs about 5.1 GB of memory once it is loaded. Below is what we measured running it in PocketWebTools for Mac on a 16 GB M1 Pro, and the settings the app uses.
Which Mac runs it
The chat models the app ships, by the memory your Mac has. Loaded is the file plus the buffers llama.cpp allocates at the app's context window. Speeds are from a 16 GB M1 Pro; the two 32 GB models ran on a 32 GB Mac mini M4.
| Model | Mac memory | Download | Loaded | Reply speed |
|---|---|---|---|---|
| Qwen3.8 27B | 32 GB and up | 17.6 GB | 20.4 GB | 5.3 to 5.8 tokens/s |
| Qwen3.6 35B‑A3B | 32 GB and up | 22.4 GB | 23.7 GB | 25.1 to 27.3 tokens/s |
| Qwen3.5 9B | 16 GB and up | 6.0 GB | 7.1 GB | 18.6 to 19 tokens/s |
| Gemma 4 12B | 16 GB and up | 7.4 GB | 8.8 GB | 13.8 to 14.7 tokens/s |
| Bonsai 27B (1‑bit) | 8 GB and up | 3.8 GB | 5.1 GB | 14 to 16.3 tokens/s |
| Qwen3.5 2B | 8 GB and up | 1.3 GB | 2.0 GB | 73 to 76 tokens/s |
Bonsai 27B speed on an M1 Pro
Measured in the app's own engine: llama.cpp on Metal, every layer on the GPU, other apps closed.
- Load, file already on disk
- 2.6 s
- First word of a short reply
- 0.9 to 1.3 s
- Reply speed
- 14 to 16 tokens/s
14.0 to 14.5 in the August 2026 bake-off, 16.3 after the October engine update.
- Reading a prompt
- 77 to 83 tokens/s
A 774-token prompt took 10.0 s before the first word appeared.
- Follow-up in the same chat
- 0.46 s to the first word
436 of 467 prompt tokens came from the cache.
- Our nine-prompt test
- 8.5 of 9
Lost half a point for an 82-word answer to an 80-word limit. Qwen3.5 9B scored 9 of 9 on the same prompts.
Settings the app runs Bonsai 27B with
The same values work in any llama.cpp front end.
- File
- Bonsai-27B-Q1_0.gguf
PrismML's own 1‑bit GGUF, checked against its SHA-256 after download.
- Context window
- 8,192 tokens
- GPU layers
- All 65, on Metal
- Batch size
- 512
- Temperature
- 0.7
- Top-p / top-k / min-p
- 0.9 / 40 / 0.05
- Thinking
- Off
The reply starts after an empty think block, so the answer begins straight away.
- Prompt cache
- Kept between turns
A follow-up only reads the new message, not the whole chat again.
What the app does with Bonsai 27B
Chat
Threads saved on your Mac, replies rendered as Markdown with code blocks and tables, and a meter that shows how much of the 8K window a thread has used.
Clip
Reads a video's transcript and picks the moments worth cutting. On our test podcast it found 2 of 4 planted highlights in 45 seconds. Qwen3.5 2B picks in about 5 seconds, so the app uses the 2B first on 8 GB Macs unless you choose Bonsai.
Extract, then ask
Send to Chat hands the text of a scan or PDF to the chat screen, so you can ask the model about the document.
Where it falls short
- Long prompts are slow. It reads 77 to 83 tokens per second, against about 220 for Qwen3.5 9B and 150 for Gemma 4 12B on the same Mac. A 2,363-token transcript took 30 seconds to read.
- Regenerating a reply or editing an earlier message reads the whole chat again. The model keeps recurrent state that cannot be rewound, so only continuing a chat uses the cache.
- 1‑bit weights cost some accuracy: 8.5 of 9 on our test against 9 of 9 for Qwen3.5 9B, a 6.0 GB download.
- Every number here comes from a 16 GB Mac. On an 8 GB Mac it is a tight fit: the app checks the memory macOS allows before loading and unloads idle models first.
Run Bonsai 27B in PocketWebTools for Mac
One app for chat, transcription, documents, voices, photos and video, with every model on your Mac. $99 once, free lifetime updates.
Bonsai 27B questions
- How much RAM does Bonsai 27B need?
- About 5.1 GB once loaded: the 3.8 GB file plus 1.2 GB of buffers at an 8,192-token context, measured from llama.cpp's own log on an M1 Pro. PocketWebTools for Mac files it under 8 GB Macs. On a 16 GB Mac it leaves room for a speech model alongside it.
- How fast is Bonsai 27B on an M1 Pro?
- 14 to 16 tokens per second when replying and 77 to 83 tokens per second when reading a prompt, on a 16 GB M1 Pro with every layer on the GPU. A short reply starts in about a second; a follow-up in the same chat starts in under half a second.
- Bonsai 27B or Qwen3.5 9B on a 16 GB Mac?
- Qwen3.5 9B. On the same Mac it replied at 19 tokens per second, read prompts about three times faster and scored 9 of 9 on our test against 8.5 for Bonsai 27B. Bonsai 27B is the pick when 6 GB is too much, which means 8 GB Macs.
- Does Bonsai 27B run on stock llama.cpp?
- The 1‑bit Q1_0 file does: support was merged upstream for the CPU on 2026-04-06 and for Metal on 2026-04-08. PrismML's ternary Q2_0 and PQ2_0 files need PrismML's own fork of llama.cpp.
- What about Bonsai 2?
- Bonsai 2 27B is a 5.95 GB ternary file. On the same M1 Pro, using PrismML's llama.cpp fork, it replied at 10.7 tokens per second and held 5.98 GB at an 8K context, against 14.6 tokens per second for Bonsai 27B. It does not load on stock llama.cpp yet, so the app does not ship it.
- Can I use Bonsai 27B commercially?
- Yes. It is released under Apache 2.0, which allows commercial use. The app ships the license text and adds no restrictions of its own.