PocketWebTools for Mac

Changelog

Every release and model refresh, dated. Every entry here was, or will be, a free update for everyone who bought the app.

  1. Pre-release

    Qwen3.5 9B becomes the recommended model for 16 GB Macs

    • Bake-off on a 16 GB M1 Pro across the three 16 GB-class models: Qwen3.5 9B answered fastest (19 tokens per second), scored best on the test prompts and is the smallest download of the three, so it is now the first pick for 16 GB Macs. Gemma 4 12B stays in the lineup as the alternative.
    • Gemma 4 works: the app now understands its chat format, which the bundled llama.cpp build did not.
    • Gemma 4 no longer runs out of GPU memory at a 16k context on 16 GB Macs (sliding-window layers use a window-sized cache: 0.7 GB instead of 5.4 GB).
    • Follow-up turns in a thread reuse the cached prompt on every model, so the second message in a long chat starts streaming in well under a second instead of re-reading the whole conversation.
  2. Pre-release

    Markdown replies, prompt caching, turn actions

    • Replies render as Markdown with syntax-highlighted code blocks and tables.
    • Copy and regenerate on every reply; stopping a reply keeps the text streamed so far.
    • A context meter shows how much of the model's window the thread has used and warns before it overflows.
    • Prompt caching across turns: the model keeps the conversation in memory between messages instead of re-processing it.
  3. Pre-release

    First build: local chat with six models

    • Chat screen with threads saved on your Mac, and a Models screen that downloads, verifies and loads models on demand.
    • Lineup: Qwen3.8 27B (flagship, 32 GB Macs), Qwen3.6 35B-A3B (fast mixture-of-experts, 32 GB), Gemma 4 12B and Qwen3.5 9B (16 GB), Bonsai 27B 1-bit and Qwen3.5 2B (8 GB). All Apache 2.0.
    • Engine: llama.cpp on Metal with every layer on the GPU, streaming tokens to the interface as they are generated.
    • Downloads resume after an interruption and are checked against the publisher's SHA-256 before a model can load.

Want these in your inbox?

Founding users get the launch-week price first.