What a Mac unlocks
Same privacy as the free tools. Different ceiling.
Every tool on pocketweb.tools already runs on your device. What a browser cannot do is hold a large model: a tab tops out at about 4 GB of memory, which is enough for 2B to 9B-class models and 1-bit tricks, not for a dense 27B. The Mac app uses your Mac's actual memory, so the lineup goes where a tab cannot.
| In your browser (free) | PocketWebTools for Mac | |
|---|---|---|
| Model size | Under about 4 GB | Up to 22 GB on a 32 GB Mac |
| Largest model in the lineup | Qwen3.5 2B, Bonsai 27B 1-bit | Qwen3.8 27B dense, Qwen3.6 35B-A3B |
| Memory available | About 4 GB per tab | Your Mac's RAM, with tiers for 8, 16 and 32 GB |
| Engine | WebGPU or WebAssembly | llama.cpp on Metal, every layer on the GPU |
| Where it runs | On your device | On your device |
| Works offline | After the model is cached | Always, after the download |
| Where models live | Browser cache, can be evicted | Application Support, until you delete them |
| Price | Free, unlimited | $29 once, free lifetime updates |
| Account | None | None. A license key from the purchase email. |
What's inside
Version 0.1.0
Chat with a local model
Threads saved on your Mac, Markdown replies with highlighted code, copy and regenerate on every turn, and a context meter so a long thread never overflows unnoticed.
Models on demand
Download only the models you want. Downloads resume after an interruption, are checked against the publisher's checksum, and stay on disk until you delete them.
Every layer on the GPU
llama.cpp on Metal with the whole model resident on Apple Silicon, plus prompt caching so follow-up messages start streaming in well under a second.
Being researched next: transcription with system-audio capture and speaker labels, the two things a browser tab categorically cannot do with audio. Whatever ships is a free update for every buyer. No dates are promised; the changelog is the record.
The lineup
All Apache 2.0. Download only what you want.
32 GB MACS AND UP
Qwen3.8 27B
Flagship17.6 GB download · 32k context · Apache 2.0
The strongest model a Mac can run today. Dense 27B, far beyond what any browser tab can hold. Best for writing, code and long documents.
Qwen3.6 35B-A3B
22.4 GB download · 32k context · Apache 2.0
Mixture-of-experts: 35B of knowledge, only 3B active per token, so replies stream noticeably faster than the flagship.
16 GB MACS
Qwen3.5 9B
Best pick for 16 GB6.0 GB download · 16k context · Apache 2.0
The fastest replies in the lineup from a 6 GB download, and still a clear step up from anything that runs in a browser.
Gemma 4 12B
7.4 GB download · 16k context · Apache 2.0
Google's 12B instruct model, the alternative for 16 GB Macs. Slower to reply than Qwen3.5 9B; pick it when you want a second model to compare answers.
8 GB MACS
Bonsai 27B (1-bit)
3.8 GB download · 8k context · Apache 2.0
A 27B model squeezed to under 4 GB with 1-bit weights. Runs on 8 GB Macs; trades some accuracy for size and speed.
Qwen3.5 2B
Instant1.3 GB download · 8k context · Apache 2.0
Small and instant. Downloads in a couple of minutes and answers fast; the right first model while a bigger one downloads.
The app checks your Mac's memory and marks which tiers fit before you download anything. Every model here was chosen because its weights may be redistributed commercially, so there is no license surprise later.
Pricing
Pay once. That is the whole model.
$29
one payment, forever.Founding users: $14 during launch week, offered once.
- Every tool in the app, today and in every future version
- Every model in the lineup, downloaded on demand
- One license key for up to 3 Macs
- Free lifetime updates, no renewal, no subscription
- 30-day money-back guarantee, no questions asked
Payments are handled by Polar as merchant of record, so VAT and sales tax are sorted at checkout. Refunds within 30 days come from your purchase page, no email required.
The only network calls it makes
Three. Listed in full, so you can hold us to it.
- 1
License activation and checks
When you enter your key, and then at most once a day while the app is open, it asks Polar (our payment provider) whether the key is still valid. The request carries the key, a device label (your Mac's name and chip) and a hashed hardware id, never the raw one. Chatting with a model already on your Mac never checks the license.
- 2
Update check
On launch the app asks whether a newer version exists so you get every update. It sends the version you are running and nothing about how you use the app.
- 3
Model downloads
When you choose a model, the weights download to your Mac once, verified against a checksum. Only the file you asked for is fetched.
Nothing else. No analytics, no crash reports, no telemetry of any kind. Your chats are stored on your Mac and never sent anywhere. Turn Wi-Fi off after a model has downloaded and everything keeps working, which is a test no cloud AI can pass.
Requirements
- Chip
- Apple Silicon (M1 or later)
- macOS
- macOS 14 Sonoma or later
- Memory
- 8 GB minimum, 16 GB recommended, 32 GB for the 27B flagship
- Disk
- The app is small; each model is its download size, from 1.3 GB to 22.4 GB
Intel Macs are not supported: these models need unified memory and Metal to run at a usable speed.
Frequently asked questions
- Is there a free trial?
- No. The free tier is pocketweb.tools itself: the same tools, in your browser, with the models a browser can run. The Mac app is for the models it cannot. If it is not for you, ask for a refund within 30 days and you get every cent back, no questions asked.
- What does free lifetime updates mean?
- Every future version of the app, including new models and new tools, at no extra cost, with no renewal. It is a promise about price, not about timing: we do not commit to a release schedule. The changelog shows what has shipped and when.
- Which Macs does it run on?
- Any Apple Silicon Mac (M1 or later) on macOS 14 Sonoma or later. Intel Macs are not supported because these models do not run usefully on them. 8 GB of memory runs the smallest tier, 16 GB is the sweet spot, and the 27B flagship needs 32 GB.
- Does it need an internet connection?
- Only for three things: activating and re-checking your license, checking for updates, and downloading a model. Everything else, including every chat, runs on your Mac and works with Wi-Fi off.
- How many Macs can I use it on?
- One key activates up to 3 Macs. You can deactivate a Mac from inside the app or from your purchase page to free a slot for another one.
- Where do my chats go?
- Nowhere. Chats are stored on your Mac and never sent to us or anyone else. The app has no analytics, no crash reporting and no telemetry of any kind. The privacy section above lists the only network calls it ever makes.
- Is it a native Mac app?
- The engine is native: a Rust core running llama.cpp on Metal, with every layer of the model on the GPU. The interface is the same one you use on pocketweb.tools, shown in a system web view. That is what keeps the app small and lets web and Mac share one design.
- Why not the Mac App Store?
- The app is signed and notarized by Apple like any other download, but sold directly. That keeps the price at $29 instead of paying a store commission, and it means multi-gigabyte model files are not squeezed through the store's sandbox rules.
- Can I use the models commercially?
- Yes. Every model in the lineup is released under the Apache 2.0 license, and we add no restrictions of our own to what you write with them.
Track record
Every release and model refresh, dated.
- Qwen3.5 9B becomes the recommended model for 16 GB Macs
- Markdown replies, prompt caching, turn actions
- First build: local chat with six models
