PocketWebToolsPocketWebTools for Mac

The AI models your browser can't run, now on your Mac.

Chat, summarize and rewrite with open models up to 27B parameters, running on your own Apple Silicon. No account, no cloud, no usage meter.

  • Runs on Apple Silicon
  • Nothing leaves your Mac
  • Pay once, free lifetime updates

$29Pay once. Founding users pay $14 during launch week.

Founding users get the launch-week price first.

What a Mac unlocks

Same privacy as the free tools. Different ceiling.

Every tool on pocketweb.tools already runs on your device. What a browser cannot do is hold a large model: a tab tops out at about 4 GB of memory, which is enough for 2B to 9B-class models and 1-bit tricks, not for a dense 27B. The Mac app uses your Mac's actual memory, so the lineup goes where a tab cannot.

In your browser (free)PocketWebTools for Mac
Model sizeUnder about 4 GBUp to 22 GB on a 32 GB Mac
Largest model in the lineupQwen3.5 2B, Bonsai 27B 1-bitQwen3.8 27B dense, Qwen3.6 35B-A3B
Memory availableAbout 4 GB per tabYour Mac's RAM, with tiers for 8, 16 and 32 GB
EngineWebGPU or WebAssemblyllama.cpp on Metal, every layer on the GPU
Where it runsOn your deviceOn your device
Works offlineAfter the model is cachedAlways, after the download
Where models liveBrowser cache, can be evictedApplication Support, until you delete them
PriceFree, unlimited$29 once, free lifetime updates
AccountNoneNone. A license key from the purchase email.

What's inside

Version 0.1.0

  • Chat with a local model

    Threads saved on your Mac, Markdown replies with highlighted code, copy and regenerate on every turn, and a context meter so a long thread never overflows unnoticed.

  • Models on demand

    Download only the models you want. Downloads resume after an interruption, are checked against the publisher's checksum, and stay on disk until you delete them.

  • Every layer on the GPU

    llama.cpp on Metal with the whole model resident on Apple Silicon, plus prompt caching so follow-up messages start streaming in well under a second.

Being researched next: transcription with system-audio capture and speaker labels, the two things a browser tab categorically cannot do with audio. Whatever ships is a free update for every buyer. No dates are promised; the changelog is the record.

The lineup

All Apache 2.0. Download only what you want.

32 GB MACS AND UP

  • Qwen3.8 27B

    Flagship

    17.6 GB download · 32k context · Apache 2.0

    The strongest model a Mac can run today. Dense 27B, far beyond what any browser tab can hold. Best for writing, code and long documents.

  • Qwen3.6 35B-A3B

    22.4 GB download · 32k context · Apache 2.0

    Mixture-of-experts: 35B of knowledge, only 3B active per token, so replies stream noticeably faster than the flagship.

16 GB MACS

  • Qwen3.5 9B

    Best pick for 16 GB

    6.0 GB download · 16k context · Apache 2.0

    The fastest replies in the lineup from a 6 GB download, and still a clear step up from anything that runs in a browser.

  • Gemma 4 12B

    7.4 GB download · 16k context · Apache 2.0

    Google's 12B instruct model, the alternative for 16 GB Macs. Slower to reply than Qwen3.5 9B; pick it when you want a second model to compare answers.

8 GB MACS

  • Bonsai 27B (1-bit)

    3.8 GB download · 8k context · Apache 2.0

    A 27B model squeezed to under 4 GB with 1-bit weights. Runs on 8 GB Macs; trades some accuracy for size and speed.

  • Qwen3.5 2B

    Instant

    1.3 GB download · 8k context · Apache 2.0

    Small and instant. Downloads in a couple of minutes and answers fast; the right first model while a bigger one downloads.

The app checks your Mac's memory and marks which tiers fit before you download anything. Every model here was chosen because its weights may be redistributed commercially, so there is no license surprise later.

Pricing

Pay once. That is the whole model.

$29

one payment, forever.Founding users: $14 during launch week, offered once.

  • Every tool in the app, today and in every future version
  • Every model in the lineup, downloaded on demand
  • One license key for up to 3 Macs
  • Free lifetime updates, no renewal, no subscription
  • 30-day money-back guarantee, no questions asked

Founding users get the launch-week price first.

Payments are handled by Polar as merchant of record, so VAT and sales tax are sorted at checkout. Refunds within 30 days come from your purchase page, no email required.

The only network calls it makes

Three. Listed in full, so you can hold us to it.

  1. 1

    License activation and checks

    When you enter your key, and then at most once a day while the app is open, it asks Polar (our payment provider) whether the key is still valid. The request carries the key, a device label (your Mac's name and chip) and a hashed hardware id, never the raw one. Chatting with a model already on your Mac never checks the license.

  2. 2

    Update check

    On launch the app asks whether a newer version exists so you get every update. It sends the version you are running and nothing about how you use the app.

  3. 3

    Model downloads

    When you choose a model, the weights download to your Mac once, verified against a checksum. Only the file you asked for is fetched.

Nothing else. No analytics, no crash reports, no telemetry of any kind. Your chats are stored on your Mac and never sent anywhere. Turn Wi-Fi off after a model has downloaded and everything keeps working, which is a test no cloud AI can pass.

Requirements

Chip
Apple Silicon (M1 or later)
macOS
macOS 14 Sonoma or later
Memory
8 GB minimum, 16 GB recommended, 32 GB for the 27B flagship
Disk
The app is small; each model is its download size, from 1.3 GB to 22.4 GB

Intel Macs are not supported: these models need unified memory and Metal to run at a usable speed.

Frequently asked questions

Is there a free trial?
No. The free tier is pocketweb.tools itself: the same tools, in your browser, with the models a browser can run. The Mac app is for the models it cannot. If it is not for you, ask for a refund within 30 days and you get every cent back, no questions asked.
What does free lifetime updates mean?
Every future version of the app, including new models and new tools, at no extra cost, with no renewal. It is a promise about price, not about timing: we do not commit to a release schedule. The changelog shows what has shipped and when.
Which Macs does it run on?
Any Apple Silicon Mac (M1 or later) on macOS 14 Sonoma or later. Intel Macs are not supported because these models do not run usefully on them. 8 GB of memory runs the smallest tier, 16 GB is the sweet spot, and the 27B flagship needs 32 GB.
Does it need an internet connection?
Only for three things: activating and re-checking your license, checking for updates, and downloading a model. Everything else, including every chat, runs on your Mac and works with Wi-Fi off.
How many Macs can I use it on?
One key activates up to 3 Macs. You can deactivate a Mac from inside the app or from your purchase page to free a slot for another one.
Where do my chats go?
Nowhere. Chats are stored on your Mac and never sent to us or anyone else. The app has no analytics, no crash reporting and no telemetry of any kind. The privacy section above lists the only network calls it ever makes.
Is it a native Mac app?
The engine is native: a Rust core running llama.cpp on Metal, with every layer of the model on the GPU. The interface is the same one you use on pocketweb.tools, shown in a system web view. That is what keeps the app small and lets web and Mac share one design.
Why not the Mac App Store?
The app is signed and notarized by Apple like any other download, but sold directly. That keeps the price at $29 instead of paying a store commission, and it means multi-gigabyte model files are not squeezed through the store's sandbox rules.
Can I use the models commercially?
Yes. Every model in the lineup is released under the Apache 2.0 license, and we add no restrictions of our own to what you write with them.

Track record

Every release and model refresh, dated.

  • Qwen3.5 9B becomes the recommended model for 16 GB Macs
  • Markdown replies, prompt caching, turn actions
  • First build: local chat with six models
Read the full changelog