Skip to content
PocketWebToolsPocketWebTools
PocketWebTools for Mac

Qwen3-TTS 1.7B: RAM, speed and settings on a Mac

Qwen3-TTS 1.7B turns text into speech in ten languages and can copy a voice from a few seconds of audio. It is a 1.5 GB download in two files. Below is what we measured running it in PocketWebTools for Mac on a 16 GB M1 Pro, and the settings the app uses.

Apache 2.0Measured on a 16 GB M1 ProUpdated
1.5GB
Download, two files
2.0 to 3.4GB
Memory while speaking
1.2 to 1.5xrealtime
Speech speed
10languages
From one model

Which Mac runs it

Mac memory
8 GB and up

The app files it under 8 GB Macs. Every number here is from a 16 GB Mac.

Model file
1.04 GB

The 1.7B backbone that writes the audio frames, 4‑bit.

Companion file
0.45 GB

The speaker encoder that reads a reference clip, and the codec that turns frames into sound. The model does nothing without it.

Memory while speaking
2.0 to 3.4 GB

Peak for the whole process across our runs.

Alongside a chat model
Works

With Qwen3.5 2B loaded in the chat screen, speech ran at the same speed.

Qwen3-TTS speed on an M1 Pro

Measured in the app's own engine on a 16 GB M1 Pro.

One line, 85 characters
7.4 s of speech in 6.2 s

1.2 times realtime, with the first audio ready after 0.74 s.

A paragraph, 417 characters
20.1 s of speech in 13.8 s

1.45 times realtime. Longer text runs faster per second because the fixed start-up cost is spread out.

Three paragraphs, in the app
30.7 s of speech in 21.4 s

1.44 times realtime with the built-in voice.

Cloned voice, in the app
26.7 s of speech in 17.9 s

1.49 times realtime, from a 5-second reference clip. Cloning costs no speed.

Spanish, German, Japanese
1.31, 1.33 and 1.32 times realtime
Does it say what you wrote
Yes, checked by transcribing it back

Every output was transcribed with Parakeet and compared with the input, a Dickens paragraph word for word and Spanish and German with accents intact. Unusual names are where it slips.

Settings the app runs Qwen3-TTS with

Files
Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf + mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0.gguf

The ggml-org GGUF pair, pinned to one revision and checked against its SHA-256 after download.

Context window
4,096 tokens
Batch size / threads
512 / 8
Top-k / top-p
40 / 0.95
GPU
On, Metal
Output
24,000 Hz mono
Suppressed tokens
1,023, applied as a logit bias

The file lists the tokens the backbone must never emit. llama.cpp's own tools mask them for you; if you call the library directly you have to do it yourself, or the output is noise.

Long text
Split into chunks of about 400 characters

Sentences are kept whole. The cache is cleared before each chunk; without that the second chunk fails.

Built-in voice
A 6.4-second reference clip

The Base model has no fixed default speaker. Without a reference each chunk picked its own loudness, and one paragraph came out 19 dB quieter than the next.

What the app does with Qwen3-TTS

Speak

Paste or type any amount of text and hear it read aloud. Blank lines become pauses between paragraphs. Save the result as M4A or WAV.

Clone a voice

Add a few seconds of audio, a Voice Memo is enough, and the model reads your text in that voice. The app asks you to confirm you have the speaker's permission.

Saved voices

Voices you add stay on your Mac and are there the next time you open the app. Nothing is uploaded.

Where it falls short

  • It is faster than realtime, not instant: at 1.4 times realtime an hour of narration takes about 40 minutes.
  • How close a cloned voice sounds to the original is a judgement by ear. We measured that the words are right, not the likeness.
  • We timed English, Spanish, German and Japanese. The other six languages the model supports run through the same path but have no timing of ours.
  • A recording can end with about two seconds of near silence after the last sentence. It is the model's own output and the app does not trim it yet.

Run Qwen3-TTS in PocketWebTools for Mac

One app for chat, transcription, documents, voices, photos and video, with every model on your Mac. $99 once, free lifetime updates.

See everything in the Mac app

Qwen3-TTS questions

How much RAM does Qwen3-TTS need?
Between 2.0 and 3.4 GB while it is speaking, measured as the peak for the whole process on an M1 Pro. The two files are 1.5 GB on disk together. PocketWebTools for Mac files it under 8 GB Macs.
How fast is Qwen3-TTS on an M1 Pro?
1.2 times realtime on a single line and 1.45 times on a paragraph: 20.1 seconds of speech took 13.8 seconds to make. The first audio is ready 0.74 seconds after you start. A cloned voice runs at the same speed.
Does Qwen3-TTS run on llama.cpp?
Yes. It runs through llama.cpp's audio generation path with two GGUF files: the 1.7B model and a companion holding the speaker encoder and the codec. The app runs both on Metal.
How much audio does voice cloning need?
A few seconds. Our tests cloned from a 5-second clip, and the speech came out at the same speed as the built-in voice.
Which languages does Qwen3-TTS speak?
Ten: English, Chinese, Japanese, Korean, German, French, Spanish, Italian, Portuguese and Russian, all from the one model.
Can I use Qwen3-TTS commercially?
Yes. Qwen3-TTS is released under Apache 2.0, which allows commercial use. The app ships the license text and adds no restrictions of its own. Cloning someone's voice needs their permission whatever the license says.