Qwen3-TTS 1.7B: RAM, speed and settings on a Mac
Qwen3-TTS 1.7B turns text into speech in ten languages and can copy a voice from a few seconds of audio. It is a 1.5 GB download in two files. Below is what we measured running it in PocketWebTools for Mac on a 16 GB M1 Pro, and the settings the app uses.
Which Mac runs it
- Mac memory
- 8 GB and up
The app files it under 8 GB Macs. Every number here is from a 16 GB Mac.
- Model file
- 1.04 GB
The 1.7B backbone that writes the audio frames, 4‑bit.
- Companion file
- 0.45 GB
The speaker encoder that reads a reference clip, and the codec that turns frames into sound. The model does nothing without it.
- Memory while speaking
- 2.0 to 3.4 GB
Peak for the whole process across our runs.
- Alongside a chat model
- Works
With Qwen3.5 2B loaded in the chat screen, speech ran at the same speed.
Qwen3-TTS speed on an M1 Pro
Measured in the app's own engine on a 16 GB M1 Pro.
- One line, 85 characters
- 7.4 s of speech in 6.2 s
1.2 times realtime, with the first audio ready after 0.74 s.
- A paragraph, 417 characters
- 20.1 s of speech in 13.8 s
1.45 times realtime. Longer text runs faster per second because the fixed start-up cost is spread out.
- Three paragraphs, in the app
- 30.7 s of speech in 21.4 s
1.44 times realtime with the built-in voice.
- Cloned voice, in the app
- 26.7 s of speech in 17.9 s
1.49 times realtime, from a 5-second reference clip. Cloning costs no speed.
- Spanish, German, Japanese
- 1.31, 1.33 and 1.32 times realtime
- Does it say what you wrote
- Yes, checked by transcribing it back
Every output was transcribed with Parakeet and compared with the input, a Dickens paragraph word for word and Spanish and German with accents intact. Unusual names are where it slips.
Settings the app runs Qwen3-TTS with
- Files
- Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf + mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0.gguf
The ggml-org GGUF pair, pinned to one revision and checked against its SHA-256 after download.
- Context window
- 4,096 tokens
- Batch size / threads
- 512 / 8
- Top-k / top-p
- 40 / 0.95
- GPU
- On, Metal
- Output
- 24,000 Hz mono
- Suppressed tokens
- 1,023, applied as a logit bias
The file lists the tokens the backbone must never emit. llama.cpp's own tools mask them for you; if you call the library directly you have to do it yourself, or the output is noise.
- Long text
- Split into chunks of about 400 characters
Sentences are kept whole. The cache is cleared before each chunk; without that the second chunk fails.
- Built-in voice
- A 6.4-second reference clip
The Base model has no fixed default speaker. Without a reference each chunk picked its own loudness, and one paragraph came out 19 dB quieter than the next.
What the app does with Qwen3-TTS
Speak
Paste or type any amount of text and hear it read aloud. Blank lines become pauses between paragraphs. Save the result as M4A or WAV.
Clone a voice
Add a few seconds of audio, a Voice Memo is enough, and the model reads your text in that voice. The app asks you to confirm you have the speaker's permission.
Saved voices
Voices you add stay on your Mac and are there the next time you open the app. Nothing is uploaded.
Where it falls short
- It is faster than realtime, not instant: at 1.4 times realtime an hour of narration takes about 40 minutes.
- How close a cloned voice sounds to the original is a judgement by ear. We measured that the words are right, not the likeness.
- We timed English, Spanish, German and Japanese. The other six languages the model supports run through the same path but have no timing of ours.
- A recording can end with about two seconds of near silence after the last sentence. It is the model's own output and the app does not trim it yet.
Run Qwen3-TTS in PocketWebTools for Mac
One app for chat, transcription, documents, voices, photos and video, with every model on your Mac. $99 once, free lifetime updates.
Qwen3-TTS questions
- How much RAM does Qwen3-TTS need?
- Between 2.0 and 3.4 GB while it is speaking, measured as the peak for the whole process on an M1 Pro. The two files are 1.5 GB on disk together. PocketWebTools for Mac files it under 8 GB Macs.
- How fast is Qwen3-TTS on an M1 Pro?
- 1.2 times realtime on a single line and 1.45 times on a paragraph: 20.1 seconds of speech took 13.8 seconds to make. The first audio is ready 0.74 seconds after you start. A cloned voice runs at the same speed.
- Does Qwen3-TTS run on llama.cpp?
- Yes. It runs through llama.cpp's audio generation path with two GGUF files: the 1.7B model and a companion holding the speaker encoder and the codec. The app runs both on Metal.
- How much audio does voice cloning need?
- A few seconds. Our tests cloned from a 5-second clip, and the speech came out at the same speed as the built-in voice.
- Which languages does Qwen3-TTS speak?
- Ten: English, Chinese, Japanese, Korean, German, French, Spanish, Italian, Portuguese and Russian, all from the one model.
- Can I use Qwen3-TTS commercially?
- Yes. Qwen3-TTS is released under Apache 2.0, which allows commercial use. The app ships the license text and adds no restrictions of its own. Cloning someone's voice needs their permission whatever the license says.