An Experiment with OpenMOSS
A one scene audiobook sample from a new coming of age romance novel
I had started first with Qwen3-TTS which provided exquisite audio without the ability to control pacing or pronunciation with tags. I had started first with a 15s audio sample with a transcript and it was good. You can hear it here:
An Experiment with Qwen3-TTS
With the assistance of Claude, I installed Qwen3-TTS onto my NVIDIA DGX Spark with the intention of creating a voice clone to use for future audiobooks. The cloning was a success; the use of Qwen3-TTS text to speech less so, as it slightly mispronounced the name “Mai Lin” and there is no mechanism currently within Qwen3-TTS for pronunciations or for voi…
The audio in the article above was done with a 15s audio sample and an associated transcript. I did another rendering of the text from the scene with a 45s sample and transcript, and that longer sample improved the output audio pacing. You can listen here:
OpenMOSS supports pronunciation and pacing/emotion tags, so I first gave it a go without any controls, just to see how the base model performs versus Qwen3-TTS and I am happy that it, too, provides exquisite audio, and even better pacing out of the model without any controls whatsoever. Here is that raw audio, unprocessed, no controls. “Mai Lin” is still mis-pronounced in the same manner. The pacing is superior, though it resulted in a reading that is one minute and thirty seconds longer, but I like it better.
OpenMOSS, however, is about 4x slower than Qwen3-TTS on my DGX Spark, though there are some optimizations possible that I have not yet implemented.
I will follow up at a later time to try out emotion, pacing, and pronunciation tags.


