A new entrant in the speech-synthesis space, Voxtral TTS is a 4-billion-parameter (4B) model designed for production use and optimized for natural-sounding speech.
Human preference tests and quality comparison
In side-by-side human evaluations against ElevenLabs Flash v2.5, Voxtral TTS was preferred for naturalness: it won 58.3% of flagship voice preference tests. For voice customization, Voxtral achieved a 68.4% win rate compared with Flash v2.5.
The developers state that Voxtral matches the quality of ElevenLabs v3 while adding emotion-aware control, allowing output to be shaped by contextual cues (for example neutral, happy, sarcastic, and other tones) so that the speech sounds more considered and less robotic.
Performance and integration
The model is reported to have about 70 ms model latency and an approximate ~9.7× real-time factor. It supports native streaming and can be integrated into existing speech-to-text (STT) and large language model (LLM) stacks.
Voice cloning and adaptation
Voxtral can clone a voice from three seconds of audio, adapting tone, personality, rhythm, and intonation. The cloning is described as zero-shot and requires no fine-tuning.
Licensing and deployment
The model weights are released under the CC BY-NC 4.0 license, allowing deployment on private infrastructure and extension to custom voice libraries, subject to the license’s non-commercial terms.
Summary
Positioned as a lightweight production-ready TTS, Voxtral claims more natural output than ElevenLabs Flash v2.5, v3-level quality with emotion steering, low latency, and easy integration. The open weights and short-sample voice cloning may appeal to users who want to run and extend speech models on their own infrastructure.



