Model launches

Voxtral TTS: lightweight 4B model rivals ElevenLabs, adds emotion steering

Voxtral TTS, a 4-billion-parameter text-to-speech model, outperformed ElevenLabs Flash v2.5 in human preference tests for naturalness and matched ElevenLabs v3 quality while offering emotion-aware control.

Voxtral TTS: lightweight 4B model rivals ElevenLabs, adds emotion steering

A new entrant in the speech-synthesis space, Voxtral TTS is a 4-billion-parameter (4B) model designed for production use and optimized for natural-sounding speech.

Human preference tests and quality comparison

In side-by-side human evaluations against ElevenLabs Flash v2.5, Voxtral TTS was preferred for naturalness: it won 58.3% of flagship voice preference tests. For voice customization, Voxtral achieved a 68.4% win rate compared with Flash v2.5.

The developers state that Voxtral matches the quality of ElevenLabs v3 while adding emotion-aware control, allowing output to be shaped by contextual cues (for example neutral, happy, sarcastic, and other tones) so that the speech sounds more considered and less robotic.

Performance and integration

The model is reported to have about 70 ms model latency and an approximate ~9.7× real-time factor. It supports native streaming and can be integrated into existing speech-to-text (STT) and large language model (LLM) stacks.

Voice cloning and adaptation

Voxtral can clone a voice from three seconds of audio, adapting tone, personality, rhythm, and intonation. The cloning is described as zero-shot and requires no fine-tuning.

Licensing and deployment

The model weights are released under the CC BY-NC 4.0 license, allowing deployment on private infrastructure and extension to custom voice libraries, subject to the license’s non-commercial terms.

Summary

Positioned as a lightweight production-ready TTS, Voxtral claims more natural output than ElevenLabs Flash v2.5, v3-level quality with emotion steering, low latency, and easy integration. The open weights and short-sample voice cloning may appeal to users who want to run and extend speech models on their own infrastructure.