Best zero-shot TTS models" or "MELLE vs VALL-E comparison"
Text-to-speech (TTS) has come a long way. In the past few years, we've witnessed a revolution in how machines generate human-like speech. Three models have stood out in the zero-shot TTS space—where a system can clone a
Text-to-speech (TTS) has come a long way. In the past few years, we've witnessed a revolution in how machines generate human-like speech. Three models have stood out in the zero-shot TTS space—where a system can clone a voice it has never heard before using just a few seconds of audio.
VALL-E (2023) pioneered the use of neural codec language models for TTS, treating speech synthesis as a language modeling task. VALL-E 2 (2024) improved upon its predecessor with smarter sampling and faster inference, achieving human parity for the first time. And now MELLE (2024) takes a radically different approach—eliminating vector quantization entirely and predicting continuous mel-spectrograms directly.
The question isn't which one is "best," but which approach makes the most sense for your use case. Let's dive in.
What is VALL-E?
VALL-E, introduced by Microsoft Research in January 2023, was a breakthrough in zero-shot TTS. It treats TTS as a conditional language modeling task—predicting discrete audio codes just like predicting words in a sentence.
How it works:
-
Uses a neural audio codec (like EnCodec) to convert speech into discrete codes (vector-quantized tokens)
-
Trains an autoregressive (AR) language model to generate coarse codes
-
Uses a non-autoregressive (NAR) model to generate the remaining fine codes
-
Trained on 60,000 hours of English speech
Key strengths:
-
Can synthesize speech with just a 3-second recording of an unseen speaker
-
Preserves the speaker's emotion and acoustic environment
-
Significantly outperformed previous state-of-the-art zero-shot TTS systems
Key weaknesses:
-
Stability issues: Random sampling can cause instability, while nucleus sampling with small top-p values may create infinite loops
-
Efficiency problems: The AR architecture is bound to the high frame rate of the audio codec
-
Fidelity loss: Vector quantization sacrifices audio fidelity compared to continuous representations
-
Complex two-stage pipeline: Requires both AR and NAR models, increasing computational and storage costs
What is VALL-E 2?
VALL-E 2, released in June 2024, builds directly on VALL-E. It's the first zero-shot TTS system to achieve human parity on LibriSpeech and VCTK benchmarks.
Two major enhancements:
-
Repetition Aware Sampling: Refines the original nucleus sampling by tracking token repetition in the decoding history, stabilizing decoding and preventing infinite loops
-
Grouped Code Modeling: Organizes codec codes into groups to shorten sequence length, boosting inference speed and addressing long-sequence modeling challenges
Key strengths:
-
Human-level speech quality, naturalness, and speaker similarity
-
Handles traditionally challenging sentences (complex or repetitive phrases) well
-
Faster than the original VALL-E
Key weaknesses:
-
Still relies on discrete codec codes with inherent fidelity limitations
-
Maintains the two-stage AR + NAR architecture
-
Requires manual configuration of sampling parameters
-
High computational requirements for training and inference
What is MELLE?
MELLE, introduced in July 2024 by researchers from Microsoft and CUHK, takes a fundamentally different approach. Instead of predicting discrete codes, it directly generates continuous mel-spectrogram frames.
How it works:
-
Single-stage, decoder-only Transformer architecture
-
Directly predicts continuous-valued mel-spectrograms from text prompts
-
Uses latent sampling (inspired by VAEs) to predict mean and log-variance of a Gaussian distribution
-
Employs a spectrogram flux loss to encourage dynamic variation between frames
-
Uses a reduction factor to predict multiple frames at once for speed
Key strengths:
-
No vector quantization → higher fidelity and no fidelity loss
-
Single-stage → simpler, more efficient, faster inference
-
47.9% relative reduction in WER_H on continuation tasks vs. VALL-E
-
64.4% relative reduction on cross-sentence tasks vs. VALL-E
-
MOS of 4.20 (ground truth: 4.29) and SMOS of 4.40 (ground truth: 3.94)
-
Generates 10 seconds of speech in 5.49 seconds (1.40 seconds with r=4)
-
Outperforms VALL-E 2 in subjective metrics
Key weaknesses:
-
Quality depends on the vocoder used for final audio generation
-
Currently evaluated on English only (multi-lingual support is future work)

The evolution from VALL-E to VALL-E 2 to MELLE represents a fundamental shift in how we think about speech synthesis. VALL-E proved that language modeling could work for TTS. VALL-E 2 fixed the stability issues and achieved human parity. But MELLE asks a more fundamental question: why compress speech at all?
By eliminating vector quantization and working directly with continuous representations, MELLE achieves superior quality, faster inference, and a dramatically simpler architecture—all while matching or exceeding the performance of its more complex predecessors.
For developers and researchers looking to build or deploy zero-shot TTS systems, MELLE represents the most efficient, high-quality, and deployable option available today. And as the technology continues to evolve, the trend toward continuous, single-stage modeling seems likely to define the next generation of speech synthesis.
Last updated 2026-08-19
Frequently asked questions
What's the main difference between VALL-E and VALL-E 2 models?+
VALL-E and VALL-E 2 use discrete tokens (like words in a language) to represent speech. They compress audio into codes using vector quantization, then predict those codes autoregressively. MELLE skips the compression step entirely and predicts continuous mel-spectrograms directly—like painting a picture instead of describing it with words.
Which model sounds most natural?+
MELLE is significantly faster. It generates 10 seconds of speech in 5.49 seconds (single pass), and with a reduction factor of 4, it takes just 1.40 seconds. VALL-E and VALL-E 2 both take 7.32 seconds.
Why does MELLE avoid vector quantization?+
Vector quantization is designed for audio compression—it sacrifices fidelity to reduce file size. MELLE argues that for high-quality speech synthesis, this trade-off is unnecessary. By working directly with continuous mel-spectrograms, MELLE preserves more acoustic detail and achieves better naturalness and speaker similarity.
What is the "spectrogram flux loss" in MELLE?+
It's a novel loss function that rewards variation between consecutive speech frames. If frames are too similar, the model gets penalized. This prevents the model from generating flat, monotone, or repetitive speech—common problems in TTS that can lead to long silences or looping.
What is "latent sampling" and why does it matter?+
Inspired by variational autoencoders (VAEs), MELLE's latent sampling module automatically samples from a learned distribution unique to each input. This replaces the manual sampling strategies (like top-p sampling) that VALL-E requires. The result is adaptive, consistent sampling without human tweaking—and it significantly improves speaker similarity.
Which model is easiest to deploy?+
MELLE offers the most streamlined deployment. It's a single-stage model with no need for a separate NAR model. It doesn't require manual sampling parameter tuning. VALL-E and VALL-E 2 require more complex two-stage pipelines and careful configuration.
What are the training data requirements?+
VALL-E: 60,000 hours of English speech VALL-E 2: Similar scale (built on VALL-E's foundation) MELLE (full): 50,000 hours (LibriHeavy corpus) MELLE-limited: Just 960 hours (LibriSpeech) with competitive results
What are the ethical concerns with these models?+
All three models raise concerns about voice spoofing and impersonation. The MELLE paper explicitly acknowledges this risk, stating that real-world deployment should require user consent for voice cloning and include a synthesized speech detection model. These models are powerful—with great power comes great responsibility.
More from Blog
Best zero-shot TTS models" or "MELLE vs VALL-E comparison"
Sign up free and get $0.98 in credit — no card required. Connect your number, pick a template, and go live in minutes.